Method, apparatus, electronic device and storage medium for processing comments
By determining the correlation between comments and video clips in video playback technology and performing aggregation processing, the problem of low interaction enthusiasm for users when viewing comments at different video playback locations is solved, more relevant comment display is achieved, and user participation is improved.
Patent Information
- Application Number
- CN202110925325.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-08-12
AI Technical Summary
In existing video playback technology, when users trigger comment viewing operations at different video playback locations, the comments displayed are always the same (such as the most popular comments), resulting in a decrease in user enthusiasm for participating in video comments.
By obtaining multiple video clips and comments in the video, the correlation between the comments and each video clip is determined separately, and the comments are aggregated according to the correlation degree to obtain target comments whose correlation degree corresponds to each video clip meets the conditions. When playing a target video clip, display the target comments related to the video clip.
It realizes aggregation of comments based on the degree of correlation between comments and video clips, which enhances users' understanding of video content and interactive enthusiasm for comments.
Smart Images

Figure CN115706833B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of information processing and artificial intelligence, and particularly relates to a method, apparatus, electronic device, and storage medium for processing comments. Background Art
[0002] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0003] With the continuous development of artificial intelligence technology, artificial intelligence technology has gradually been applied to video playback. In related technologies, during video playback, the comments that users see are sorted and displayed based on the popularity, quality, etc. of the comments, so that when users trigger the operation of viewing comments at different video playback positions, the comments that can be displayed each time are the same (such as the most popular comment), which reduces the enthusiasm of users to participate in the interaction of video comments. Summary of the Invention
[0004] Embodiments of the present application provide a method, apparatus, electronic device, and storage medium for processing comments, which can achieve the effect of aggregating comments related to each video segment according to the relevance degree of the comments and the video segments in content, so as to display the aggregated comments related to the video segments when playing each video segment.
[0005] The technical solution of the embodiments of the present application is implemented as follows:
[0006] Embodiments of the present application provide a method for processing comments, including:
[0007] Obtain at least two video segments included in the video, and multiple comments of the video;
[0008] Respectively determine the relevance degree of the comments and each video segment, where the relevance degree is used to indicate the relevance degree of the comment and the video segment in content;
[0009] Aggregate the multiple comments according to the relevance degree of the comments and each video segment to obtain target comments corresponding to each video segment and satisfying the relevance degree condition;
[0010] When playing a target video segment among the at least two video segments, display a target comment corresponding to the target video segment.
[0011] In the above solution, the method further includes:
[0012] During the process of playing the video through the media play interface, present a comment viewing function item in the media play interface;
[0013] When playing a target video segment among the at least two video segments, in response to a trigger operation on the comment viewing function item, display a target comment corresponding to the target video segment.
[0014] An embodiment of the present application further provides a comment processing device, including:
[0015] An acquisition module, configured to acquire at least two video segments included in a video, and multiple comments of the video;
[0016] A determination module, configured to respectively determine the relevance between the comment and each video segment, where the relevance is used to indicate the degree of relevance between the comment and the video segment in terms of content;
[0017] An aggregation module, configured to aggregate the multiple comments according to the relevance between the comment and each video segment, to obtain a target comment whose relevance corresponding to each video segment meets the relevance condition;
[0018] A display module, configured to display a target comment corresponding to the target video segment when playing a target video segment among the at least two video segments.
[0019] In the above solution, the determination module is further configured to respectively perform the following processing for each of the video segments to determine the relevance between the comment and each video segment:
[0020] Acquire video frame images included in the video segment, video texts associated with the video segment, and comment texts of the comment;
[0021] Input the video frame images included in the video segment, the video texts associated with the video segment, and the comment texts into a first machine learning model for relevance prediction, to obtain the relevance between the comment and the video segment.
[0022] In the above solution, the determination module is further configured to respectively perform encoding processing on the video frame images included in the video segment, the video texts associated with the video segment, and the comment texts through a feature encoding layer of the first machine learning model, to obtain corresponding encoded features;
[0023] Through the feature interaction layer of the first machine learning model, perform feature interaction processing on the encoded features of the video text and the encoded features of the comment text to obtain first interaction features, and
[0024] perform feature interaction processing on the encoded features of the video frame image and the encoded features of the comment text to obtain second interaction features;
[0025] Through the feature concatenation layer of the first machine learning model, perform concatenation processing on the first interaction features and the second interaction features to obtain concatenated features;
[0026] Through the feature prediction layer of the first machine learning model, perform relevance prediction on the concatenated features to obtain the relevance between the comment and the video segment.
[0027] In the above solution, the aggregation module is further configured to perform the following processing for each of the video segments:
[0028] Select comments with a relevance reaching the relevance threshold from the multiple comments as candidate comments for the video segment, and the relevance distribution score is used to indicate the distribution of the relevance between the corresponding candidate comments and each video segment;
[0029] Determine the relevance distribution score corresponding to the candidate comment based on the relevance between the candidate comment and each video segment;
[0030] Select candidate comments with a relevance distribution score lower than the score threshold from the candidate comments as target comments whose relevance to the video segment meets the relevance condition.
[0031] In the above solution, when the number of the target comments is at least two, the display module is further configured to, when playing the target video segment among the at least two video segments, obtain the target comments corresponding to the target video segment;
[0032] Determine the sorting score corresponding to each of the target comments, and the sorting score is used to indicate the expectation degree of the viewing user of the target comment for the corresponding target comment;
[0033] Sort the target comments in descending order based on the sorting score to obtain the sorted target comments, and display the sorted target comments.
[0034] In the above solution, the display module is further configured to perform the following processing for each of the target comments to determine the sorting score corresponding to each of the target comments:
[0035] Obtain the first interest tags of the viewing users of the target video segment, the second interest tags of the publishing users of the target comment, the comment text of the target comment, and the interaction features associated with the target comment;
[0036] Input the first interest tags, the second interest tags, the comment text, and the interaction features into a second machine learning model for sorting score prediction to obtain the sorting score corresponding to the target comment.
[0037] In the above solution, when the number of the interaction features is at least two, the display module is further configured to respectively encode the first interest tags, the second interest tags, and the comment text through the feature encoding layer of the second machine learning model to obtain corresponding encoded features;
[0038] Concatenate the at least two interaction features through the first feature connection layer of the second machine learning model to obtain a first concatenated feature;
[0039] Perform feature interaction processing on the encoded features of the first interest tags and the encoded features of the second interest tags through the feature interaction layer of the second machine learning model to obtain a first interaction feature, and perform feature interaction processing on the encoded features of the first interest tags and the encoded features of the comment text to obtain a second interaction feature;
[0040] Concatenate the first interaction feature, the second interaction feature, and the first concatenated feature through the second feature connection layer of the second machine learning model to obtain a second concatenated feature;
[0041] Perform sorting score prediction on the second concatenated feature through the feature prediction layer of the second machine learning model to obtain the sorting score corresponding to the target comment.
[0042] In the above solution, the display module is further configured to respectively perform the following processing on each of the target comments to determine the sorting score corresponding to each of the target comments:
[0043] Respectively determine the sub-sorting scores of the target comment under each of the sorting rules according to a plurality of preset sorting rules;
[0044] Obtain the weight values corresponding to each of the sorting rules;
[0045] Based on the sub-sorting scores of the target comment under each of the sorting rules and the weight values corresponding to each of the sorting rules, determine the sorting score corresponding to the target comment.
[0046] In the above solution, the display module is further configured to present an information flow card for displaying comments on the video during the playback of the video;
[0047] When playing the target video segment among the at least two video segments, the information flow card is used to display the target comment corresponding to the target video segment.
[0048] In the above solution, the display module is further configured to present a comment display area for displaying comments of the video during the process of playing the video.
[0049] When playing the target video segment among the at least two video segments, the target comment corresponding to the target video segment is displayed through the comment display area.
[0050] In the above solution, the display module is further configured to, when the number of the target comments is multiple and the number of the target comments reaches the target number, present a scroll bar corresponding to the multiple target comments on one side of the comment display area.
[0051] When a trigger operation for the scroll bar is received, the display position of the target comment in the comment display area is adjusted to adjust the target comment displayed through the comment display area.
[0052] In the above solution, the display module is further configured to display the comments of the video in a bullet screen scrolling display manner during the process of playing the video.
[0053] When playing the target video segment among the at least two video segments, the target comment corresponding to the target video segment is displayed in a bullet screen scrolling display manner.
[0054] In the above solution, the display module is further configured to present a comment viewing function item in the media play interface during the process of playing the video through the media play interface.
[0055] When playing the target video segment among the at least two video segments, in response to a trigger operation for the comment viewing function item, the target comment corresponding to the target video segment is displayed.
[0056] An embodiment of the present application further provides an electronic device, including:
[0057] A memory for storing executable instructions;
[0058] A processor, configured to implement the comment processing method provided by the embodiment of the present application when executing the executable instructions stored in the memory.
[0059] An embodiment of the present application further provides a computer-readable storage medium, storing executable instructions, where when the executable instructions are executed by a processor, the comment processing method provided by the embodiment of the present application is implemented.
[0060] The embodiments of the present application have the following beneficial effects:
[0061] Obtain multiple video segments and multiple comments included in the video, and respectively determine the relevance between the comments and each video segment. Since the relevance is used to indicate the degree of relevance in content between the comment and the video segment, multiple comments can be aggregated according to the relevance, so as to obtain target comments whose relevance to each video segment meets the relevance condition. In this way, the effect of aggregating relevant comments for each video segment according to the degree of relevance in content between the comment and the video segment is achieved. When playing the target video segment, the aggregated target comments corresponding to the target video segment can be displayed. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a schematic structural diagram of a comment processing system 100 provided by an embodiment of the present application;
[0063] Figure 2 is a schematic flowchart of a comment processing method provided by an embodiment of the present application;
[0064] Figure 3 is a schematic diagram of implementing relevance prediction through a first machine learning model provided by an embodiment of the present application;
[0065] Figure 4 is a schematic diagram of implementing sorting score prediction through a second machine learning model provided by an embodiment of the present application;
[0066] Figure 5 is a schematic diagram of displaying target comments provided by an embodiment of the present application;
[0067] Figure 6 is a schematic diagram of displaying target comments provided by an embodiment of the present application;
[0068] Figure 7 is a schematic diagram of displaying target comments provided by an embodiment of the present application;
[0069] Figure 8 is a schematic diagram of displaying target comments provided by an embodiment of the present application;
[0070] Figure 9 is a schematic flowchart of a comment processing method provided by an embodiment of the present application;
[0071] Figure 10 is a schematic structural diagram of a comment processing device 600 provided by an embodiment of the present application;
[0072] Figure 11 is a schematic structural diagram of an electronic device 500 provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail in conjunction with the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0074] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0075] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0077] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0078] 1) Client: An application program running on a terminal for providing various services, such as an instant messaging client or a video client.
[0079] 2) In response to: Used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more of the executed operations can be real-time or can have a set delay; without special instructions, there is no limitation on the execution order of the multiple executed operations.
[0080] Based on the above explanations of the nouns and terms involved in the embodiments of this application, the following describes the comment processing system provided by the embodiments of this application. See Figure 1 , Figure 1FIG. 0 is a schematic architecture diagram of a comment processing system 100 provided by an embodiment of the present application. To support an exemplary application, a terminal 400 is connected to a server 200 through a network 300. The network 300 can be a wide area network, a local area network, or a combination of the two, and uses a wireless or wired link to implement data transmission. In practical applications, the terminal 400 can be provided with a client for video playback, such as a video client, an instant messaging client, a browser client, an information flow client, etc.
[0081] The terminal 400 is configured to send a retrieval request for at least two video segments included in a video and multiple comments of the video to the server 200;
[0082] The server 200 is configured to, upon receiving the retrieval request, return at least two video segments included in the video and multiple comments of the video to the terminal 400;
[0083] The terminal 400 is configured to receive at least two video segments included in the video and multiple comments of the video; respectively determine the relevance between the comments and each video segment; aggregate the multiple comments according to the relevance between the comments and each video segment to obtain target comments corresponding to each video segment and meeting the relevance condition; and when playing a target video segment among the at least two video segments, display the target comments corresponding to the target video segment.
[0084] In practical applications, the server 200 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, etc., but is not limited thereto. The terminal 400 and the server 200 can be directly or indirectly connected through a wired or wireless communication method, which is not limited in this application.
[0085] Based on the above description of the comment processing system and electronic device provided by the embodiments of the present application, the comment processing method provided by the embodiments of the present application will be described below. In some embodiments, the comment processing method provided by the embodiments of the present application can be implemented independently by the server or the terminal, or jointly implemented by the server and the terminal. The comment processing method provided by the embodiments of the present application will be described below taking the terminal as an example. Refer to Figure 2 , Figure 2 FIG. 19 is a schematic flowchart of the comment processing method provided by the embodiments of the present application. The comment processing method provided by the embodiments of the present application includes:
[0086] Step 101: The terminal obtains at least two video segments included in the video, as well as multiple comments on the video.
[0087] Here, the terminal is provided with a client, such as a video client. The terminal responds to the running instruction for the client and runs the provided client to play the video through the running client. Specifically, when the terminal plays the video, it needs to obtain the video data of the video. Here, the video data may include the content data of multiple video segments in the video, or may also include the comment data of multiple comments associated with the video, etc. The terminal sends a request for obtaining video data to the server to receive the video data returned by the server based on the obtaining request, so as to obtain at least two video segments included in the video, as well as multiple comments on the video. The multiple comments are comment contents posted by users when watching the video (such as the video content, the plot of the video, characters, etc.).
[0088] Step 102: Determine the relevance between the comments and each video segment respectively.
[0089] Among them, the relevance is used to indicate the degree of relevance between the comment and the video segment in terms of content.
[0090] In the implementation of this application, after the terminal obtains at least two video segments included in the video and multiple comments on the video, it is also necessary to determine the relevance between the comments and each video segment. The relevance is used to indicate the degree of relevance between the comment and the video segment in terms of content, so that when playing each video segment, relevant comments can be displayed to the user to enhance the user's understanding and interest in the comments and improve the user's interaction with the comments.
[0091] In some embodiments, the terminal can perform the following processing for each video segment respectively to determine the relevance between the comment and each video segment: obtain the video frame images included in the video segment, the video text associated with the video segment, and the comment text of the comment; input the video frame images included in the video segment, the video text associated with the video segment, and the comment text into the first machine learning model for relevance prediction to obtain the relevance between the comment and the video segment.
[0092] Here, the terminal can calculate the relevance between comments and each video segment through a machine learning model (i.e., the above-mentioned first machine learning model). In actual implementation, the first machine learning model is constructed through a machine learning network, such as Convolutional Neural Networks (CNN), Deep Neural Networks (DNN), etc., and initial model parameters, the activation function of the model (such as the rectified linear unit function, etc.), and the loss function of the model (such as the cross-entropy loss function, logarithmic loss function, etc.) are set for the first machine learning model. The first machine learning model can be a classification model or a regression model. After the initial first machine learning model is constructed, the constructed first machine learning model is trained.
[0093] Specifically, training samples for training the first machine learning model can be obtained, and the training samples are labeled with corresponding sample labels; then the training samples are input into the constructed initial first machine learning model for prediction, and corresponding prediction results are output; finally, the difference between the prediction results and the sample labels of the corresponding training samples is obtained, so as to update the model parameters of the first machine learning model based on the difference. Specifically, the value of the loss function of the first machine learning model can be determined based on the difference. When the value of the loss function exceeds a preset threshold, the error signal of the recommendation model is determined based on the loss function, the error signal is backpropagated in the first machine learning model, and the model parameters of each layer are updated during the propagation process. The above steps are continuously iterated throughout the training process until the loss function of the first machine learning model converges, and the updated model parameters of the first machine learning model when the loss function converges are used as the model parameters of the finally trained first machine learning model to reduce the error of the model output.
[0094] After the first machine learning model is trained, the terminal can perform the following processing for each video segment respectively to determine the relevance between the comment and each video segment: First, obtain the video frame images included in the video segment, the video text associated with the video segment (such as the line information, introduction information, etc. of the video segment), and the comment text of the comment; then input the video frame images included in the video segment, the video text associated with the video segment, and the comment text into the first machine learning model, so as to perform relevance prediction through the first machine learning model to obtain the relevance between the comment and the video segment.
[0095] In some embodiments, the terminal may input the video frame images included in the video clip, the video text associated with the video clip, and the comment text into the first machine learning model for relevance prediction in the following manner to obtain the relevance between the comment and the video clip: through the feature encoding layer of the first machine learning model, encode the video frame images included in the video clip, the video text associated with the video clip, and the comment text respectively to obtain the corresponding encoded features; through the feature interaction layer of the first machine learning model, perform feature interaction processing on the encoded features of the video text and the encoded features of the comment text to obtain the first interaction feature, and perform feature interaction processing on the encoded features of the video frame images and the encoded features of the comment text to obtain the second interaction feature; through the feature concatenation layer of the first machine learning model, concatenate the first interaction feature and the second interaction feature to obtain the concatenated feature; through the feature prediction layer of the first machine learning model, perform relevance prediction on the concatenated feature to obtain the relevance between the comment and the video clip.
[0096] See Figure 3 , Figure 3 is a schematic diagram of implementing relevance prediction through the first machine learning model provided by an embodiment of the present application. In practical applications, the structure of the first machine learning model involved in the above embodiments is as Figure 3 shown. The first machine learning model includes a feature encoding layer 310, a feature interaction layer 320, a feature concatenation layer 330, and a feature prediction layer 340. Among them, the feature encoding layer for the video text and the comment text can be implemented based on the Language Representation Model (A Lite Bidirectional Encoder Representations from Transformers, ALBERT), and the feature encoding layer for the video frame images can be implemented based on the Deep Residual Network (Residual Network, ResNet) and the Self-Attention mechanism.
[0097] The processing flow for calculating the relevance between the comment and each video clip through the first machine learning model is as follows: input the video frame images included in the video clip, the video text associated with the video clip, and the comment text into the feature encoding layer of the first machine learning model. Through the feature encoding layer, encode the video frame images included in the video clip, the video text associated with the video clip, and the comment text respectively to obtain the corresponding encoded features, that is, Figure 3 the encoded features shown include: the video clip text representation corresponding to the video text, the comment text representation corresponding to the comment text, and the video clip image representation corresponding to the video frame images; then through the feature interaction layer, perform feature interaction processing on the encoded features of the video text and the encoded features of the comment text to obtain the first interaction feature (that is, Figure 3The interactive representation of the video clip text and the comment text shown in FIG), and the encoding features of the video frame image and the encoding features of the comment text are subjected to feature interaction processing to obtain the second interactive feature (i.e. Figure 3 The interactive representation of the video clip image and the comment text shown in the figure); then the first interactive feature and the second interactive feature are spliced together through the feature splicing layer to obtain the splicing feature; finally, the splicing feature is predicted through the feature prediction layer to obtain the correlation between the comment and the video clip.
[0098] Step 103: according to the relevance between the comments and each video clip, a plurality of comments are aggregated to obtain target comments corresponding to each video clip whose relevance meets the relevance condition.
[0099] Here, after determining the relevance between the comments and each video clip, multiple comments on the video can be aggregated according to the relevance between the comments and each video clip, so as to obtain target comments corresponding to each video clip whose relevance meets the relevance condition. Specifically, the comments corresponding to the same video clip whose relevance meets the relevance condition can be aggregated together to obtain target comments corresponding to the video clip whose relevance meets the relevance condition. In practical applications, the target comments whose relevance meets the relevance condition can be target comments whose relevance reaches the relevance threshold, or a preset number of target comments that are ranked top after sorting in descending order of relevance, etc.
[0100] In some embodiments, the terminal may aggregate multiple comments according to the relevance between the comments and each video segment to obtain a target comment corresponding to each video segment that satisfies the relevance condition in the following manner: perform the following processing for each video segment: select comments whose relevance reaches a relevance threshold from multiple comments as candidate comments for the video segment; determine the relevance distribution score corresponding to the candidate comments based on the relevance between the candidate comments and each video segment; select candidate comments whose relevance distribution score is lower than the score threshold from the candidate comments as target comments whose relevance meets the relevance condition corresponding to the video segment. The relevance distribution score is used to indicate the distribution of the relevance between the corresponding candidate comments and each video segment.
[0101] Here, firstly, comments whose relevance reaches a relevance threshold are selected from multiple comments as candidate comments for the video clip. In actual implementation, comments whose relevance reaches the relevance threshold and whose ranking is at the top after sorting in descending order based on relevance can be further selected.
[0102] Then, based on the relevance between the candidate comments and each video segment, the relevance distribution score corresponding to the candidate comment is determined. This relevance distribution score is used to indicate the distribution of the relevance between the corresponding candidate comment and each video segment. In practical applications, a comment should have a high local relevance, that is, a comment should not be relevant to too many video segments. If a comment is relevant to too many video segments, it means that the comment is a relatively global or general comment, and it is meaningless to perform relevance aggregation on such a comment with video segments. Therefore, in the embodiments of the present application, constraints are imposed by limiting the relevance distribution between the comment and each video segment. For example, the relevance between the c-th comment of the video and the s-th video segment is P_rc[c][s], and the relevance distribution score of comment c with respect to each video segment is Sum_s(-P_rc[c][s]*log(P_rc[c][s])); here, the larger the relevance distribution score corresponding to a certain comment, the more evenly the relevance between the comment and each video segment is distributed, and the worse the local relevance of the comment is, that is, it may be a relatively global comment with a lower relevance to the video segment, and such a comment may not be aggregated. That is, it is necessary to restrict the target comment for aggregation, and the relevance distribution score of the target comment with respect to each video segment should be lower than a preset score threshold. Therefore, subsequently, candidate comments with a relevance distribution score lower than the score threshold are selected from the candidate comments as the target comments whose relevance to the video segment meets the relevance condition.
[0103] Step 104: When playing the target video segment among at least two video segments, display the target comment corresponding to the target video segment.
[0104] Here, after aggregating multiple comments according to the relevance to obtain the target comments whose relevance to each video segment meets the relevance condition, when playing the target video segment among at least two video segments, the terminal displays the target comment corresponding to the target video segment. Since the target comment is a comment whose relevance to the target video segment meets the relevance condition, it can enhance the user's understanding and interest in the comment and improve the user's interaction with the comment.
[0105] In some embodiments, when playing the target video segment among at least two video segments and when the number of target comments is at least two, the terminal can display the target comment corresponding to the target video segment in the following manner: when playing the target video segment among at least two video segments, obtain the target comment corresponding to the target video segment; determine the sorting score corresponding to each target comment, where this sorting score is used to indicate the expected degree of the viewing user of the target comment for the corresponding target comment; sort the target comments in descending order based on the sorting score to obtain the sorted target comments, and display the sorted target comments.
[0106] Here, when the terminal displays the target comments corresponding to the target video segment, if there are at least two such target comments, the terminal can also sort these multiple target comments. Specifically, when playing the target video segment among at least two video segments, the terminal can obtain the target comments corresponding to the target video segment, and then determine the sorting scores corresponding to each target comment, so as to perform a descending sort on the target comments based on the sorting scores, obtain the sorted target comments and display them. Here, the sorting score is used to indicate the expected degree of the viewing users of the target comments for the corresponding target comments, that is, the sorting score can be used to indicate the degree of interest of the viewing users in each target comment. Therefore, performing a descending sort based on the sorting scores and displaying the sorted target comments can give users a better visual experience and improve user comment participation.
[0107] In some embodiments, the terminal can perform the following processing for each target comment respectively to determine the sorting score corresponding to each target comment: obtain the first interest tag of the viewing user of the target video segment, the second interest tag of the publishing user of the target comment, the comment text of the target comment, and the interaction features associated with the target comment; input the first interest tag, the second interest tag, the comment text, and the interaction features into the second machine learning model for sorting score prediction to obtain the sorting score corresponding to the target comment.
[0108] Here, the terminal can calculate the sorting score of the target comment through a machine learning model (i.e., the above-mentioned second machine learning model). In actual implementation, the second machine learning model is constructed through a machine learning network (such as a convolutional neural network CNN, a deep neural network DNN, etc.), and initial model parameters, the activation function of the model (such as a rectified linear unit function, etc.), and the loss function of the model (such as a cross-entropy loss function, a logarithmic loss function, etc.) are set for the second machine learning model. The first machine learning model can be a classification model or a regression model. After the initial second machine learning model is constructed, the constructed second machine learning model is trained.
[0109] Specifically, training samples for training the second machine learning model can be obtained, and the training samples are labeled with corresponding sample labels; then the training samples are input into the constructed initial second machine learning model for prediction, and corresponding prediction results are output; finally, the difference between the prediction results and the sample labels of the corresponding training samples is obtained, so as to update the model parameters of the second machine learning model based on the difference. Specifically, the value of the loss function of the second machine learning model can be determined based on the difference. When the value of the loss function exceeds a preset threshold, the error signal of the recommendation model is determined based on the loss function, the error signal is backpropagated in the second machine learning model, and the model parameters of each layer are updated during the propagation process. The above steps are continuously iterated throughout the training process until the loss function of the second machine learning model converges. The updated model parameters of the second machine learning model when the loss function converges are used as the model parameters of the finally trained second machine learning model to reduce the error of the model output.
[0110] After the second machine learning model is trained, the terminal can implement the prediction of the sorting scores corresponding to each target comment based on the second machine learning module. Specifically, first, the first interest labels of the viewing users of the target video segment (such as movies and TV shows, post-00s, idols, etc.), the second interest labels of the publishing users of the target comment (such as TV dramas, post-00s, idols, etc.), the comment text of the target comment, and the interaction features associated with the target comment (such as comment like data, comment reply data, relevance between the comment and the video segment, etc.) are obtained; then the first interest labels, the second interest labels, the comment text, and the interaction features are input into the second machine learning model, and the sorting scores corresponding to the target comments are obtained through the sorting score prediction by the second machine learning model.
[0111] In some embodiments, when the number of interaction features is at least two, the terminal can input the first interest label, the second interest label, the comment text, and the interaction features into the second machine learning model in the following manner to predict the sorting score and obtain the sorting score corresponding to the target comment: Through the feature encoding layer of the second machine learning model, encode the first interest label, the second interest label, and the comment text respectively to obtain the corresponding encoded features; Through the first feature connection layer of the second machine learning model, splice at least two interaction features to obtain the first spliced feature; Through the feature interaction layer of the second machine learning model, perform feature interaction processing on the encoded features of the first interest label and the second interest label to obtain the first interaction feature, and perform feature interaction processing on the encoded features of the first interest label and the comment text to obtain the second interaction feature; Through the second feature connection layer of the second machine learning model, splice the first interaction feature, the second interaction feature, and the first spliced feature to obtain the second spliced feature; Through the feature prediction layer of the second machine learning model, perform sorting score prediction on the second spliced feature to obtain the sorting score corresponding to the target comment.
[0112] See Figure 4 , Figure 4 FIG. is a schematic diagram of implementing sorting score prediction through the second machine learning model provided by an embodiment of the present application. In practical applications, the structure of the first machine learning model involved in the above embodiment is as Figure 4 shown. The second machine learning model includes a feature encoding layer 410, a first feature connection layer 420, a feature interaction layer 430, a second feature splicing layer 440, and a feature prediction layer 450. Among them, the feature encoding layer for the first interest label, the second interest label, and the comment text can be implemented based on a language representation model (Bidirectional Encoder Representations from Transformers, BERT).
[0113] The processing flow for calculating the sorting score corresponding to the target comment through the second machine learning model is as follows: Input the first interest label, the second interest label, the comment text, and the interaction features into the feature encoding layer of the second machine learning model, and encode the first interest label, the second interest label, and the comment text respectively to obtain the corresponding encoded features; Through the first feature connection layer (i.e., Figure 4 the fully connected network 1 shown), splice at least two interaction features to obtain the first spliced feature; Through the feature interaction layer, perform feature interaction processing on the encoded features of the first interest label and the second interest label to obtain the first interaction feature (i.e., Figure 4(the interest interaction representation between the current user and the comment publisher shown), and perform feature interaction processing on the encoded features of the first interest tag and the encoded features of the comment text to obtain a second interaction feature (i.e., Figure 4 (the interaction representation between the current user and the comment shown); through a second feature connection layer, splice the first interaction feature, the second interaction feature, and the first splicing feature to obtain a second splicing feature; through a feature prediction layer, perform sorting score prediction on the second splicing feature to obtain the sorting score corresponding to the target comment.
[0114] In some embodiments, the terminal can determine the sorting scores corresponding to each target comment in the following manner: for each target comment, perform the following processing respectively to determine the sorting scores corresponding to each target comment: according to a plurality of preset sorting rules, respectively determine the sub-sorting scores of the target comment under each sorting rule; obtain the weight values corresponding to each sorting rule; based on the sub-sorting scores of the target comment under each sorting rule and the weight values corresponding to each sorting rule, determine the sorting score corresponding to the target comment.
[0115] Here, the terminal can also determine the sorting scores corresponding to each target comment based on the following manner: obtain the preset sorting rules, such as sorting according to the relevance of each target comment to the video segment, or sorting according to the number of likes / like rate of each target comment, or sorting according to the number of replies / reply rate of each target comment, or sorting according to the degree of relevance of each target comment to the interest profile of the viewing user, etc.; then according to the obtained multiple sorting rules, respectively determine the sub-rule scores of the target comment under each sorting rule. In practical applications, each sorting rule also corresponds to a corresponding weight value. After determining the sub-rule scores of the target comment under each sorting rule, obtain the weight values corresponding to each sorting rule, and then perform weighted average processing based on the sub-sorting scores of the target comment under each sorting rule and the weight values corresponding to each sorting rule to determine the sorting score corresponding to the target comment.
[0116] In some embodiments, when playing the target video segment among at least two video segments, the terminal can display the target comment corresponding to the target video segment in the following manner: during the process of playing the video, present an information flow card for displaying the comments of the video; when playing the target video segment among at least two video segments, use the information flow card to display the target comment corresponding to the target video segment.
[0117] Here, during the process of playing the video, the terminal presents an information flow card for displaying the comments of the video; when playing the target video segment among at least two video segments, then use the information flow card to display the target comment corresponding to the target video segment.
[0118] As an example, see Figure 5 ,Figure 5 This is a schematic diagram showing the display of the target comment provided by an embodiment of the present application. Here, the terminal presents an information flow card 50 for displaying comments on the video, as shown in Figure 5 Figure A; when the terminal plays the target video clip "Video Clip 2", it uses the information flow card to display the target comment corresponding to the target video clip "Video Clip 2 (describing the video content of meeting again after separation)", such as "Finally met again, so happy for them!", as shown in Figure 5 Figure B.
[0119] In some embodiments, when playing the target video clip among at least two video clips, the terminal can display the target comment corresponding to the target video clip in the following manner: during the process of playing the video, present a comment display area for displaying comments on the video; when playing the target video clip among at least two video clips, display the target comment corresponding to the target video clip through the comment display area.
[0120] Here, during the process of playing the video, the terminal presents a comment display area for displaying comments on the video; thus, when playing the target video clip among at least two video clips, the target comment corresponding to the target video clip is displayed through the comment display area.
[0121] In some embodiments, when the number of target comments is multiple and the number of target comments reaches the target number, the terminal can present a scroll bar corresponding to the multiple target comments on one side of the comment display area; when receiving a trigger operation for the scroll bar, adjust the display position of the target comments in the comment display area to adjust the target comments displayed through the comment display area.
[0122] In practical applications, when the number of target comments is multiple and the number of target comments reaches the target number, the terminal can present a scroll bar corresponding to the multiple target comments on one side of the comment display area, such as presenting a scroll bar corresponding to the multiple target comments on the right side or below the comment display area. The user can adjust the target comments displayed in the comment display area based on this scroll bar. When the terminal receives a trigger operation for the scroll bar, such as a sliding operation, a scrolling operation, a click operation, etc., it adjusts the display position of the target comments in the comment display area to adjust the target comments displayed through the comment display area.
[0123] As an example, see Figure 6 , Figure 6It is a schematic diagram showing the display of the target comment provided by the embodiment of the present application. Here, when playing the target video segment (describing the video content of meeting again after separation) among at least two video segments, the terminal displays the target comment corresponding to the target video segment through the comment display area 60, that is, "Finally met again, so happy for them!", "It's been so long, so sad!", etc., as Figure 6 shown in Figure A of
[0124] When the number of target comments reaches the target number, a scroll bar 61 corresponding to multiple target comments appears on the right side of the comment display area, as Figure 6 shown in Figure B of Figure 6 When responding to the sliding operation on the scroll bar, the display position of the target comment in the comment display area is adjusted to realize adjusting the target comment displayed through the comment display area. As
[0125] shown in Figure C of
[0126] The target comment displayed is adjusted from "Finally met again, so happy for them!" and "It's been so long, so sad!" to "Long time no see!" and "Meet, sprinkle flowers!". Figure 7 , Figure 7 It is a schematic diagram showing the display of the target comment provided by the embodiment of the present application. Here, when the terminal plays the target video segment (describing the video content of meeting again after separation), the target comment corresponding to the target video segment is displayed in a bullet screen scrolling display mode, such as "Finally met again, so happy for them!" and "It's been so long, so sad!".
[0127] In some embodiments, the terminal can display the target comment corresponding to the target segment in the following way: during the process of playing the video through the media player interface, a comment viewing function item is presented in the media player interface; when playing the target video segment among at least two video segments, in response to the trigger operation on the comment viewing function item, the target comment corresponding to the target video segment is displayed.
[0128] Here, during the process of playing a video through the media playback interface, the terminal can present a comment viewing function item in the media playback interface, so that the user can trigger a viewing operation for comments based on this comment viewing function item. In this way, during the video playback process, the display and hiding of comments can be controlled based on this comment viewing function item, thereby improving the user viewing experience, that is, the comments can be viewed through the comment viewing function item only when the user needs to view the comments, and the comments can be not displayed at other times.
[0129] When the terminal plays the target video segment, when a trigger operation for the comment viewing function item is received, in response to the trigger operation for the comment viewing function item, the target comments corresponding to the target video segment are displayed.
[0130] As an example, see Figure 8 , Figure 8 is a schematic diagram showing the display of the target comments provided in the embodiments of the present application. Here, the terminal presents a comment viewing function item "Comments" in the media playback interface, as shown in Figure 8 Figure A; when playing the target video segment (the video content describing meeting again after separation), in response to the trigger operation for the comment viewing function item "Comments", the target comments corresponding to the target video segment are displayed, such as "Finally met again, so happy for them!", "It's been so long, so sad", etc., as shown in Figure 8 Figure B.
[0131] Applying the above embodiments of the present application, multiple video segments and multiple comments included in the video are obtained, and the relevance between the comments and each video segment is determined respectively. Since the relevance is used to indicate the degree of content relevance between the comment and the video segment, the multiple comments can be aggregated according to the relevance, so as to obtain the target comments whose relevance to each video segment meets the relevance condition. In this way, the effect of aggregating the comments related to each video segment according to the content relevance between the comment and the video segment is achieved. When playing the target video segment, the aggregated target comments related to the target video segment can be displayed, enhancing the user's understanding of the video content and improving the enthusiasm of the user to participate in comment interaction.
[0132] The following will illustrate the exemplary application of the embodiments of the present application in an actual application scenario.
[0133] In the related art, the comments under the video are sorted and displayed based on the popularity, quality, etc. of the comments, so that when the user triggers a viewing operation for the comments at different video playback positions, the comments that can be displayed each time are the same (such as the most popular comments), and will not change dynamically as the user switches or plays different video contents, resulting in the displayed comments may not be very relevant to the video content currently viewed by the user, affecting the user's video comment viewing and interaction experience.
[0134] Based on this, an embodiment of the present application provides a method for processing comments to at least solve the above existing problems. Specifically, in the embodiment of the present application, through multi-dimensional in-depth understanding of the video, the comments of the video are intelligently aggregated according to relevant plot segments. When the user plays to different video playing positions, the aggregated comments related to the current plot are dynamically and preferentially displayed, thereby enhancing the user's understanding and interest in the comments and improving the user's interaction with the comments.
[0135] Here, comment aggregation means that a video can generally be segmented into multiple video segments in chronological order. The comments made by users are generally for a certain video segment. Comments related to the same video segment can be aggregated together. When the video plays to this video segment, the aggregated comments related to this video segment are preferentially displayed, so that the displayed comments are related to the playing progress of the video.
[0136] Next, a detailed description will be given of the method for processing comments provided by the embodiment of the present application. The implementation process of the comment aggregation display provided by the embodiment of the present application is as Figure 9 shown. Figure 9 is a schematic flowchart of the method for processing comments provided by the embodiment of the present application, including:
[0137] Step 201: The terminal intelligently aggregates video comments.
[0138] Here, multiple comments corresponding to the video can be aggregated based on the relevance between the comments and each video segment included in the video. Specifically, the video can be segmented into multiple video segments. For example, the video is segmented with 10 seconds as a time segment to obtain multiple video segments, and then the relevance between the comments and each video segment is calculated, so as to aggregate multiple comments based on the relevance. This relevance is used to indicate the degree of relevance between the comment and the video segment in terms of content.
[0139] In practical applications, the calculation of the relevance between the comments and each video segment can be implemented through a machine learning model. The model structure of this machine learning model is as Figure 3 shown. Figure 3 The model for calculating the relevance between the comment and the video segment shown can be trained on a constructed dataset of relevant and irrelevant video segment-comments. Based on the input image frames of the video segment, video segment text (including text recognized from images and speech), and comment text, the model can output the relevance P_rc between the comment and the video segment. In actual implementation, the training dataset can be constructed manually, or videos with relatively small differences in video length and video segment segmentation length on the platform can be mined, and their own comments are used as positive examples with high relevance, and negative examples are constructed through negative sampling.
[0140] In practical applications, to ensure that the aggregated comments obtained after aggregation have a good relevance to each video segment, the following requirements are imposed on the relevance between comments and video segments: 1) The relevance between a comment and a video segment needs to meet a preset relevance threshold, and the top S video segments that meet the relevance threshold and are the most relevant are selected as the video segments related to the comment. 2) A comment should have a high local relevance, that is, a comment should not be related to too many video segments. If a comment is related to too many video segments, it means that the comment is a relatively global or general comment, and it is meaningless to perform relevance aggregation on such a comment with video segments. Therefore, in the embodiments of the present application, constraints are imposed by limiting the relevance distribution between a comment and each video segment. For example, the relevance between the c-th comment of a video and the s-th video segment is P_rc[c][s], and the entropy of the relevance distribution between comment c and each video segment is Sum_s(-P_rc[c][s]*log(P_rc[c][s])); here, the larger the entropy, the more uniform the relevance between the comment and each segment, and the worse the local relevance of the comment. Such a comment may not be aggregated. That is, it is necessary to limit the entropy of the relevance distribution between a comment and each video segment to be lower than a preset entropy threshold.
[0141] Comments that meet the above requirements are aggregated according to their relevance to each video segment, that is, comments related to the same video segment are aggregated together. The results are shown in Table 1. Among them, the aggregated comments corresponding to segment 2 include comment 11 and comment 14; the aggregated comments corresponding to segment 7 include comment 12 and comment 22; the aggregated comments corresponding to segment 3 include comment 16, comment 13, and comment 21, and so on.
[0142] Video Comment Comment-related segment Comment-segment relevance Video 1 Comment 11 Segment 2 0.512 Video 1 Comment 16 Segment 3 0.217 Video 1 Comment 12 Segment 7 0.321 Video 1 Comment 12 Segment 10 0.462 Video 1 Comment 13 Segment 3 0.531 Video 1 Comment 14 Segment 2 0.672 Video 2 Comment 21 Segment 3 0.815 Video 2 Comment 22 Segment 7 0.663 … … … … Video v Comment vc Segment 3 0.856
[0143] Table 1
[0144] Step 202: Select and display aggregated comments based on the played target video segment.
[0145] When displaying the aggregated comments related to the video segment constructed in step 201 to the user, first, based on the user's current playback position, the target video segment of the video corresponding to the current playback position needs to be determined. Based on the target video segment, the aggregated comments related to the target video segment are selected, and then the aggregated comments related to the target video segment are sorted internally.
[0146] In practical applications, sorting can be performed using a rule fusion method. For example, the relevance between a comment and a video segment, the like / reply interaction feature of the comment, the relevance between the comment and the user's interest, etc. are weighted and fused by rules, and each comment in the aggregated comments is sorted based on the fusion score to display the sorted aggregated comments.
[0147] In practical applications, it can also be done through, for exampleFigure 4 For each comment in the aggregated comments on the target video segment by the shown machine learning model, calculate the user expectation score, and sort each comment in the aggregated comments based on the obtained user expectation score.
[0148] Figure 4 The shown machine learning model can calculate the user expectation score of the current user for each comment in the aggregated comments on the current target video segment. This model integrates the interest relevance between the current user and the comment publisher, the relevance between the current user's interest and the comment, the text representation of the comment, the interaction features of the comment, and the relevance between the comment and the video segment calculated in step 201.
[0149] Among them, the user's interest tag sequence can be statistically obtained from the user's historical valid played video sequence. Migrate the tags of the valid played videos to the user's interest tags and retain interest tags above a certain threshold. The like rate of a comment is the number of likes / the number of video plays, and the reply rate of a comment is the number of replies / the number of video plays. Through Figure 4 The shown machine learning model enables the aggregated comments to not only focus on the relevance between the comments and the target video segment, but also capture the interest degree between the current user and the comments and the content quality of the comments themselves, making the display order of the aggregated comments more in line with the current user's needs.
[0150] Step 203: Display the aggregated comments.
[0151] After the above steps 201 - 202, construct the aggregated comments for the target video segment of the current user, and after proper sorting, display the sorted aggregated comments. Specifically, it can be displayed through an information flow card, or a dynamic aggregated comment display area can be placed above the comment area and below the player. In this way, display the comments more relevant to the current playback scenario of the user, enabling the user to deepen the interest in the comments based on the current video content and enhancing the user's comment interaction.
[0152] Applying the above embodiments of the present application, when the user is watching a video, at different viewing positions, preferentially and dynamically display the aggregated comments more relevant to the current scenario, making the comments more in line with the context environment of the video playback and enhancing the product experience of comment interaction.
[0153] Next, continue to describe the comment processing device 600 provided by the embodiments of the present application. Refer to Figure 10 , Figure 10 is a schematic structural diagram of the comment processing device 600 provided by the embodiments of the present application. The comment processing device 600 provided by the embodiments of the present application includes:
[0154] An obtaining module 610, configured to obtain at least two video segments included in the video, and multiple comments on the video;
[0155] A determination module 620, configured to respectively determine the relevance between the comment and each video clip, where the relevance is used to indicate the degree of content relevance between the comment and the video clip;
[0156] An aggregation module 630, configured to aggregate the multiple comments according to the relevance between the comment and each video clip, so as to obtain a target comment corresponding to each video clip and satisfying the relevance condition;
[0157] A display module 640, configured to display the target comment corresponding to the target video clip when playing the target video clip among the at least two video clips.
[0158] In some embodiments, the determination module 620 is further configured to respectively perform the following processing for each of the video clips to determine the relevance between the comment and each video clip:
[0159] Obtain the video frame images included in the video clip, the video text associated with the video clip, and the comment text of the comment;
[0160] Input the video frame images included in the video clip, the video text associated with the video clip, and the comment text into a first machine learning model for relevance prediction, so as to obtain the relevance between the comment and the video clip.
[0161] In some embodiments, the determination module 620 is further configured to respectively perform encoding processing on the video frame images included in the video clip, the video text associated with the video clip, and the comment text through the feature encoding layer of the first machine learning model to obtain corresponding encoded features;
[0162] Perform feature interaction processing on the encoded features of the video text and the encoded features of the comment text through the feature interaction layer of the first machine learning model to obtain a first interaction feature, and
[0163] Perform feature interaction processing on the encoded features of the video frame images and the encoded features of the comment text to obtain a second interaction feature;
[0164] Perform splicing processing on the first interaction feature and the second interaction feature through the feature splicing layer of the first machine learning model to obtain a spliced feature;
[0165] Perform relevance prediction on the spliced feature through the feature prediction layer of the first machine learning model to obtain the relevance between the comment and the video clip.
[0166] In some embodiments, the aggregation module 630 is further configured to respectively perform the following processing for each of the video clips:
[0167] Select comments with a relevance reaching the relevance threshold from the multiple comments as candidate comments for the video clip;
[0168] Based on the relevance between the candidate comments and each video clip, determine the relevance distribution score corresponding to the candidate comments, where the relevance distribution score is used to indicate the distribution of the relevance between the corresponding candidate comments and each video clip;
[0169] Select candidate comments with a relevance distribution score lower than the score threshold from the candidate comments as target comments whose relevance to the video clip meets the relevance condition.
[0170] In some embodiments, when the number of the target comments is at least two, the display module 640 is further configured to, when playing a target video clip among the at least two video clips, obtain target comments corresponding to the target video clip;
[0171] Determine the sorting scores corresponding to the target comments, where the sorting scores are used to indicate the expected degree of the viewing users of the target comments for the corresponding target comments;
[0172] Based on the sorting scores, perform a descending order sorting on the target comments to obtain the sorted target comments, and display the sorted target comments.
[0173] In some embodiments, the display module 640 is further configured to perform the following processing on each of the target comments respectively to determine the sorting scores corresponding to the target comments:
[0174] Obtain the first interest tags of the viewing users of the target video clip, the second interest tags of the publishing users of the target comments, the comment texts of the target comments, and the interaction features associated with the target comments;
[0175] Input the first interest tags, the second interest tags, the comment texts, and the interaction features into a second machine learning model for sorting score prediction to obtain the sorting scores corresponding to the target comments.
[0176] In some embodiments, when the number of the interaction features is at least two, the display module 640 is further configured to respectively encode the first interest tags, the second interest tags, and the comment texts through the feature encoding layer of the second machine learning model to obtain corresponding encoded features;
[0177] Through the first feature connection layer of the second machine learning model, splice the at least two interaction features to obtain a first spliced feature;
[0178] Through the feature interaction layer of the second machine learning model, perform feature interaction processing on the encoded features of the first interest label and the encoded features of the second interest label to obtain a first interaction feature, and perform feature interaction processing on the encoded features of the first interest label and the encoded features of the review text to obtain a second interaction feature;
[0179] Through the second feature connection layer of the second machine learning model, splice the first interaction feature, the second interaction feature, and the first spliced feature to obtain a second spliced feature;
[0180] Through the feature prediction layer of the second machine learning model, perform sorting score prediction on the second spliced feature to obtain the sorting score corresponding to the target review.
[0181] In some embodiments, the display module 640 is further configured to perform the following processing on each of the target reviews respectively to determine the sorting score corresponding to each target review:
[0182] According to a plurality of preset sorting rules, respectively determine the sub-sorting scores of the target review under each sorting rule;
[0183] Obtain the weight values corresponding to the sorting rules;
[0184] Based on the sub-sorting scores of the target review under each sorting rule and the weight values corresponding to the sorting rules, determine the sorting score corresponding to the target review.
[0185] In some embodiments, the display module 640 is further configured to present an information flow card for displaying the reviews of the video during the playback of the video;
[0186] When playing the target video segment among the at least two video segments, use the information flow card to display the target reviews corresponding to the target video segment.
[0187] In some embodiments, the display module 640 is further configured to present a review display area for displaying the reviews of the video during the playback of the video;
[0188] When playing the target video segment among the at least two video segments, display the target reviews corresponding to the target video segment through the review display area.
[0189] In some embodiments, the display module 640 is further configured to, when the number of the target reviews is multiple and the number of the target reviews reaches the target number, present a scroll bar corresponding to the multiple target reviews on one side of the review display area;
[0190] When a trigger operation for the scroll bar is received, adjust the display position of the target comment in the comment display area to adjust the target comment displayed through the comment display area.
[0191] In some embodiments, the display module 640 is further configured to, during the process of playing the video, display the comments of the video in a bullet screen scrolling display manner;
[0192] When playing the target video segment among the at least two video segments, display the target comments corresponding to the target video segment in a bullet screen scrolling display manner.
[0193] In some embodiments, the display module 640 is further configured to, during the process of playing the video through the media play interface, present a comment viewing function item in the media play interface;
[0194] When playing the target video segment among the at least two video segments, in response to a trigger operation for the comment viewing function item, display the target comments corresponding to the target video segment.
[0195] Applying the above embodiments of the present application, multiple video segments and multiple comments included in the video are obtained, and the relevance between the comments and each video segment is determined respectively. Since the relevance is used to indicate the degree of relevance between the comment and the video segment in terms of content, multiple comments can be aggregated according to the relevance, so as to obtain target comments whose relevance to each video segment meets the relevance condition. In this way, the effect of aggregating the comments related to each video segment according to the degree of relevance between the comment and the video segment in terms of content is achieved. When playing the target video segment, the target comments aggregated and related to the target video segment can be displayed, enhancing the user's understanding of the video content and improving the user's enthusiasm for participating in comment interaction.
[0196] The embodiments of the present application further provide an electronic device. Refer to Figure 11 , Figure 11 which is a schematic structural diagram of the electronic device 500 provided by the embodiments of the present application. In practical applications, the electronic device 500 can be Figure 1 the terminal or server shown. Taking the electronic device 500 as Figure 1 the terminal shown as an example, the electronic device for processing comments in the method of the embodiments of the present application is described. The electronic device 500 provided by the embodiments of the present application includes:
[0197] A memory 550 for storing executable instructions;
[0198] A processor 510, configured to implement the comment processing method provided by the embodiments of the present application when executing the executable instructions stored in the memory.
[0199] Here, the processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0200] The memory 550 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 550 optionally includes one or more storage devices that are physically remote from the processor 510.
[0201] The memory 550 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), and the volatile memory can be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0202] In some embodiments, it may further include at least one network interface 520 and a user interface 530. Each component in the electronic device 500 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 11 all kinds of buses are labeled as the bus system 540.
[0203] The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the comment processing method provided by the embodiments of the present application.
[0204] The embodiments of the present application also provide a computer-readable storage medium storing executable instructions, and when the executable instructions are executed by a processor, the comment processing method provided by the embodiments of the present application is implemented.
[0205] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0206] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0207] As an example, the executable instructions may or may not correspond to a file in the file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code).
[0208] As an example, the executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected by a communication network.
[0209] As described above, the above are only embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are all included in the protection scope of the present application.
Claims
1. A method for processing comments, characterized in that, The method includes: Obtaining at least two video segments included in the video, as well as multiple comments on the video; Respectively determining the relevance between the comments and each video segment, where the relevance is used to indicate the degree of content relevance between the comment and the video segment; Aggregating the multiple comments according to the relevance between the comments and each video segment to obtain target comments corresponding to each video segment and satisfying the relevance condition; When playing a target video segment among the at least two video segments, displaying the target comments corresponding to the target video segment; where, when playing the target video segment among the at least two video segments, displaying the target comments corresponding to the target video segment includes: during the process of playing the video, presenting an information flow card for displaying the comments of the video; when playing the target video segment among the at least two video segments and receiving a trigger operation for the comment viewing function item, in response to the trigger operation for the comment viewing function item, using the information flow card to display the target comments corresponding to the target video segment.
2. The method according to claim 1, wherein The respectively determining the relevance between the comments and each video segment includes: Performing the following processing for each of the video segments respectively to determine the relevance between the comments and each video segment: Obtaining the video frame images included in the video segment, the video text associated with the video segment, and the comment text of the comment; Inputting the video frame images included in the video segment, the video text associated with the video segment, and the comment text into a first machine learning model for relevance prediction to obtain the relevance between the comment and the video segment.
3. The method according to claim 2, wherein The inputting the video frame images included in the video segment, the video text associated with the video segment, and the comment text into a first machine learning model for relevance prediction to obtain the relevance between the comment and the video segment includes: Respectively performing encoding processing on the video frame images included in the video segment, the video text associated with the video segment, and the comment text through the feature encoding layer of the first machine learning model to obtain corresponding encoded features; Performing feature interaction processing on the encoded features of the video text and the encoded features of the comment text through the feature interaction layer of the first machine learning model to obtain a first interaction feature, and Performing feature interaction processing on the encoded features of the video frame images and the encoded features of the comment text to obtain a second interaction feature; Performing splicing processing on the first interaction feature and the second interaction feature through the feature splicing layer of the first machine learning model to obtain a spliced feature; Performing relevance prediction on the spliced feature through the feature prediction layer of the first machine learning model to obtain the relevance between the comment and the video segment.
4. The method according to claim 1, wherein The aggregating the multiple comments according to the relevance between the comments and each video segment to obtain target comments corresponding to each video segment and satisfying the relevance condition includes: Performing the following processing for each of the video segments respectively: Select comments with a relevance reaching the relevance threshold from the multiple comments as candidate comments for the video clip; Based on the relevance between the candidate comments and each video clip, determine the relevance distribution score corresponding to the candidate comments. The relevance distribution score is used to indicate the distribution of the relevance between the corresponding candidate comments and each video clip; Select candidate comments with a relevance distribution score lower than the score threshold from the candidate comments as target comments whose relevance to the video clip meets the relevance condition.
5. The method according to claim 1, characterized in that, When the number of the target comments is at least two, when playing the target video clip among the at least two video clips, displaying the target comments corresponding to the target video clip includes: When playing the target video clip among the at least two video clips, obtain the target comments corresponding to the target video clip; Determine the sorting score corresponding to each of the target comments. The sorting score is used to indicate the expected degree of the viewing users of the corresponding target comment for the target comment; Based on the sorting score, sort the target comments in descending order to obtain the sorted target comments, and display the sorted target comments.
6. The method according to claim 5, characterized in that, The determining the sorting score corresponding to each of the target comments includes: For each of the target comments, perform the following processing respectively to determine the sorting score corresponding to each of the target comments: Obtain the first interest tags of the viewing users of the target video clip, the second interest tags of the publishing users of the target comments, the comment text of the target comments, and the interaction features associated with the target comments; Input the first interest tags, the second interest tags, the comment text, and the interaction features into a second machine learning model for sorting score prediction to obtain the sorting score corresponding to the target comment.
7. The method according to claim 6, characterized in that When the number of the interaction features is at least two, the inputting the first interest tags, the second interest tags, the comment text, and the interaction features into a second machine learning model for sorting score prediction to obtain the sorting score corresponding to the target comment includes: Through the feature encoding layer of the second machine learning model, encode the first interest tags, the second interest tags, and the comment text respectively to obtain the corresponding encoded features; Through the first feature connection layer of the second machine learning model, splice the at least two interaction features to obtain a first spliced feature; Through the feature interaction layer of the second machine learning model, perform feature interaction processing on the encoded features of the first interest tags and the encoded features of the second interest tags to obtain a first interaction feature, and perform feature interaction processing on the encoded features of the first interest tags and the encoded features of the comment text to obtain a second interaction feature; Through the second feature connection layer of the second machine learning model, splice the first interaction feature, the second interaction feature, and the first spliced feature to obtain a second spliced feature; Through the feature prediction layer of the second machine learning model, perform sorting score prediction on the second spliced feature to obtain the sorting score corresponding to the target comment.
8. The method according to claim 5, characterized in that, Said determining the sorting scores corresponding to each of the target comments includes: Performing the following processing on each of the target comments respectively to determine the sorting scores corresponding to each of the target comments: Determining the sub-sorting scores of the target comments under each of the sorting rules respectively according to a plurality of preset sorting rules; Obtaining the weight values corresponding to each of the sorting rules; Determining the sorting scores corresponding to the target comments based on the sub-sorting scores of the target comments under each of the sorting rules and the weight values corresponding to each of the sorting rules.
9. The method according to claim 1, wherein Said when playing a target video segment among the at least two video segments, displaying the target comments corresponding to the target video segment includes: During the process of playing the video, presenting a comment display area for displaying the comments of the video; When playing a target video segment among the at least two video segments, displaying the target comments corresponding to the target video segment through the comment display area.
10. The method according to claim 9, wherein The method further includes: When the number of the target comments is multiple and reaches a target number, presenting a scroll bar corresponding to the multiple target comments on one side of the comment display area; When receiving a trigger operation on the scroll bar, adjusting the display position of the target comments in the comment display area so as to adjust the target comments displayed through the comment display area.
11. The method according to claim 1, characterized in that, Said when playing a target video segment among the at least two video segments, displaying the target comments corresponding to the target video segment includes: During the process of playing the video, displaying the comments of the video in a bullet screen scrolling display manner; When playing a target video segment among the at least two video segments, displaying the target comments corresponding to the target video segment in a bullet screen scrolling display manner.
12. A processing device for comments, characterized in that, The apparatus includes: An obtaining module, configured to obtain at least two video segments included in a video and multiple comments of the video; A determining module, configured to respectively determine the relevance between the comments and each video segment, where the relevance is used to indicate the degree of content relevance between the comment and the video segment; An aggregating module, configured to aggregate the multiple comments according to the relevance between the comments and each video segment to obtain target comments corresponding to each video segment and satisfying a relevance condition; A displaying module, configured to display the target comments corresponding to the target video segment when playing a target video segment among the at least two video segments; wherein, the displaying module is specifically configured to: during the process of playing the video, presenting an information flow card for displaying the comments of the video; when playing a target video segment among the at least two video segments and receiving a trigger operation on a comment viewing function item, in response to the trigger operation on the comment viewing function item, using the information flow card to display the target comments corresponding to the target video segment.
13. An electronic device, characterized in that, The electronic device includes: A memory, configured to store executable instructions; A processor, configured to implement the method for processing comments according to any one of claims 1 to 11 when executing the executable instructions stored in the memory.
14. A computer-readable storage medium, characterized in that, Stored with executable instructions, when the executable instructions are executed, they are used to implement the method for processing comments as described in any one of claims 1 to 11.
15. A computer program product, the computer program product includes executable instructions, and the executable instructions are stored in a computer-readable storage medium; When the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the method for processing comments as described in any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Bullet screen information display method and device, computer equipment and storage medium
CN112533051A