Content recommendation method, device, computer equipment and storage medium

By using the pre-trained click-through rate prediction model combined with the feature information of the video clip, the recommendation time point of the video content is determined, and the problem of low content recommendation efficiency and accuracy in the prior art is solved, and more efficient and accurate content push is achieved.

CN111708941BActive Publication Date: 2025-05-09SHENZHEN YAYUE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010535054.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-12
Publication Date
2025-05-09
Estimated Expiration
2040-06-12

AI Technical Summary

Technical Problem

The prior art provides content recommendation by manually selecting recommended time points or identifying emotional mutation points in barrage data in video content, which is inefficient and has low accuracy, and cannot effectively identify the user's interest in recommended information.

Method used

Through the pre-trained click-through rate prediction model, based on the barrage characteristics, playback behavior characteristics and user characteristics of the video clip, the click-through rate prediction value of the video clip is determined, the values ​​that meet the recommendation conditions are filtered, and the recommended time point is determined based on these values, and then the recommended content is played when the video content is played to that time point.

Benefits of technology

It improves the efficiency and accuracy of content recommendation, and can more accurately analyze the user's interest in video content, thereby pushing recommended content at the right time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111708941B_ABST
    Figure CN111708941B_ABST
Patent Text Reader

Abstract

The present application relates to artificial intelligence, and in particular to a content recommendation method, device, computer equipment and storage medium. The method comprises: obtaining at least two video segments divided from video content; determining the click rate prediction value of each video segment based on the bullet screen features, playback behavior features and user features corresponding to each video segment through a pre-trained click rate prediction model; screening the click rate prediction values ​​that meet the recommendation conditions from the click rate prediction values ​​of each video segment; determining the recommendation time point based on the video segments corresponding to the screened click rate prediction values; and playing the recommended content when the video content is played to the recommended time point. The use of this method can effectively improve the recommendation efficiency and accuracy of the recommended content, thereby realizing accurate recommendation of the recommended content in the video content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a content recommendation method, apparatus, computer device and storage medium. Background Art

[0002] With the rapid development of Internet technology, video websites are becoming more and more popular. As a new way of watching and commenting on movies, bullet screens are gradually gaining popularity. More and more users are participating in bullet screen comments while watching videos. Bullet screens are topical and interesting, and a method of pushing information in video content based on bullet screens has emerged.

[0003] In traditional methods, the recommendation time point is usually selected manually, or the emotional mutation point of the barrage data is identified and used as the recommendation time point. However, the method of manually selecting the recommendation time point requires a lot of manpower costs and is inefficient. The emotional mutation point of the barrage data is often the turning point of the wonderful content. The user's focus is mainly on the video content, and it is impossible to effectively identify the user's interest in the recommended information, resulting in low recommendation efficiency and push accuracy of the information. Summary of the invention

[0004] Based on this, it is necessary to provide a content recommendation method, device, computer equipment and storage medium that can effectively improve the recommendation efficiency and recommendation accuracy of recommended content in response to the above technical problems.

[0005] A content recommendation method, the method comprising:

[0006] Acquire at least two video segments divided from the video content;

[0007] Determine the click rate prediction value of each video clip based on the bullet screen features, playback behavior features and user features corresponding to each video clip through a pre-trained click rate prediction model;

[0008] Selecting, from the click rate prediction values ​​of the video clips, click rate prediction values ​​that meet the recommendation condition;

[0009] Determine the recommended time point based on the video clips corresponding to the filtered click rate prediction values;

[0010] When the video content is played to the recommended time point, the recommended content is played.

[0011] A content recommendation device, comprising:

[0012] An information acquisition module, used to acquire at least two video segments divided from the video content;

[0013] A click rate prediction module, used to determine the click rate prediction value of each video segment based on the bullet screen features, playback behavior features and user features corresponding to each video segment through a pre-trained click rate prediction model;

[0014] A recommendation processing module, used to filter out click rate prediction values ​​that meet the recommendation conditions from the click rate prediction values ​​of the video clips; and determine the recommendation time point based on the video clips corresponding to the filtered click rate prediction values;

[0015] The content display module is used to play the recommended content when the video content is played to the recommended time point.

[0016] In one of the embodiments, the information acquisition module is also used to obtain the barrage information, playback behavior information and user information corresponding to each of the video clips; the barrage information includes barrage content and barrage numerical information; based on the barrage content, the barrage emotional feature value of each video clip is determined; according to the barrage emotional feature value and the barrage numerical information, the barrage attribute information of each video clip is generated.

[0017] In one of the embodiments, the information acquisition module is also used to extract text vectors corresponding to each of the barrage contents; perform sentiment analysis on the text vectors to obtain content sentiment feature values ​​of each barrage content; and determine the barrage sentiment feature values ​​corresponding to each video clip based on the content sentiment feature values ​​of each barrage content.

[0018] In one of the embodiments, the click-through rate prediction module is also used to extract barrage features and playback behavior features based on the barrage attribute information and the playback behavior information through a first extraction network included in the click-through rate prediction model; extract user features based on the user information through a second extraction network included in the click-through rate prediction model; and determine the click-through rate prediction value of each video clip based on the barrage features, the playback behavior features and the user features through a prediction layer included in the click-through rate prediction model.

[0019] In one of the embodiments, the click rate prediction module is also used to extract the barrage attribute information representation from the barrage attribute information and extract the playback behavior information representation from the playback behavior information through the first extraction network; and encode the barrage attribute information representation and the playback behavior information representation respectively to obtain barrage features and playback behavior features.

[0020] In one of the embodiments, the click rate prediction module is further used to extract user-related feature representations from the user information through the second extraction network; and perform feature encoding on the user-related feature representations to obtain user features of preset dimensions.

[0021] In one of the embodiments, the click rate prediction module is also used to fuse the barrage features, the playback behavior features and the user features through the prediction layer to obtain target multimodal features; and determine the click rate prediction value of each video clip based on the target multimodal features.

[0022] In one of the embodiments, the content recommendation device further includes a content generation module, which is used to obtain the bullet screen content of the video clip corresponding to the recommendation time point; and generate the recommended content corresponding to the recommendation time point based on the bullet screen content.

[0023] In one of the embodiments, the content generation module is also used to obtain description information of the object to be recommended; extract semantic features of the barrage content to obtain barrage semantic features; and generate recommended content corresponding to the recommendation time point based on the barrage semantic features and the description information.

[0024] In one of the embodiments, the recommended content is bullet screen recommended content; the content display module is further configured to play the bullet screen recommended content in the bullet screen area of ​​the video content when the video content is played to the recommended time point.

[0025] In one of the embodiments, the click-through rate prediction model is obtained through a training step, and the content recommendation device also includes a model training module for obtaining training samples and training labels; the training samples include sample barrage attribute information, sample playback behavior information and sample user information corresponding to each sample video clip in the sample video content; the training label is the historical click-through rate of the sample recommended content in the sample video content; the click-through rate prediction model is trained based on the training samples and the training labels.

[0026] In one of the embodiments, the model training module is also used to extract sample barrage features of the sample barrage attribute information and sample playback behavior features of the sample playback behavior information through a first extraction network included in the click-through rate prediction model; extract sample user features of the sample user information through a second extraction network included in the click-through rate prediction model; determine the sample click-through rate of each sample video clip based on the sample barrage features, the sample playback behavior features and the sample user features through the prediction layer included in the click-through rate prediction model; based on the difference between the sample click-through rate and the training label, adjust the parameters of the click-through rate prediction model and continue training until the training conditions are met and stop training.

[0027] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0028] Acquire at least two video segments divided from the video content;

[0029] Determine the click rate prediction value of each video clip based on the bullet screen features, playback behavior features and user features corresponding to each video clip through a pre-trained click rate prediction model;

[0030] Selecting, from the click rate prediction values ​​of the video clips, click rate prediction values ​​that meet the recommendation condition;

[0031] Determine the recommended time point based on the video clips corresponding to the filtered click rate prediction values;

[0032] When the video content is played to the recommended time point, the recommended content is played.

[0033] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0034] Acquire at least two video segments divided from the video content;

[0035] Determine the click rate prediction value of each video clip based on the bullet screen features, playback behavior features and user features corresponding to each video clip through a pre-trained click rate prediction model;

[0036] Selecting, from the click rate prediction values ​​of the video clips, click rate prediction values ​​that meet the recommendation condition;

[0037] Determine the recommended time point based on the video clips corresponding to the filtered click rate prediction values;

[0038] When the video content is played to the recommended time point, the recommended content is played.

[0039] The above-mentioned content recommendation method, device, computer equipment and storage medium, after obtaining at least two video clips divided from the video content, determine the click rate prediction value of each video clip through the pre-trained click rate prediction model based on the bullet screen features, playback behavior features and user features corresponding to each video clip; because the bullet screen features, playback behavior features and user features can reflect the user's viewing emotions, the browsing degree of the video clip and the main user group. By combining and analyzing the bullet screen features, playback behavior features and user features of each video clip, it is possible to accurately and effectively analyze the video clips in the video content that are more suitable for content push, thereby accurately analyzing the click rate prediction value of each video clip. Then, from the click rate prediction values ​​of each video clip, the click rate prediction values ​​that meet the recommendation conditions are screened; the recommended time point is determined based on the video clip corresponding to the screened click rate prediction value, and the recommended content is played when the video content is played to the recommended time point. In this way, content recommendation can be accurately performed at the recommended time point of the analyzed video content, thereby effectively improving the push efficiency and push accuracy of information. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A diagram of an application environment of a content recommendation method in an embodiment;

[0041] Figure 2 A schematic diagram of a flow chart of a content recommendation method in one embodiment;

[0042] Figure 3 A flowchart of performing sentiment analysis on a text vector in one embodiment;

[0043] Figure 4 This is an interface diagram of video content including bullet screen content in one embodiment;

[0044] Figure 5 This is an interface diagram of video content including bullet screen content in another embodiment;

[0045] Figure 6 A schematic diagram of the structure of a click rate prediction model in one embodiment;

[0046] Figure 7 A schematic diagram of a process of determining a click rate prediction value of each video clip by using a click rate prediction model in one embodiment;

[0047] Figure 8 A flowchart of a content recommendation method in another embodiment;

[0048] Fig. 9 A schematic diagram of an interface for playing bullet screen recommended content in video content in one embodiment;

[0049] Fig.10A schematic diagram of a flow chart of the steps of training a click rate prediction model in one embodiment;

[0050] Fig.11 It is a flowchart of a content recommendation method in a specific embodiment;

[0051] Fig.12 is a structural block diagram of a content recommendation device in one embodiment;

[0052] Fig.13 It is a structural block diagram of a content recommendation device in another embodiment;

[0053] Fig.14 is a structural block diagram of a content recommendation device in yet another embodiment;

[0054] Fig.15 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0056] The solution provided in the embodiment of the present application involves technologies such as artificial intelligence, machine learning (ML) and computer vision (CV) and image processing. Artificial intelligence is a theory, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results, so that the machine has the functions of perception, reasoning and decision-making. Machine learning involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc., and studies how computers simulate or realize human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve their performance. Computer vision and image processing technology is to use computer equipment to replace human eyes to identify, track and measure machine vision such as targets, and further do graphic processing, trying to establish an artificial intelligence system that can obtain information from images or multidimensional data. By processing various information corresponding to the video content based on machine learning and image processing technology, it is possible to effectively realize intelligent recommendation of recommended content.

[0057] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources in a wide area network or local area network to realize data calculation, storage, processing and sharing. It is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model application. It can form a resource pool, which is used on demand and is flexible and convenient. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, various types of data usually need to be transmitted to the background system for logical processing, and data of different levels will be processed separately. All kinds of industry data require strong system backing support. Cloud computing distributes computing tasks on a resource pool composed of a large number of computers, so that various application systems can obtain computing power, storage space and information services as needed. The content recommendation method provided in this application can be based on cloud technology for computing and processing, so that intelligent recommendation of recommended content can be realized efficiently.

[0058] The content recommendation method provided by the present application can be applied to a computer device. The computer device can be a terminal or a server. It is understandable that the content recommendation method provided by the present application can be applied to a terminal, a server, or a system including a terminal and a server, and is implemented through the interaction between the terminal and the server.

[0059] In one embodiment, the computer device may be a server. The content recommendation method provided in this application may be applied to Figure 1 In the application environment shown, the application environment includes a system of terminals and servers, and is implemented through the interaction between the terminals and the servers. Among them, the terminal 102 communicates with the server 104 through the network. After the server 104 obtains at least two video segments divided from the video content, it determines the click rate prediction value of each video segment based on the bullet screen features, playback behavior features and user features corresponding to each video segment through the pre-trained click rate prediction model. The server 104 then selects the click rate prediction value that meets the recommendation conditions from the click rate prediction values ​​of each video segment; determines the recommended time point based on the video segment corresponding to the selected click rate prediction value, and plays the recommended content when the video content is played to the recommended time point. The terminal 102 plays and displays the recommended content in the video content when the video content is played to the recommended time point. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented with an independent server or a server cluster composed of multiple servers.

[0060] In one embodiment, Figure 2As shown, a content recommendation method is provided, which is illustrated by applying it to a computer device, and the computer device can specifically be a terminal or a server. In this embodiment, the method includes the following steps:

[0061] S202: Obtain at least two video segments divided from video content.

[0062] Video refers to various technologies that capture, record, process, store, transmit and reproduce a series of static images in the form of electrical signals. The development of network technology has also enabled video documentary clips to exist on the Internet in the form of streaming media and can be received and played by computers. Video content is video data, which is a stream of images that changes over time and contains richer information and content that cannot be expressed by other media. Transmitting information in the form of video can express the content to be transmitted intuitively, vividly, realistically and efficiently.

[0063] The video content may be a video played on a video website or a video inserted in a web page. For example, it may be various film and television videos, live videos, program videos, and self-media videos. The video content includes at least two video clips. The video content to be processed may be obtained from a video website or from a video database.

[0064] Before processing the video content, the computer device can process the video content to be processed based on the video processing instruction. The video processing instruction can be automatically generated by the system. For example, when it is necessary to push the recommended object, the description information of the recommended object can be uploaded to the video website, and the backend server corresponding to the video website can automatically generate the video processing instruction. The video processing instruction can also be manually triggered by the user. For example, when the user browses the video content through the terminal, the video processing instruction can be triggered.

[0065] Specifically, after the computer device obtains the video content to be processed, it divides the video content according to a preset division method, and divides the video content into at least two video clips. The preset division method may be to divide the video content into equal parts according to the total duration of the video content, for example, to determine the number of video clips according to the total duration of the video content, and to divide the video content into equal parts. The video content may also be divided according to the preset segment duration, for example, to divide the video content according to a preset duration t, and obtain multiple video clips of the same duration t, for example, t may be 10 seconds, 15 seconds, etc.

[0066] S204, determining a click rate prediction value of each video segment through a pre-trained click rate prediction model based on the bullet screen features, playback behavior features, and user features corresponding to each video segment.

[0067] When the computer device obtains the video content, it also obtains various information associated with the video content, including barrage information, playback behavior information and user information corresponding to the video content.

[0068] Among them, barrage is a form of interaction. Users can enter their own comments in the comment box when watching videos, which is pop-up information. Barrage information, that is, video barrage, refers to the commentary subtitles played in pop-up form when watching videos on the Internet. These barrage information will be saved. When the video content is played again by the browsing user, the barrage information will be loaded at the same time as the player loads the video file, and will appear at the corresponding time point in the video content. Browsing users can also choose to turn off the barrage, or choose to only browse specific barrage information. Barrage information can include information such as barrage content, number of barrage likes, and number of barrages.

[0069] Playback behavior refers to the behavior of users performing various operations on video content when browsing video content, such as behavior information including play, stop, pause, fast forward, skip, and replay. Playback behavior information can be record information corresponding to various playback behaviors.

[0070] User information refers to user information corresponding to each user who browses the video content, for example, it can be user information corresponding to users who have browsed the current video content on the video website platform. User information can be user portrait information, such as including information such as gender and age.

[0071] In one embodiment, the video website platform may also pre-configure at least one object to be recommended, which may be a target object such as a product, application software, or user object, and the object to be recommended may also correspond to a corresponding application platform. The user information may also include user portrait information in the application platform of the object to be recommended. For example, the corresponding user information may be obtained in the application platform corresponding to the object to be recommended according to the user identifier. In this way, comprehensive user information associated with the object to be recommended may be obtained.

[0072] Among them, the click-through rate (CTR) refers to the click-through rate of network information (such as picture information, video information, advertising information, etc.) on the Internet, that is, the ratio of the actual number of clicks on the information content to the display volume (i.e., exposure). The click-through rate can usually reflect the quality effect of the recommended content, which can be used as an indicator to measure the quality effect of the recommended content. Taking advertising recommendation content as an example, CTR is an important indicator to measure the effect of Internet advertising. The click-through rate in this embodiment refers to the click-through rate of the recommended content in the video content.

[0073] The click-through rate prediction model is a model that has the ability to predict click-through rate after training. Specifically, it can be a neural network model based on logistic regression, a deep neural network model based on machine learning, or a neural network model that combines the two.

[0074] After the computer device obtains at least two video segments divided from the video content, it determines the click rate prediction value of each video segment through a pre-trained click rate prediction model based on the barrage features, playback behavior features and user features corresponding to each video segment.

[0075] Specifically, after the computer device obtains multiple video clips corresponding to the video content, it obtains the bullet screen information, playback behavior information and user information corresponding to each video clip. The computer device inputs the bullet screen information, playback behavior information and user information corresponding to each video clip into the pre-trained click rate prediction model, and extracts features of the bullet screen information, playback behavior information and user information through the click rate prediction model to obtain the bullet screen features, playback behavior features and user features corresponding to each video clip. The click rate prediction model determines the click rate prediction value of each video clip based on the extracted bullet screen features, playback behavior features and user features.

[0076] Since the barrage information is interesting and topical, it can reflect the emotions expressed by the viewing users when browsing the video content. The playback behavior information can reflect the exciting parts of the video content or the parts with high browsing rate. The user information can reflect the main user groups watching the video content. By combining and analyzing the barrage information, playback behavior information and user information of each video clip, it is possible to analyze which clips in the video content are more suitable for content push based on the user's viewing emotions, browsing rate and user group, thereby accurately analyzing the click-through rate prediction value of each video clip.

[0077] S206 , selecting click rate prediction values ​​that meet the recommendation condition from the click rate prediction values ​​of the video clips.

[0078] The click rate prediction is to predict the click situation of the recommended content, which is used to determine the probability of the recommended content being clicked by the user. The click rate prediction value is used to push the recommended content. The click rate prediction value that meets the recommendation conditions in a video content can be one or more. Among them, more than one means more than two.

[0079] After the computer device determines the click rate prediction value of each video clip in the video content through the click rate model, the click rate prediction value that meets the recommendation condition is screened from the click rate prediction values ​​of each video clip. Specifically, the computer device can sort the click rate prediction values ​​of each video clip in descending order, and screen out a preset number of click rate prediction values ​​according to the sorting result, that is, the preset number of click rate prediction values ​​with the highest numerical values ​​are determined as the click rate prediction values ​​that meet the recommendation condition. The click rate prediction values ​​can also be directly sorted from large to small, and the click rate prediction value with the largest numerical value is determined as the click rate prediction value that meets the recommendation condition.

[0080] The preset number may also be determined according to the total duration of the video content. It may also be determined according to the prediction threshold corresponding to the click rate prediction value. For example, the preset number may be a preset value range. When there are multiple click rate prediction values ​​that reach the prediction threshold, the click rate prediction value that meets the condition may be selected according to the preset number.

[0081] S208: Determine a recommended time point based on the screened video clips corresponding to the predicted click rate values.

[0082] The video content includes a corresponding video timeline, and the timeline refers to a recording system connected in time sequence. The video timeline means connecting multiple consecutive frames of images in time tracks. Each video clip in the video content is divided according to the video timeline of the video content. Each video clip has a corresponding time period on the video timeline of the video content. The recommended time point refers to a time point on the video timeline in the video content, and is used to insert the recommended content to be recommended at the recommended time point in the video content.

[0083] After the computer device selects the click rate prediction values ​​that meet the recommendation conditions from the click rate prediction values ​​of each video clip, it determines the recommended time point in the video content according to the video clip corresponding to the selected click rate prediction value. The recommended time point in the video content can be one or more. When there are multiple video clips selected, the recommended time points correspond to the corresponding video clips, that is, there are also multiple.

[0084] Specifically, the computer device may also determine the starting point of the screened video segment, that is, the time point of the starting point of the video segment in the video content, as the recommended time point. For example, the recommended time point in the video content may also be determined based on the middle point or end point of the video segment.

[0085] S210, playing the recommended content when the video content is played to the recommended time point.

[0086] The recommended content may be pre-configured content corresponding to the recommended object, and the recommended content may be pre-configured information. The recommended object refers to the thing that is the target of the recommendation, for example, the recommended object may include products, application software, users, promotional information, etc. The recommended content may include information in various forms such as plain text, plain pictures, icons, or a combination of pictures and text. The recommended content may also include attribute information such as playback duration and playback position. For example, the recommended content may include user push information, resource promotion information, and various advertising information, etc.

[0087] After the computer device determines the recommended time point in the video content, it obtains the recommended content corresponding to the object to be recommended, and plays the recommended content when the video content is played to the recommended time point, thereby realizing content recommendation in the video content. Among them, the recommended content can generate corresponding information according to a preset format, such as text, graphics, icons, and a combination of graphics and text. The recommended content also includes attribute information such as a preset display position, display form, and display duration. For example, the display form includes corner marks, screen pressure bars, etc. The recommended content can be inserted into the video content in an embedded manner for playback without affecting the playback of the video content itself, thereby effectively realizing content recommendation in the video content.

[0088] After the user loads the video content with the added recommended content through the corresponding user terminal, when the video content is played to the recommended time point in the video display interface of the user terminal, the corresponding recommended content is played. The user can also click on the recommended content in the video display interface to jump to the relevant page of the recommended object to recommend the recommended content.

[0089] In the above content recommendation method, after the computer device obtains at least two video clips divided from the video content, the click rate prediction value of each video clip is determined through the pre-trained click rate prediction model based on the bullet screen features, playback behavior features and user features corresponding to each video clip; because the bullet screen features, playback behavior features and user features can reflect the user's viewing emotions, the browsing degree of the video clip and the main user group. By combining and analyzing the bullet screen features, playback behavior features and user features of each video clip, it is possible to accurately and effectively analyze the video clips in the video content that are more suitable for content push, thereby accurately analyzing the click rate prediction value of each video clip. The computer device then selects the click rate prediction value that meets the recommendation conditions from the click rate prediction values ​​of each video clip; determines the recommendation time point based on the video clip corresponding to the selected click rate prediction value, and plays the recommended content when the video content is played to the recommended time point. In this way, content recommendation can be accurately performed at the recommended time point of the analyzed video content, thereby effectively improving the push efficiency and push accuracy of information.

[0090] In one embodiment, after obtaining at least two video segments divided from the video content, the above-mentioned content recommendation method also includes: obtaining the barrage information, playback behavior information and user information corresponding to each video segment; the barrage information includes barrage content and barrage numerical information; based on the barrage content, determining the barrage emotional feature value of each video segment; generating the barrage attribute information of each video segment according to the barrage emotional feature value and the barrage numerical information.

[0091] The barrage information includes barrage content and barrage value information. The barrage content may include barrage text, pictures, icons, or a combination of pictures and texts, and the barrage value information includes barrage likes and the number of barrages. The number of barrages may be the ratio of the number of barrages in each video clip to the total number of barrages in the entire video content. The number of barrage likes may be the ratio of the number of likes in each video clip to the total number of all likes in the entire video content.

[0092] The emotional feature value of the barrage refers to the emotional expression of the user's barrage content, which can be specifically expressed as a barrage emotional score. The barrage emotional feature value can reflect whether the user's emotional expression of the barrage content is positive or negative. For example, the barrage content of a video clip in the video content can reflect the user's interest in this video clip. Generally speaking, for video clips that users are not interested in, the emotional expression of the barrage content is more negative; and for video clips that users are more interested in, the emotional expression of the barrage content is more positive.

[0093] After obtaining at least two video segments divided from the video content, the computer device obtains the bullet screen information, playback behavior information and user information corresponding to each video segment, that is, obtains all the bullet screen information, playback behavior information included in each video segment, and user information who has browsed the video segment.

[0094] After obtaining the bullet screen information of each video clip, the computer device also performs emotional feature analysis on the bullet screen content in the bullet screen information, and specifically extracts text emotional features from the bullet screen text in the bullet screen content, thereby obtaining the bullet screen emotional feature value of each video clip. The computer device then uses the bullet screen emotional feature value and the bullet screen numerical information to generate bullet screen attribute information of each video clip.

[0095] In this embodiment, after obtaining the barrage information, playback behavior information and user information corresponding to each video clip, the barrage emotional feature value of each video clip is determined according to the barrage content, and based on the barrage emotional feature value and the barrage numerical information, the barrage attribute information of each video clip can be effectively obtained, so that each video clip can be more accurately analyzed and processed.

[0096] In one embodiment, based on the barrage content, the barrage emotional feature value of each video clip is determined, including: extracting the text vector corresponding to each barrage content; performing sentiment analysis on the text vector to obtain the content emotional feature value of each barrage content; and determining the barrage emotional feature value corresponding to each video clip based on the content emotional feature value of each barrage content.

[0097] Among them, the sentiment analysis of the barrage content can be processed by a pre-trained sentiment analysis model. The sentiment analysis model can be a text sentiment feature extraction based on an LSTM (Long Short-Term Memory) model. In addition, a text sentiment feature extraction based on a DNN (Deep Neural Networks) model or a CNN (Convolutional Neural Networks) model can also be used, which is not limited here.

[0098] Specifically, the computer device inputs the bullet content corresponding to each video clip into the sentiment analysis model. The bullet text in the bullet content is segmented to obtain the word vector corresponding to the bullet text, and the text vector corresponding to each bullet content is extracted according to the word vector. The computer device then performs sentiment analysis on the text vector through the sentiment analysis model to obtain the content sentiment feature value of each bullet content. Among them, the computer device can perform sentiment analysis on each bullet content in each video clip one by one through the sentiment analysis model to obtain the content sentiment feature value of each bullet content.

[0099] After the sentiment analysis model is used to analyze the content sentiment feature values ​​of each barrage content, the comprehensive barrage sentiment feature value corresponding to each video clip is determined based on the content sentiment feature values ​​of all barrage contents included in the video clip.

[0100] In one embodiment, if Figure 3 As shown, it is a flowchart of performing sentiment analysis on text vectors in one embodiment to obtain the content sentiment feature value of each barrage content. The computer device first performs word segmentation on the barrage text in the barrage content and generates a word vector corresponding to the barrage content. Then, the sentiment feature of the word vector of each barrage content is extracted through the pre-trained sentiment analysis model to obtain the sentiment analysis result corresponding to each barrage content. In this way, the content sentiment feature value of each barrage content can be accurately and effectively obtained.

[0101] For example, the value range of the content emotion feature value can be -1.0-1.0. The emotion feature of each barrage content is extracted through the emotion analysis model, and a normalized value of [0-1] is output. The computer device further normalizes the content emotion feature value of each barrage content in each video clip. For example, the content emotion feature values ​​of all barrage contents in the video clip are normalized to [0~1.0] and then added and averaged to obtain the comprehensive barrage emotion feature value corresponding to each video clip. For example, the specific calculation formula can be as follows:

[0102]

[0103] Where S represents the emotional score of the bullet screen in the video clip of length t, that is, the comprehensive emotional feature value of the bullet screen corresponding to the video clip, v i represents the sentiment score of the i-th comment before normalization, and n represents the total number of comments under the video clip.

[0104] For example, Figure 4 As shown in FIG. 1 , it is an interface diagram of video content including bullet screen content in a specific embodiment. Figure 4 From the barrage content sent by users shown in , it can be seen that the users’ emotional evaluation of watching the movie is low, so this part of the barrage can be determined as the barrage content with a low emotional score.

[0105] like Figure 5 As shown in FIG. 1 , it is an interface diagram of video content including bullet screen content in another specific embodiment. Figure 5 From the barrage content sent by users shown in , it can be seen that users have high emotional evaluations of watching movies, so this part of the barrage can be determined as barrage content with a higher emotional score.

[0106] In this embodiment, the sentiment feature extraction of the barrage content of each video clip is performed through the sentiment analysis model, which can accurately and effectively identify the content sentiment feature value of each barrage content, and then based on the content sentiment feature value of all barrage content in each video clip, the comprehensive barrage sentiment feature value of each video clip can be accurately obtained.

[0107] In one embodiment, a pre-trained click-through rate prediction model is used to determine the click-through rate prediction value of each video clip based on the barrage features, playback behavior features, and user features corresponding to each video clip, including: extracting barrage features and playback behavior features based on barrage attribute information and playback behavior information through a first extraction network included in the click-through rate prediction model; extracting user features based on user information through a second extraction network included in the click-through rate prediction model; and determining the click-through rate prediction value of each video clip based on the barrage features, playback behavior features, and user features through a prediction layer included in the click-through rate prediction model.

[0108] Among them, the click-through rate prediction model is a model that has been pre-trained and has the ability to predict click-through rates, and can specifically be a neural network model based on machine learning. The click-through rate prediction model includes a first extraction network, a second extraction network and a prediction layer, that is, the click-through rate prediction model is a combination model including a first extraction network and a second extraction network. Among them, the first extraction network can be a network structure based on a regression model, which is used to extract barrage features and playback behavior features. For example, the first extraction network can be a meta-model in a logistic regression model, that is, a partial network structure included in the logistic regression model for extracting specific feature vectors. Among them, the meta-model describes the elements in the model, the relationships between elements, and the representations, and the model includes the meta-model. Taking the neural network model as an example, the meta-model can be regarded as a part of the neural network structure of the model, which is used to extract specific feature representations.

[0109] Similarly, the second extraction network may be a network structure based on a deep neural network model, a network structure for extracting user feature vectors, for example, a meta-model in a deep neural network model, that is, a partial network structure included in the deep neural network model for extracting user feature vectors. Figure 6 FIG. 1 is a schematic diagram of the structure of a click rate prediction model in an embodiment.

[0110] The computer device obtains at least two video segments divided from the video content, and obtains the barrage attribute information, playback behavior information and user information corresponding to each video segment, and then inputs the barrage attribute information, playback behavior information and user information corresponding to each video segment into a pre-trained click rate prediction model.

[0111] Specifically, the barrage attribute information and playback behavior information of each video clip are input into the first extraction network of the click rate prediction model. Through the first extraction network, the barrage features and playback behavior features are extracted based on the barrage attribute information and playback behavior information, thereby obtaining the barrage features and playback behavior features corresponding to each video clip.

[0112] The user information corresponding to each video clip is input into the second extraction network included in the click rate prediction model, and the user features are extracted based on the user information by the second extraction network, thereby obtaining the user features corresponding to each video clip.

[0113] After extracting the barrage features, playback behavior features and user features corresponding to each video clip, the prediction layer included in the click-through rate prediction model is used to determine the click-through rate prediction value of each video clip based on the barrage features, playback behavior features and user features, thereby accurately and effectively obtaining the click-through rate prediction value of each video clip.

[0114] In this embodiment, the pre-trained click-through rate prediction model can accurately extract the barrage features and playback behavior features corresponding to the barrage attribute information and playback behavior information in each video clip, as well as the user characteristics corresponding to the user information, thereby effectively capturing the user's viewing emotion characteristics, browsing characteristics and user group characteristics corresponding to each video clip, and determining the click-through rate prediction value of each video clip based on the barrage features, playback behavior characteristics and user characteristics, thereby accurately analyzing the click-through rate prediction value of each video clip for the recommended content.

[0115] In one embodiment, the first extraction network included in the click-through rate prediction model extracts barrage features and playback behavior features based on barrage attribute information and playback behavior information, including: extracting barrage attribute information representation from the barrage attribute information and extracting playback behavior information representation from the playback behavior information through the first extraction network; encoding the barrage attribute information representation and the playback behavior information representation respectively to obtain barrage features and playback behavior features.

[0116] The first extraction network may be a pre-trained linear model included in the click rate prediction model. The generalized linear model is a mathematical model that quantitatively describes statistical relationships and is used to analyze the relationship between dependent variables (targets) and independent variables (predictors), such as the significant relationship between independent variables and dependent variables and the influence of multiple independent variables on a dependent variable. The first extraction network is used to extract various feature representations corresponding to the bullet comment attribute information and the playback behavior information.

[0117] For example, the first extraction network can use a meta-model based on a logistic regression model to extract features from the bullet comment attribute information and the playback behavior information to obtain corresponding bullet comment features and playback behavior features. In addition, the first extraction network can also use a meta-model such as a linear regression model or a stepwise regression model to extract features from the bullet comment attribute information and the playback behavior information, which is not limited here.

[0118] After the computer device inputs the bullet screen attribute information and playback behavior information of each video clip and the user information into the click-through rate prediction model, the bullet screen attribute information and playback behavior information are input into the first extraction network. The first extraction network first performs feature extraction on the bullet screen attribute information and the playback behavior information, extracts the bullet screen attribute information representation from the bullet screen attribute information, and extracts the playback behavior information representation from the playback behavior information. For example, the information representation obtained can be each feature vector corresponding to the bullet screen attribute information and the playback behavior information, such as a bullet screen emotion vector, a bullet screen like vector, a bullet screen number vector, a video replay vector, a video skip vector and other feature vectors, each of which also includes a corresponding vector value. The bullet screen attribute information representation and the playback behavior information representation are encoded and processed respectively to obtain bullet screen features and playback behavior features. Specifically, the first extraction network performs linear processing according to each feature vector and vector value corresponding to the bullet screen attribute information and the playback behavior information to obtain the corresponding bullet screen features and playback behavior features. The bullet screen features and playback behavior features can specifically include feature values ​​in a preset numerical range.

[0119] Taking the first extraction network as a logistic regression model as an example, after extracting each feature vector corresponding to the bullet comment attribute information and the playback behavior information through the logistic regression model, the vector value corresponding to each feature vector is normalized to a feature value within a preset numerical range, for example, each vector value is normalized to a feature value in the size interval of [0-1]. Then, the vector value corresponding to each feature vector is linearly output through the logistic regression model to obtain the bullet comment features and playback behavior features corresponding to each video clip. Among them, the logistic regression formula can be as follows:

[0120] y=W T X+b

[0121] Among them, y is the feature value used for prediction, or it can be the probability of each vector used for prediction, X represents the feature vector, W represents the parameters of the model, that is, the weight corresponding to each feature vector finally trained, and b is the bias term, that is, the constant term.

[0122] In this embodiment, the first extraction network included in the click-through rate prediction model is used to extract barrage features and playback behavior features based on the barrage attribute information and playback behavior information, thereby effectively analyzing the linear relationship between the barrage attribute information and playback behavior information and the recommended content, thereby accurately extracting the barrage features and playback behavior features corresponding to each video clip, which can be used to accurately predict the click-through rate prediction value of each video clip.

[0123] In one embodiment, user features are extracted based on user information through a second extraction network included in a click rate prediction model, including: extracting user-related feature representations from user information through the second extraction network; and feature encoding the user-related feature representations to obtain user features of preset dimensions.

[0124] Among them, the second extraction network is a pre-trained deep neural network model (Deep Models), and the second extraction network includes at least two layers of network structure, which is used to extract various feature representations corresponding to various association vectors included in the user information. The second extraction network can be based on a DNN (deep neural network) model for user feature extraction. In addition, user feature extraction can also be based on an LSTM (long short-term memory network) model or a CNN (convolutional neural network) model, which is not limited here.

[0125] After the computer device inputs the bullet comment attribute information and playback behavior information of each video clip and the user information into the click rate prediction model, the user information is input into the second extraction network. The second extraction network first extracts features from the user information and extracts user-related feature representations, that is, user-related features, such as gender, age, interests, hobbies, etc. from the user information. The second extraction network further performs feature encoding on the user-related feature representation through the encoding network layer therein, thereby obtaining user features of preset dimensions.

[0126] Taking the second extraction network based on the DNN model as an example, the DNN model includes an input layer, an embedding layer (Embedding) and several hidden layers. After the user information is input through the input layer of the second extraction network, the high-dimensional vector in the user information is converted into a low-dimensional embedding representation, that is, a user-related feature representation, through the embedding layer. For example, the high-dimensional vector representing the user id (if there are 1,000 users, the one hot vector corresponding to the user id is 0,0,…1,…0) is converted into a low-dimensional and dense user embedding (for example, 0.33458763, 0.69234245, 0.1034593…), and the user embedding vector represents the relevant features of the user to a certain extent, such as gender, age, interests, hobbies, etc. The user-related feature representation is feature encoded through the hidden layer of the second extraction network to obtain user features of preset dimensions, thereby accurately and effectively extracting the user features of each user in each video clip.

[0127] In one embodiment, the video website platform may pre-store the user-related feature representations of each user obtained through training. When processing the video content, the user-related feature representations of the corresponding user may be directly obtained from the video website platform for processing. In this way, the user-related feature representations of each user may be obtained quickly and effectively, thereby effectively improving the processing efficiency and speed of data.

[0128] In this embodiment, the user information is subjected to feature extraction through the second extraction network in the click rate prediction model, thereby being able to accurately and effectively obtain the user features corresponding to each video clip.

[0129] In one embodiment, Figure 7 As shown, the steps of determining the click rate prediction value of each video clip through the click rate prediction model specifically include the following contents:

[0130] S702, extracting bullet comment features and playback behavior features based on bullet comment attribute information and playback behavior information through a first extraction network included in the click rate prediction model.

[0131] S704: extracting user features based on user information through a second extraction network included in the click rate prediction model.

[0132] S706, through the prediction layer included in the click rate prediction model, the bullet screen features, the playback behavior features and the user features are fused to obtain the target multimodal features.

[0133] S708: Determine a click rate prediction value of each video segment based on the target multimodal feature.

[0134] Among them, the first extraction network can be a linear model, the second extraction network can be a deep neural network model, and the prediction layer included in the click-through rate prediction model includes preset prediction functions and weights for predicting the click-through rate of recommended content in each video clip.

[0135] After the computer device obtains at least two video segments divided from the video content, as well as the barrage information, playback behavior information and user information corresponding to each video segment, the barrage information, playback behavior information and user information are input into a pre-trained click-through rate prediction model, and the barrage features and playback behavior features are extracted based on the barrage attribute information and playback behavior information through the first extraction network of the click-through rate prediction model, and user features are extracted based on the user information through the second extraction network included in the click-through rate prediction model.

[0136] After the click-through rate prediction model extracts the bullet screen features, playback behavior features, and user features respectively, the bullet screen features, playback behavior features, and user features are input into the prediction layer included in the click-through rate prediction model. Among them, the prediction layer can also include a feature connection layer for fusing various features. Specifically, through the feature connection layer of the prediction layer, the bullet screen features, playback behavior features, and user features are feature fused to obtain the target multimodal features. The prediction layer then performs regression prediction on the click-through rate of the recommended content in the video clip based on the obtained target multimodal features, thereby obtaining the click-through rate prediction value of each video clip.

[0137] For example, the prediction layer can use logistic loss as the loss function, which can be expressed as follows:

[0138]

[0139] Among them, W represents the weight of the model, T represents the transpose of the weight, b represents the bias, x represents the feature, a represents the sigmoid function, φ(x) represents the cross product feature, and a lf represents the activation value of the last layer of the neural network, and p represents the predicted value of the click rate.

[0140] In this embodiment, the click-through rate prediction model constructed by a combined model including a linear model and a deep neural network model can effectively extract the user's behavior characteristics and barrage characteristics, as well as user characteristics in the video content, and can accurately and effectively capture the relationship between user behavior characteristics and barrage characteristics and user characteristics and the click-through rate of recommended content in each video clip, so as to accurately predict the click-through rate prediction value of each video clip in the video content, and then accurately and effectively analyze the video clips in the video content that are more suitable for content push.

[0141] In one embodiment, after determining the recommended time point based on the screened video clip corresponding to the click rate prediction value, the above-mentioned content recommendation method also includes: obtaining the barrage content of the video clip corresponding to the recommended time point; based on the barrage content, generating recommended content corresponding to the recommended time point.

[0142] The recommended content is the content corresponding to the preset object to be recommended. The object to be recommended includes descriptive information such as the recommended object identifier, the recommended object name, and the recommended object attributes.

[0143] The computer device uses a pre-trained click rate prediction model to determine the click rate prediction value corresponding to each video clip according to the barrage information, playback behavior information and user information corresponding to each video clip in the video content, and screens the click rate prediction value that meets the recommendation conditions from the click rate prediction values ​​of each video clip.

[0144] After determining the recommended time point based on the screened video clip corresponding to the predicted click rate value, the computer device further obtains the barrage content of the video clip corresponding to the recommended time point, and generates recommended content corresponding to the recommended time point based on the barrage content.

[0145] Specifically, the computer device extracts semantic features of the bullet-screen content corresponding to the video clip to obtain the semantic features of the bullet-screen, and then generates recommended content related to the bullet-screen content according to the semantic features of the bullet-screen. When generating the recommended content, the recommended object identifier or the recommended object name according to the recommended object can also be combined to generate the recommended content including the recommended object identifier or the recommended object name.

[0146] The computer device can also specifically extract semantic features of the barrage content through a pre-trained content generation model to obtain barrage semantic features, and generate recommended content related to the barrage content based on the barrage semantic features.

[0147] In this embodiment, by generating recommended content corresponding to the recommended time point based on the barrage content of the video clip corresponding to the recommended time point, it is possible to effectively generate recommended content related to the barrage content, thereby making the generated recommended content more consistent with the user's viewing mood, thereby effectively improving the click-through rate of the recommended content in the video content.

[0148] In one embodiment, based on the barrage content, determining the recommended content corresponding to the recommended time point includes: obtaining description information of the object to be recommended; extracting semantic features of the barrage content to obtain barrage semantic features; and generating the recommended content corresponding to the recommended time point based on the barrage semantic features and the description information.

[0149] When generating recommended content, the computer device also obtains description information of the object to be recommended. The computer device extracts semantic features from the barrage content, and after obtaining the barrage semantic features, it also generates recommended content corresponding to the recommended time point based on the combination of the barrage semantic features and the description information. Specifically, the computer device extracts semantic features from all barrage content in the video clip corresponding to the recommended time point to obtain the barrage semantic features corresponding to the video clip. The computer device also extracts semantic features from the description information of the object to be recommended to obtain the semantic features of the recommended object. Then, the barrage semantic features are combined with the semantic features of the recommended object to generate the corresponding recommended content.

[0150] The computer device can also generate a pre-trained content generation model to generate recommended content corresponding to the recommended time point according to the semantic features of the barrage and the semantic features of the recommended object, thereby accurately and efficiently generating recommended content that is compatible with the barrage content and the object to be recommended.

[0151] In one of the embodiments, there may be multiple objects to be recommended, and the objects to be recommended include description information, and the description information also includes category attribute information. When there are multiple objects to be recommended, the computer device extracts semantic features of the barrage content, and after obtaining the barrage semantic features, it can further filter out the most matching recommended objects from the objects to be recommended based on the barrage semantic features of the barrage content for recommendation. Specifically, the computer device can determine the degree of match between the barrage content and each object to be recommended based on the barrage semantic features and the category attribute information or description information of each object to be recommended, and filter out the recommended object with the highest match as the object to be recommended.

[0152] In this embodiment, the recommended content corresponding to the recommended time point is generated by combining the barrage content of the video clip corresponding to the recommended time point with the description information of the object to be recommended. This can accurately and efficiently generate recommended content that is suitable for both the barrage content and the object to be recommended. Therefore, the generated recommended content can be more in line with the user's viewing mood and user characteristics, thereby effectively improving the click-through rate of the recommended content in the video content.

[0153] In one embodiment, Figure 8 As shown, the recommended content is the bullet screen recommended content; another content recommendation method is provided, the method comprising the following steps:

[0154] S802: Obtain at least two video segments divided from video content.

[0155] S804, determining the click rate prediction value of each video segment through the pre-trained click rate prediction model based on the bullet comment features, playback behavior features and user features corresponding to each video segment.

[0156] S806: Filter the predicted click rate values ​​that meet the recommendation condition from the predicted click rate values ​​of the video clips.

[0157] S808: Determine a recommended time point based on the screened video clips corresponding to the predicted click rate values.

[0158] S810, when the video content is played to the recommended time point, the bullet screen recommendation content is played in the bullet screen area of ​​the video content.

[0159] The bullet screen recommendation content refers to the recommended content in the form of bullet screen, that is, the recommended content displayed in the bullet screen area of ​​the video content when the video content is played. The bullet screen recommendation content is at least one of text, picture, icon or a combination of picture and text.

[0160] After the computer device obtains at least two video clips divided from the video content, and the bullet screen information, playback behavior information and user information corresponding to each video clip, the computer device extracts the bullet screen features, playback behavior features and user features corresponding to each video clip based on the bullet screen information, playback behavior information and user information through the pre-trained click rate prediction model, and determines the click rate prediction value bullet screen information, playback behavior information and user information of each video clip. The computer device then selects the click rate prediction value that meets the recommendation condition from the click rate prediction value of each video clip. The recommended time point is determined based on the video clip corresponding to the selected click rate prediction value, and the bullet screen recommendation content corresponding to the recommended time point is generated, and then when the video content is played to the recommended time point, the recommended content is played in the bullet screen area of ​​the video content.

[0161] Since the generated bullet screen recommendation content is pushed in the bullet screen area of ​​the video content and played together with other bullet screen content, it can effectively reduce the user's aversion to the recommended content when pushing the recommended content. Therefore, it is possible to accurately recommend content at the recommended time point of the analyzed video content, thereby effectively improving the efficiency and accuracy of information push.

[0162] like Fig. 9 FIG. 1 is a schematic diagram of an interface for playing bullet screen recommended content in a bullet screen area of ​​a video content in an embodiment. Fig. 9 , Fig. 9 The video playback interface is shown here. The bottom of the playback interface is the playback function box, and the upper area of ​​the playback interface is the bullet screen area. Fig. 9 In 902, the upper area of ​​the playback interface is the subtitle area. For example, the barrage content in the barrage area includes: "I have never seen a firefly", "What kind of bug is this", "There were so many in the countryside of my hometown when I was a child", "Shiny", etc. It can be seen from the video content and subtitle information that the video content is popular science content. After determining the click-through rate prediction value of each video clip by combining and analyzing the barrage features, playback behavior features and user features corresponding to each video clip. If the current video screen is a click-through rate prediction value that meets the recommendation conditions, the barrage recommendation content is played when the video content is played to the corresponding recommended time point. For example, the barrage recommendation content can be "XXX, find it more interesting". And when the video content is played to the corresponding recommended time point, the barrage recommendation content is played in the barrage area. Refer to Fig. 9904 in the figure is the recommended content of the barrage. After analyzing the predicted click rate of each video clip by combining the barrage features, playback behavior features and user features, and determining the recommended time point in the video content, the barrage recommended content associated with the barrage content is generated and played and pushed in the barrage area along with other barrage content. This can effectively reduce the contrast effect and user aversion of information push in the video content, and can accurately recommend content at the recommended time point of the analyzed video content, thereby effectively improving the efficiency and accuracy of information push.

[0163] In one embodiment, the click-through rate prediction model is obtained through a training step, and the training step includes: obtaining training samples and training labels; the training samples include sample barrage attribute information, sample playback behavior information and sample user information corresponding to each sample video clip in the sample video content; the training label is the historical click-through rate of the sample recommended content in the sample video content; and the click-through rate prediction model is trained based on the training samples and training labels.

[0164] The click rate prediction model is trained using training sample data. Before processing the video content using the click rate prediction model, the required click rate prediction model needs to be trained in advance.

[0165] The training samples may be sample video content within a historical time period, and the sample video content includes sample bullet-screen attribute information, sample playback behavior information, and sample user information corresponding to each sample video clip. That is, the bullet-screen attribute information, playback behavior information, and user information of the sample video content within the past period of time. The sample video content includes the historical sample recommended content released within the historical time period, and the sample video content also includes the real historical click-through rate of the sample recommended content within the historical time period.

[0166] In the process of training the click-through rate prediction model, the sample bullet comment attribute information, sample playback behavior information and sample user information corresponding to the sample video clip are used as training samples for training, and the historical click-through rate of the sample recommended content in the sample video content is used as the training label. The training label is used to adjust the parameters of each training result to further train and optimize the click-through rate prediction model.

[0167] The training samples can be obtained from a preset sample library or from various platforms, such as video content published or shared on video playback networks, video sharing networks, various web pages, etc. User information of users who browse sample video content on the corresponding platform can also be obtained.

[0168] Specifically, after the computer device obtains the training sample, the sample bullet comment attribute information, sample playback behavior information and sample user information in the training sample are input into the preset click rate prediction model for training, and the click rate prediction model is adjusted and optimized using the training label to train a click rate prediction model that meets the conditions. By using the training sample and the training label to train the click rate prediction model, a click rate prediction model with prediction ability can be effectively obtained.

[0169] In one embodiment, Fig.10 FIG. 1 is a step of training a click rate prediction model in an embodiment, which specifically includes the following contents:

[0170] S1002, obtaining training samples and training labels; the training samples include sample barrage attribute information, sample playback behavior information and sample user information corresponding to each sample video clip in the sample video content; the training label is the historical click rate of the sample recommended content in the sample video content.

[0171] S1004, extracting sample barrage features of sample barrage attribute information and sample playback behavior features of sample playback behavior information through a first extraction network included in the click rate prediction model.

[0172] S1006: extracting sample user features of the sample user information through a second extraction network included in the click rate prediction model.

[0173] S1008, determining the sample click rate of each sample video clip based on the sample bullet comment features, the sample playback behavior features and the sample user features through the prediction layer included in the click rate prediction model.

[0174] S1010, based on the difference between the sample click rate and the training label, adjust the parameters of the click rate prediction model and continue training until the training condition is met and the training is stopped.

[0175] The click rate prediction model includes a first extraction network and a second extraction network. The first extraction network can be a linear model, and the second extraction network can be a deep neural network model. The first extraction network and the second extraction network can be used as encoder layers in the click rate prediction model.

[0176] After the computer device inputs the sample bullet attribute information, sample playback behavior information and sample user information in the training sample into the preset click rate prediction model, the sample bullet attribute information and the sample playback behavior information are feature extracted through the first extraction network included in the click rate prediction model, and the sample bullet attribute information and the sample playback behavior information are respectively extracted. At the same time, the sample user features corresponding to the sample user information are extracted through the second extraction network included in the click rate prediction model. Further, through the prediction layer of the click rate prediction model, the sample bullet feature, the sample playback behavior feature and the sample user feature, the click rate of the sample recommended content in the sample video content is regressively predicted to obtain the sample click rate of each sample video clip. Then, based on the difference between the sample click rate and the sample label, the parameters of the click rate prediction model are adjusted and training is continued until the training conditions are met and the training is stopped.

[0177] The difference between the sample click rate and the efficiency label can be measured by a loss function, such as the mean absolute value loss function (MAE), the smoothed mean absolute error (Huber loss), the cross entropy loss function, etc. The training condition is the condition for ending the model training. The training stop condition can be reaching a preset number of iterations, or the prediction performance index of the click rate prediction model after adjusting the parameters reaches the preset index.

[0178] In one embodiment, the parameters of the first extraction network and the second extraction network can be transferred and learned during the process of training the click-through rate prediction model to fine-tune the parameters, for example, by using a fine-tune method.

[0179] The computer device can quickly and accurately extract the sample bullet screen features and sample playback behavior features of the sample video content through the first extraction network; and can quickly and accurately extract the sample user features of the sample video content through the second extraction network. Click-through rate prediction training is performed based on the sample bullet screen features, sample playback behavior features and sample user features to obtain the sample click-through rate. The computer device can then gradually adjust the parameters in the click-through rate prediction model according to the difference between the obtained sample click-through rate and the training label. Therefore, during the parameter adjustment process, the click-through rate prediction model can simultaneously combine the sample bullet screen features, sample playback behavior features and sample user features to capture the implicit relationship between the click-through rate of the sample video content and the recommended content. When predicting the click-through rate of the recommended content in the video content based on the click-through rate prediction model, multiple guidance of the sample bullet screen features, sample playback behavior features and sample user features is obtained, thereby being able to train a click-through rate prediction model with high prediction accuracy, thereby improving the accuracy of the click-through rate prediction of the recommended content in the video content.

[0180] In a specific embodiment, Fig.11 As shown, a specific content recommendation method is provided, comprising the following steps:

[0181] S1102: Obtain at least two video segments divided from video content.

[0182] S1104, obtaining the barrage information, playback behavior information and user information corresponding to each video clip; the barrage information includes barrage content and barrage value information.

[0183] S1106, extracting the text vector corresponding to each barrage content; performing sentiment analysis on the text vector to obtain the content sentiment feature value of each barrage content.

[0184] S1108, determine the barrage emotional feature value corresponding to each video clip according to the content emotional feature value of each barrage content, and generate barrage attribute information of each video clip according to the barrage emotional feature value and barrage numerical information.

[0185] S1110, extracting a barrage attribute information representation from the barrage attribute information through a first extraction network included in the click rate prediction model, and extracting a playback behavior information representation from the playback behavior information.

[0186] S1112, respectively encode the barrage attribute information representation and the playback behavior information representation to obtain barrage features and playback behavior features.

[0187] S1114, extracting user-related feature representation from the user information through a second extraction network included in the click rate prediction model.

[0188] S1116, feature encoding is performed on the user-related feature representation to obtain user features of preset dimensions.

[0189] S1118, through the prediction layer included in the click rate prediction model, the bullet screen features, the playback behavior features and the user features are feature-fused to obtain the target multimodal features; based on the target multimodal features, the click rate prediction value of each video clip is determined.

[0190] S1120 , selecting click rate prediction values ​​that meet the recommendation condition from the click rate prediction values ​​of the video clips.

[0191] S1122: Determine a recommended time point based on the screened video clips corresponding to the predicted click rate values.

[0192] S1124, obtaining the bullet screen content of the video clip corresponding to the recommended time point; obtaining the description information of the object to be recommended.

[0193] S1126, extract semantic features from the barrage content to obtain barrage semantic features.

[0194] S1128, based on the semantic features and description information of the bullet comment, generate recommended content corresponding to the recommended time point.

[0195] S1130, playing the recommended content when the video content is played to the recommended time point.

[0196] In this embodiment, the click rate prediction value of each video clip is determined based on the bullet screen features, playback behavior features and user features corresponding to each video clip through a pre-trained click rate prediction model; because the bullet screen features, playback behavior features and user features can reflect the user's viewing emotions, the browsing degree of the video clip and the main user group. By combining and analyzing the bullet screen features, playback behavior features and user features of each video clip, it is possible to accurately and effectively analyze the video clips in the video content that are more suitable for content push, thereby accurately analyzing the click rate prediction value of each video clip. Determine the recommended time point based on the video clip corresponding to the screened click rate prediction value, and play the recommended content when the video content is played to the recommended time point. In this way, content recommendation can be accurately performed at the recommended time point of the analyzed video content, thereby effectively improving the information push efficiency and push accuracy.

[0197] The present application also provides an application scenario, and the application scenario applies the above-mentioned content recommendation method. Specifically, the application of the content recommendation method in the application scenario is as follows:

[0198] After the computer device obtains the video content to be processed, it divides at least two video clips from the video content, and the bullet screen information, playback behavior information and user information corresponding to each video clip. Then, the computer device extracts the bullet screen features, playback behavior features and user features corresponding to each video clip based on the bullet screen information, playback behavior information and user information through the pre-trained click rate prediction model, and determines the click rate prediction value of each video clip. The computer device then selects the click rate prediction value that meets the recommendation conditions from the click rate prediction values ​​of each video clip. The recommended time point is determined based on the video clip corresponding to the selected click rate prediction value, and the recommended content corresponding to the recommended time point is generated, and the recommended content is added to the position corresponding to the recommended time point of the video content.

[0199] The content of the bullet comment can be information in a preset format, such as text, graphics, icons, or a combination of text and graphics. The content of the bullet comment also includes attribute information such as a preset display position, display format, and display duration. For example, the display format includes a corner mark, a screen pressure bar, and the like.

[0200] When the user browses the video content through the corresponding user terminal, after the video content is loaded and played in the user terminal, when the video content is played to the recommended time point, the recommended content is played in the video content in an interstitial manner at the preset position of the video content in the corresponding display format. In this way, content recommendation can be made in the video content accurately and effectively.

[0201] The present application also provides an application scenario, and the application scenario applies the above-mentioned content recommendation method. Specifically, the application of the content recommendation method in the application scenario is as follows:

[0202] After the computer device obtains the video content to be processed, it divides at least two video clips from the video content, and the bullet screen information, playback behavior information and user information corresponding to each video clip. Then, through the pre-trained click rate prediction model, the bullet screen features, playback behavior features and user features corresponding to each video clip are extracted based on the bullet screen information, playback behavior information and user information, and the click rate prediction value bullet screen information, playback behavior information and user information of each video clip are determined. The computer device then selects the click rate prediction value that meets the recommendation conditions from the click rate prediction values ​​of each video clip. The recommended time point is determined based on the video clip corresponding to the selected click rate prediction value, and the bullet screen recommendation content corresponding to the recommended time point is generated, and the bullet screen recommendation content is added to the position corresponding to the recommended time point of the video content.

[0203] When the user browses the video content through the corresponding user terminal, after the video content is loaded and the user terminal turns on the bullet screen display function, when the video content is played to the recommended time point, the recommended content is played in the bullet screen area of ​​the video content. In this way, content recommendation can be made accurately and effectively in the bullet screen area of ​​the video content.

[0204] It should be understood that although Figure 2 , 7 The steps in the flowcharts of , 8, and 11 are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 , 7 At least part of the steps in , 8, and 11 may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0205] In one embodiment, Fig.12As shown, a content recommendation device 1200 is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an information acquisition module 1202, a click rate prediction module 1204, a recommendation processing module 1206 and a content display module 1208, wherein:

[0206] The information acquisition module 1202 is used to acquire at least two video segments divided from the video content;

[0207] The click rate prediction module 1204 is used to determine the click rate prediction value of each video segment based on the bullet comment features, playback behavior features and user features corresponding to each video segment through a pre-trained click rate prediction model;

[0208] The recommendation processing module 1206 is used to filter the click rate prediction values ​​that meet the recommendation conditions from the click rate prediction values ​​of each video segment; and determine the recommendation time point based on the video segment corresponding to the filtered click rate prediction value;

[0209] The content display module 1208 is used to play the recommended content when the video content is played to the recommended time point.

[0210] In one embodiment, the information acquisition module 1202 is also used to obtain the barrage information, playback behavior information and user information corresponding to each video clip; the barrage information includes barrage content and barrage numerical information; based on the barrage content, the barrage emotional feature value of each video clip is determined; according to the barrage emotional feature value and barrage numerical information, the barrage attribute information of each video clip is generated.

[0211] In one embodiment, the information acquisition module 1202 is also used to extract the text vector corresponding to each barrage content; perform sentiment analysis on the text vector to obtain the content sentiment feature value of each barrage content; and determine the barrage sentiment feature value corresponding to each video clip based on the content sentiment feature value of each barrage content.

[0212] In one embodiment, the click-through rate prediction module 1204 is also used to extract barrage features and playback behavior features based on barrage attribute information and playback behavior information through the first extraction network included in the click-through rate prediction model; extract user features based on user information through the second extraction network included in the click-through rate prediction model; and determine the click-through rate prediction value of each video clip based on the barrage features, playback behavior features and user features through the prediction layer included in the click-through rate prediction model.

[0213] In one embodiment, the click rate prediction module 1204 is also used to extract the barrage attribute information representation from the barrage attribute information and extract the playback behavior information representation from the playback behavior information through the first extraction network; the barrage attribute information representation and the playback behavior information representation are respectively encoded to obtain the barrage features and the playback behavior features.

[0214] In one embodiment, the click rate prediction module 1204 is further used to extract user-related feature representations from user information through a second extraction network; perform feature encoding on the user-related feature representations to obtain user features of preset dimensions.

[0215] In one embodiment, the click rate prediction module 1204 is also used to fuse the bullet screen features, playback behavior features and user features through the prediction layer to obtain target multimodal features; and determine the click rate prediction value of each video clip based on the target multimodal features.

[0216] In one embodiment, Fig.13 As shown, the above-mentioned content recommendation device 1200 also includes a content generation module 1207, which is used to obtain the bullet screen content of the video clip corresponding to the recommended time point; based on the bullet screen content, generate the recommended content corresponding to the recommended time point.

[0217] In one embodiment, the content generation module 1207 is also used to obtain description information of the object to be recommended; extract semantic features of the barrage content to obtain barrage semantic features; and generate recommended content corresponding to the recommended time point based on the barrage semantic features and description information.

[0218] In one embodiment, the recommended content is bullet screen recommended content; the content display module 1208 is also used to play the bullet screen recommended content in the bullet screen area of ​​the video content when the video content is played to the recommended time point.

[0219] In one embodiment, the click rate prediction model is obtained by training through a training step, such as Fig.14 As shown, the above-mentioned content recommendation device 1200 also includes a model training module 1201, which is used to obtain training samples and training labels; the training samples include sample barrage attribute information, sample playback behavior information and sample user information corresponding to each sample video clip in the sample video content; the training label is the historical click-through rate of the sample recommended content in the sample video content; the click-through rate prediction model is trained based on the training samples and training labels.

[0220] In one embodiment, the model training module 1201 is also used to extract sample barrage features of sample barrage attribute information and sample playback behavior features of sample playback behavior information through a first extraction network included in the click-through rate prediction model; extract sample user features of sample user information through a second extraction network included in the click-through rate prediction model; determine the sample click-through rate of each sample video clip based on the sample barrage features, sample playback behavior features and sample user features through the prediction layer included in the click-through rate prediction model; based on the difference between the sample click-through rate and the training label, adjust the parameters of the click-through rate prediction model and continue training until the training conditions are met and stop training.

[0221] For the specific definition of the content recommendation device, please refer to the definition of the content recommendation method above, which will not be repeated here. Each module in the above content recommendation device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0222] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.15 As shown. The computer device includes a processor, a memory and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as video content, barrage information, playback behavior information, user information and recommended content. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a content recommendation method is implemented.

[0223] Those skilled in the art will understand that Fig.15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0224] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0225] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0226] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0227] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0228] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A content recommendation method, characterized in that: The method comprises: Acquire at least two video segments divided from the video content; Extracting bullet screen features and playback behavior features based on bullet screen attribute information and playback behavior information of each video clip through a first extraction network included in the click rate prediction model; Extracting user features based on user information through a second extraction network included in the click rate prediction model; Determining the click rate prediction value of each video clip according to the bullet comment feature, the playback behavior feature and the user feature through the prediction layer included in the click rate prediction model; Selecting, from the click rate prediction values ​​of the video clips, click rate prediction values ​​that meet the recommendation condition; Determine the recommended time point based on the video clips corresponding to the filtered click rate prediction values; When the video content is played to the recommended time point, the recommended content is played.

2. The method according to claim 1, characterized in that After obtaining at least two video segments divided from the video content, the method further includes: Obtaining the barrage information, playback behavior information and user information corresponding to each of the video clips; the barrage information includes barrage content and barrage value information; Based on the content of the bullet screen, determining the emotional feature value of the bullet screen of each video clip; The barrage attribute information of each video clip is generated according to the barrage emotional feature value and the barrage numerical information.

3. The method according to claim 2, characterized in that The determining of the emotional feature value of the bullet screen of each video clip based on the bullet screen content includes: Extracting text vectors corresponding to the contents of each bullet comment; Performing sentiment analysis on the text vector to obtain sentiment feature values ​​of the content of each bullet comment; According to the content emotion feature values ​​of each barrage content, the barrage emotion feature values ​​corresponding to each video clip are determined.

4. The method according to claim 1, characterized in that: The first extraction network included in the click rate prediction model extracts bullet screen features and playback behavior features based on bullet screen attribute information and playback behavior information of each video clip, including: extracting, by the first extraction network, a barrage attribute information representation from the barrage attribute information, and extracting a playback behavior information representation from the playback behavior information; The barrage attribute information representation and the playback behavior information representation are respectively encoded to obtain barrage features and playback behavior features.

5. The method according to claim 1, characterized in that The extracting user features based on user information through the second extraction network included in the click rate prediction model includes: extracting user-related feature representation from the user information through the second extraction network; Feature encoding is performed on the user-related feature representation to obtain user features of preset dimensions.

6. The method according to claim 1, characterized in that The prediction layer included in the click rate prediction model determines the click rate prediction value of each video clip according to the bullet screen feature, the playback behavior feature and the user feature, including: Through the prediction layer, the bullet screen features, the playback behavior features and the user features are fused to obtain target multimodal features; A click rate prediction value of each video segment is determined based on the target multimodal feature.

7. The method according to claim 1, characterized in that After determining the recommended time point based on the video clips corresponding to the filtered click rate prediction values, the method further includes: Obtaining the bullet screen content of the video clip corresponding to the recommended time point; Based on the bullet screen content, the recommended content corresponding to the recommended time point is generated.

8. The method according to claim 7, characterized in that The determining, based on the bullet screen content, the recommended content corresponding to the recommended time point includes: Get the description information of the object to be recommended; Extracting semantic features from the barrage content to obtain barrage semantic features; Based on the barrage semantic features and the description information, the recommended content corresponding to the recommended time point is generated.

9. The method according to any one of claims 1 to 8, characterized in that: The recommended content is the barrage recommended content; The playing of the recommended content when the video content is played to the recommended time point includes: When the video content is played to the recommended time point, the bullet screen recommended content is played in the bullet screen area of ​​the video content.

10. The method according to any one of claims 1 to 8, characterized in that: The click rate prediction model is obtained through training steps, and the training steps include: Acquire training samples and training labels; the training samples include sample bullet screen attribute information, sample playback behavior information and sample user information corresponding to each sample video clip in the sample video content; the training label is the historical click rate of the sample recommended content in the sample video content; A click rate prediction model is trained based on the training samples and the training labels.

11. The method according to claim 10, characterized in that The step of training the click rate prediction model based on the training samples and the training labels includes: Extracting, by means of a first extraction network included in the click rate prediction model, sample bullet-screen features of the sample bullet-screen attribute information and sample playback behavior features of the sample playback behavior information; Extracting sample user features of the sample user information through a second extraction network included in the click rate prediction model; Determining the sample click rate of each sample video clip based on the sample bullet comment feature, the sample playback behavior feature and the sample user feature through the prediction layer included in the click rate prediction model; Based on the difference between the sample click rate and the training label, the parameters of the click rate prediction model are adjusted and training is continued until the training condition is met and the training is stopped.

12. A content recommendation device, characterized in that: The device comprises: An information acquisition module, used to acquire at least two video segments divided from the video content; A click rate prediction module, configured to extract bullet screen features and playback behavior features based on bullet screen attribute information and playback behavior information of each video segment through a first extraction network included in the click rate prediction model; extract user features based on user information through a second extraction network included in the click rate prediction model; and determine a click rate prediction value of each video segment based on the bullet screen features, the playback behavior features, and the user features through a prediction layer included in the click rate prediction model; A recommendation processing module, used to filter out click rate prediction values ​​that meet the recommendation conditions from the click rate prediction values ​​of the video clips; and determine the recommendation time point based on the video clips corresponding to the filtered click rate prediction values; The content display module is used to play the recommended content when the video content is played to the recommended time point.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Information recommendation method and device, server and user terminal

    CN106572399A

  • Method and device for inserting recommendation information in video

    CN108235126A

  • Barrage-based video recommendation method and device

    CN108737859A

  • Content recommendation method and device, electronic equipment and storage medium

    CN111177575A