A method and device for processing bullet screen and electronic equipment

By identifying audience category characteristics and bullet screen characteristics based on user and video information during video playback, the system can filter out target bullet screens, solving the problem of poor video bullet screen filtering and improving the user viewing experience.

CN116320631BActive Publication Date: 2025-12-19BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310192656.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-12-19
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

In existing technologies, the filtering effect of video bullet comments is not good, making it difficult to effectively filter out bullet comments that users are interested in, resulting in a poor user viewing experience.

Method used

Based on user and video information, the target audience category features are determined from the set of audience category features, the features of video bullet comments are extracted, and the target bullet comments are filtered out according to the matching degree and displayed to the user.

Benefits of technology

It improves the accuracy of video bullet comment filtering, effectively displaying bullet comments that are highly relevant to users, thus enhancing the user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116320631B_ABST
    Figure CN116320631B_ABST
Patent Text Reader

Abstract

The application provides a method and device for processing barrage and electronic equipment, comprising: determining a target crowd category feature from a crowd category feature set corresponding to a target video based on user information of a user and video information of the target video; performing feature extraction on barrage information corresponding to each video barrage in a video barrage set of the target video to obtain video barrage features corresponding to the video barrage; determining a barrage matching degree according to the target crowd category feature and the video barrage feature; screening a target barrage from the video barrage set based on the barrage matching degree; and outputting the target barrage. In the application, the same user can correspond to different crowd category features for different target videos, and the target barrage is matched from the barrage of the target video and displayed to the user through the targeted crowd category feature, which effectively improves the accuracy when the video barrage is screened and can effectively screen the target barrage with a higher matching degree with the user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of information processing, and particularly relates to a bullet screen processing method and device and electronic equipment. BACKGROUND

[0002] With the continuous development of Internet technology, there are more and more video contents on the network. When watching a video, a user often leaves a bullet screen (such as a video message, a video bullet screen, etc.) for the video, and interacts through the video bullet screen on the video page. With the increase in the number of video viewers, a video can have a large number of video bullet screens, which are filled with a large amount of bullet screen content that is not interesting to users, and can easily disturb the users.

[0003] In related technologies, in order to reduce the display amount of video bullet screens, a bullet screen filtering function is usually provided for a user, and the user can perform a filtering operation on the video bullet screen through filtering options such as setting a filtering keyword and filtering a user level. For example, the user can set the filtering user level to level 10, and the bullet screen left by other users at a level below 10 will not be displayed, so as to reduce the interference on the user and reduce the display amount of video bullet screens.

[0004] However, it is difficult to effectively filter the video bullet screen through simple filtering options, and there can still be a lot of bullet screen that is not interesting to the user in the video bullet screen displayed to the user, and the video bullet screen that is interesting to the user can also be filtered, which can cause poor filtering effect of the video bullet screen. SUMMARY

[0005] Therefore, the present application provides a bullet screen processing method and device and electronic equipment, which can solve the problem of poor filtering effect when filtering a video bullet screen.

[0006] According to a first aspect of the present application, a bullet screen processing method is provided, and the method comprises the following steps.

[0007] When a user plays a target video, a target crowd category feature is determined from a crowd category feature set corresponding to the target video based on user information of the user and video information of the target video.

[0008] Features of bullet screen information corresponding to each video bullet screen in a video bullet screen set of the target video are extracted to obtain video bullet screen features corresponding to the video bullet screen.

[0009] A bullet screen matching degree is determined according to the target crowd category feature and the video bullet screen feature.

[0010] A target bullet screen is filtered from the video bullet screen set based on the bullet screen matching degree.

[0011] The target bullet screen is output.

[0012] According to a second aspect of the present application, a device for processing bullet screen is provided, the device comprising:

[0013] a category feature module configured to determine a target crowd category feature from a crowd category feature set corresponding to a target video based on user information of a user and video information of the target video when the user plays the target video;

[0014] a bullet screen feature module configured to extract features from bullet screen information corresponding to each video bullet screen in a video bullet screen set of the target video to obtain video bullet screen features corresponding to the video bullet screen;

[0015] a bullet screen matching degree module configured to determine a bullet screen matching degree based on the target crowd category feature and the video bullet screen features;

[0016] a target bullet screen module configured to screen a target bullet screen from the video bullet screen set based on the bullet screen matching degree;

[0017] a bullet screen output module configured to output the target bullet screen.

[0018] In a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the first aspect.

[0019] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to enable a computer to perform the steps of the first aspect.

[0020] In a fifth aspect, a computer program product is provided, comprising a computer program, and the computer program, when executed by a processor, implements the steps of the first aspect.

[0021] Compared with the prior art, the present application has the following advantages:

[0022] The barrage processing method, device and electronic equipment provided in the application can determine a target crowd category feature from a crowd category feature set corresponding to a target video based on user information of a user and video information of the target video when the user plays the target video, perform feature extraction on barrage information corresponding to each video barrage in a video barrage set of the target video to obtain video barrage features corresponding to the video barrage, determine a barrage matching degree according to the target crowd category feature and the video barrage features, filter target barrages from the video barrage set based on the barrage matching degree, and output the target barrages. The user can be determined to have a crowd category feature for the target video, and the video barrage can be filtered according to the crowd category feature, so that the same user can have different crowd category features for different target videos. The target barrage is matched from the barrage of the target video and displayed to the user through the targeted crowd category feature, which effectively improves the accuracy when the video barrage is filtered and can effectively filter the target barrage with a higher matching degree to the user.

[0023] The above description is only a summary of the technical solutions of the application. In order to make the technical solutions of the application more clearly understood and implemented according to the content of the description, and in order to make the above and other purposes, features and advantages of the application more obvious and easy to understand, the following detailed description of the specific embodiments of the application is given. BRIEF DESCRIPTION OF DRAWINGS

[0024] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not meant to limit the present application. Moreover, the same reference numerals in the attached drawings indicate the same or similar components. In the drawings:

[0025] Figure 1 FIG. 1 is a step flowchart of a barrage processing method provided by an embodiment of the application;

[0026] Figure 2 FIG. 2 is a model structure diagram of an initial neural network model provided by an embodiment of the application;

[0027] Figure 3 FIG. 3 is a step flowchart of another barrage processing method provided by an embodiment of the application;

[0028] Figure 4 FIG. 4 is a step flowchart of a model training method provided by an embodiment of the application;

[0029] Figure 5 FIG. 5 is an architecture diagram of a barrage processing system provided by an embodiment of the application;

[0030] Figure 6 FIG. 6 is a block diagram of a barrage processing device provided by an embodiment of the application. DETAILED DESCRIPTION

[0031] Exemplary embodiments of the present application will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it is understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thoroughly and completely understood, and will fully convey the scope of the application to those skilled in the art.

[0032] Figure 1 is a step flow chart of a bullet screen processing method provided by an embodiment of the present application. As shown in Figure 1 , the method can include:

[0033] Step 101, when a user plays a target video, determining a target crowd category feature from a crowd category feature set corresponding to the target video based on user information of the user and video information of the target video.

[0034] The user information of the user can include user identification information and user account information, wherein the user identification information can include a user name or a user identity document (UID), etc., and the user account information can include gender information, age information, etc. set by the user for the account.

[0035] The target video can represent a video being watched by the user, can also represent a video collected by the user, and can also represent a video added to a “want to watch list” by the user. The target video is not specifically limited by the embodiments of the present application, and the range of the target video can be set according to actual needs.

[0036] The video information of the target video can include video identification information and video content information, wherein the video identification information can include a video name or a video identification code, etc., and the video content information can include video introduction information, video tag information, video category information, etc.

[0037] In the embodiments of the present application, each video can correspond to a crowd category feature set, and the crowd category feature set stores a plurality of crowd category features. The crowd category feature set corresponding to the target video can be first determined from the crowd category feature set library according to the video representation information of the target video, and then the video user feature can be determined according to the user information of the user and the video information of the target video, and then the video user feature is matched with each crowd category feature in the crowd category feature set corresponding to the target video, and the target crowd category feature corresponding to the user is determined according to the feature similarity. The crowd category feature is obtained by classifying the audience of the video and extracting the personal information and / or behavior information of different categories of crowds, for example, the audience of a video can be classified according to age, video watching time and other information, and the age features and video watching time features of different crowd categories are extracted and fused to obtain the crowd category feature of the video, which is not limited in the embodiments of the present application. In the embodiments of the present application, different videos can correspond to different crowd category feature sets, and the crowd category feature set library can store the crowd category feature sets corresponding to each video.

[0038] It should be noted that the video user feature is determined by the user information of the user and the video information of the target video, that is, for the same user, the video user feature determined for different target videos can also be different, and further, the target crowd category feature corresponding to the same user for different target videos can also be different.

[0039] For example, the crowd category feature set corresponding to a video A can include three crowd category features, namely crowd category feature A1, crowd category feature A2 and crowd category feature A3, and the crowd category feature set corresponding to another video B can also include three crowd category features, namely crowd category feature B1, crowd category feature B2 and crowd category feature B3. If the target video is video B, the video user feature determined by the user information of the user and the video information of the target video is closest to the crowd category feature B2 of video B, and then the crowd category feature B2 can be determined as the target crowd category feature corresponding to the user.

[0040] Step 102, extracting features of the comment information corresponding to each video comment in the video comment set of the target video to obtain the video comment feature corresponding to the video comment.

[0041] Due to the large number of viewers of the same video, a video often has a large number of video bullet screens, and if all video bullet screens of the video are displayed to the user, the information amount will be too large, and a large number of video bullet screens contain content that is not interesting to the user, which is not conducive to the user to obtain useful information. For video bullet screens, one moment of a video can correspond to hundreds or thousands of video bullet screens, and simultaneously displaying these video bullet screens can easily block the video screen, resulting in poor user viewing experience.

[0042] Therefore, in the embodiments of the present application, the target video bullet screens can be filtered according to the target user group category features determined in step 101, and a part of the video bullet screens that are most interesting to the user among all the video bullet screens of the target video are displayed to the user, so as to improve the user's experience in watching the video bullet screens of the target video, and avoid interference of a large number of video bullet screens on the content of the target video.

[0043] Specifically, the bullet screen information of each video bullet screen in the video bullet screen set of the target video can be obtained first, and then the bullet screen information is feature extracted to obtain the video bullet screen features corresponding to each video bullet screen. The bullet screen information of the video bullet screen can include any information related to the video bullet screen in addition to the bullet screen content, such as the font of the video bullet screen, the color of the video bullet screen, the video information of the video to which the video bullet screen belongs (the video information of the target video), the user information of the user to which the video bullet screen belongs, the number of likes of the video bullet screen, and the number of dislikes of the video bullet screen, etc.

[0044] After obtaining the bullet screen information of the video bullet screen, the bullet screen information is feature extracted to obtain the video bullet screen features corresponding to each video bullet screen of the target video. It should be noted that if the bullet screen information only includes the bullet screen content, the bullet screen content can be feature extracted to obtain the video bullet screen features, if the video bullet screen also includes other information, the bullet screen content and the other information can be feature extracted respectively, and then the extracted features are fused to obtain the video bullet screen features, or the bullet screen content and the other information can be feature extracted as a whole to obtain the video bullet screen features. The above feature extraction methods can include semantic feature extraction, sparse feature extraction, dense feature extraction, etc., and the embodiments of the present application do not limit the above feature extraction methods, and the user can select a suitable feature extraction method according to actual needs.

[0045] Step 103, determining a bullet screen matching degree according to the target user group category features and the video bullet screen features.

[0046] After obtaining the target crowd category feature of the target video and the video comment features of the target video, in order to determine the matching degree between the video comment and the user, the comment matching degree of each video comment can be determined according to the target crowd category feature and the video comment feature. The greater the comment matching degree of the video comment, the higher the matching degree between the video comment and the user, and the more interested the user is in the video comment.

[0047] Specifically, in an embodiment, the target crowd category feature and the video comment feature of the video comment can be compared to obtain the similarity between them, and the similarity can be taken as the comment matching degree of the video comment. In another embodiment, the target crowd category feature and the video comment feature of the video comment can be mathematically operated to obtain the comment matching degree. For example, if the target crowd category feature and the video comment feature are both in the form of a vector, the dot product operation can be performed on the target crowd category feature and the video comment feature to obtain the comment matching degree. The skilled person can also determine the above-mentioned comment matching degree in other ways, which are not specifically limited in the embodiments of the present application.

[0048] Step 104: screening the target comment from the video comment set based on the comment matching degree.

[0049] After obtaining the comment matching degree corresponding to each video comment in the video comment set, the video comment set can be screened according to the comment matching degree corresponding to each video comment, so as to obtain the target comment that the user is more interested in.

[0050] Specifically, in an embodiment, the video comment set can be screened according to the size of the comment matching degree, and the fixed screening number of video comments with the largest comment matching degree can be screened from the video comment set as the target video comment. For example, the video comment set of the target video contains five video comments: video comment 1 to video comment 5, wherein the comment matching degree of the video comment 1 is 0.2, the comment matching degree of the video comment 2 is 0.3, the comment matching degree of the video comment 3 is 0.6, the comment matching degree of the video comment 4 is 0.8, and the comment matching degree of the video comment 5 is 0.9. In the case of a fixed screening number of 2, the target video comment can contain the video comment 4 and the video comment 5.

[0051] In another implementation, the video barrage with the largest barrage matching degree can be selected from the video barrage set as the target video barrage according to the size of the barrage matching degree. For example, the video barrage set of the target video contains five video barrages: video barrage 1 to video barrage 5, wherein the barrage matching degree of the video barrage 1 is 0.2, the barrage matching degree of the video barrage 2 is 0.3, the barrage matching degree of the video barrage 3 is 0.6, the barrage matching degree of the video barrage 4 is 0.8, and the barrage matching degree of the video barrage 5 is 0.9. In the case of a fixed screening ratio of 10%, the target video barrage can be determined to contain the video barrage 5.

[0052] In step 105, the target barrage is output.

[0053] During the process of watching the target video, the target video barrage can be displayed to the user, and the barrages other than the target video barrage can be hidden to avoid interference with the user watching the target video.

[0054] Specifically, since each target barrage corresponds to a barrage content and a display time, when the video playing progress matches the display time of the target barrage, the target barrage is superimposed and displayed in the playing picture of the target video.

[0055] In the above example, if it is determined that the target barrage contains the video barrage 4 and the video barrage 5, and the display time of the video barrage 4 is 1 minute and 18 seconds and the display time of the video barrage 5 is 2 minutes and 32 seconds, the video barrage 4 is displayed when the playing progress of the target video reaches 1 minute and 18 seconds, and the video barrage 5 is displayed when the playing progress of the target video reaches 2 minutes and 32 seconds.

[0056] In summary, in the barrage processing method provided by the embodiments of the present application, the target crowd category feature of the user can be determined for the target video based on the user information of the user and the video information of the target video when the user plays the target video, the barrage information corresponding to each video barrage in the video barrage set of the target video is subjected to feature extraction to obtain the video barrage feature corresponding to the video barrage, the barrage matching degree is determined according to the target crowd category feature and the video barrage feature, the target barrage is selected from the video barrage set based on the barrage matching degree, and the target barrage is output. The user can be determined to belong to the crowd category feature for the target video, and the video barrage is filtered according to the crowd category feature, so that the same user can correspond to different crowd category features for different target videos. The target barrage is matched from the barrage of the target video and displayed to the user through the targeted crowd category feature, which effectively improves the accuracy when the video barrage is screened and can effectively screen the target barrage with a higher matching degree for the user.

[0057] In the embodiment of the present application, the user barrage matching model can also be obtained by training the initial neural network model. The user barrage matching model can be used to implement the above barrage processing method, that is, to filter the target barrage of interest of the user from the video barrage set of the target video.

[0058] Referring to Figure 2 , Figure 2 A model structure diagram of a user barrage matching model provided by an embodiment of the present application is shown. As shown in Figure 2 , the user barrage matching model can include a first sub-network model, a second sub-network model, a third sub-network model, and a feature matching sub-model. The first sub-network model is used to determine video user features according to video information and user information. The second sub-network model is used to determine target crowd category features from a crowd category feature set corresponding to a target video. The third sub-network model is used to determine video barrage features corresponding to the video barrage of the target video. The feature matching sub-model is used to determine a barrage matching degree between the target crowd category features and the video barrage features.

[0059] Figure 3 is a step flowchart of another barrage processing method provided by an embodiment of the present application. As shown in Figure 3 , the method can include:

[0060] In step 201, the video information and the user information are input into the first sub-network model of the user barrage matching model, and the video user features output by the first sub-network model are obtained.

[0061] The above feature extraction process can be based on a feature extraction model, which can include a sparse feature extraction model, a dense feature extraction model, a semantic feature extraction model, a feature encoding model (such as an embedding model), etc. For example, an embedding model can be used to extract features of the user identification information and the video identification information, and obtain the user identification information features and the video identification information features represented in an encoding form. A sparse feature extraction model can be used to extract features of the gender information of the user account.

[0062] In the embodiment of the present application, the first sub-network model can be used to extract user features based on the user information, extract video features based on the video information, and fuse the user features and the video features to obtain the video user features.

[0063] In the embodiment of the present application, as shown in Figure 2 , the first sub-network model can include a first feature splicing layer and a first normalization layer. The first feature splicing layer is used to receive the input sample video information and sample user information and output first spliced features. The first normalization layer is used to normalize the first spliced features to obtain the video user features.

[0064] Sub-step 2011, input the video information and the user information into the first feature splicing layer to obtain the first splicing feature output by the first feature splicing layer; wherein the first feature splicing layer is configured to determine the video feature according to the video information, determine the user feature according to the user information, and splice the video feature and the user feature into the first splicing feature.

[0065] The user feature and the video feature can be spliced first to obtain the first splicing feature. For example, the features are usually described in the form of vectors, and the vectors usually appear in the form of arrays. The concat function can be used to splice the user feature and the video feature. The concat function is used to concatenate two or more arrays, and this method can return a new array, i.e., the first splicing feature. In addition, other methods can also be used to splice the user feature and the video feature, and the embodiments of the present application do not make specific limitations in this regard.

[0066] Specifically, the first feature splicing layer can be composed of a first feature extraction layer and a first splicing layer. The first feature extraction layer can further include a plurality of feature extraction modules and / or feature encoding modules. For example, as shown in FIG. 2, the first feature extraction layer can include a video identification information encoding module, a user identification information encoding module, and a user account information feature extraction module. The video identification information encoding module is configured to perform feature encoding on the video identification information in the sample video information. The user identification information encoding module is configured to perform feature encoding on the user identification information in the sample user information. The user account information feature extraction module is configured to perform feature extraction on the user account information in the sample user information. The above-mentioned encoding modules can perform embedding encoding. Figure 2 The first splicing layer can be constructed based on the concat function, and is configured to combine the features output by the feature extraction modules and / or the feature encoding modules in the first feature extraction layer to obtain the first splicing feature.

[0067] Sub-step 2012, input the first splicing feature into the first normalization layer to obtain the video user feature output by the first normalization layer.

[0068] In the embodiments of the present application, after obtaining the user feature of the user and the video feature of the target video, the user feature and the video feature can be fused to obtain the video user feature of the user for the target video.

[0069] The following provides a specific feature fusion method.

[0070]

[0071] ​To simplify subsequent operations and improve operation efficiency, the first spliced feature can be input into a first normalization layer for normalization processing to obtain a video user feature. The normalization layer can be implemented by a normalization function nn.LayerNorm.

[0072] Further, the first normalization layer can be further provided with a first fully connected layer. The first spliced feature can be first input into the first fully connected layer, and the output of the first fully connected layer can be taken as the input of the first normalization layer, and finally the video user feature output by the first normalization layer can be obtained. The fully connected layer (FC) can integrate the multi-dimensional features inputted, and output the one-dimensional feature to the subsequent first normalization layer after integrating the features of each dimension, so as to reduce the influence of feature position on the subsequent processing result. By increasing the first fully connected layer, the accuracy of the video user feature can be significantly improved.

[0073] In the embodiment of the present application, the first normalization layer can perform normalization operation on the first spliced feature to obtain the final video user feature. Specifically, the first normalization layer can include an nn.LayerNorm function layer, and the normalization operation can be performed by the nn.LayerNorm function.

[0074] Further, the first normalization layer can further include an L2 regularization (L2 Regularization) sublayer. The L2 regularization sublayer can be arranged at the end of the first normalization layer. Through the L2 regularization sublayer, an additional penalty term can be added to the model to prevent overfitting.

[0075] Further, as shown in Figure 2 The first spliced feature can be first input into the first fully connected layer, and the output of the first fully connected layer can be taken as the input of the first normalization layer, and then the video user feature output by the first normalization layer can be obtained.

[0076] In step 202, the video user feature and the video information are input into a second sub-network model of the user barrage matching model to obtain the target crowd category feature output by the second sub-network model.

[0077] As shown in Figure 2As shown, the second sub-network model can further include a three-dimensional feature matrix, a crowd category feature module, and a one-hot encoding module. The one-hot encoding module is configured to determine a corresponding target one-hot encoding of the target video according to the target video information. The crowd category feature module is configured to determine a crowd category feature set corresponding to the target video from the three-dimensional feature matrix according to the target one-hot encoding module. The crowd category feature module is further configured to determine a target crowd category feature of the user for the target video from the crowd category feature set according to the video user feature.

[0078] In sub-step 2021, the video information is input into the one-hot encoding module to obtain a target one-hot encoding output by the one-hot encoding module.

[0079] One-hot encoding is a way of encoding objects, also known as one-bit valid encoding. The method is to use N-bit status bits to encode N objects, each object has its own valid bit, and only one valid bit in each one-hot encoding.

[0080] For example, if two videos are one-hot encoded, the one-hot encodings of the two videos are "10" and "01" respectively. If three videos are one-hot encoded, the one-hot encodings of the three videos are "100", "010" and "001" respectively. If four videos are one-hot encoded, the one-hot encodings of the four videos are "1000", "0100", "0010" and "0001" respectively. And so on.

[0081] All videos in the video library can be one-hot encoded, and a one-hot encoding library corresponding to the video library can be established according to the video identification of the video and the one-hot encoding corresponding to the video. The video library can be all videos of a video provider, or all videos under a certain video tag or video category, for example, the video library can be a movie video library. Then, in the embodiments of the present application, the one-hot encoding library can be queried according to the video identification of the target video to obtain the one-hot encoding corresponding to the target video.

[0082] In the embodiments of the present application, the target one-hot encoding module corresponding to the target video can be determined by the one-hot encoding module, and the video information of the target video can be input into the target one-hot encoding to obtain the corresponding target one-hot encoding output by the target one-hot encoding module. It should be noted that in the training stage of the model, the number of state bits of the one-hot encoding can be determined according to the number of sample videos. If there are 1000 sample videos, the number of state bits of the one-hot encoding is 1000.

[0083] In actual application, as the use time increases, new videos are continuously added to the video library. Even if all the videos in the video library are used as sample videos during model training, the newly added videos in the video library cannot be matched to the one-hot encoding. Therefore, in the embodiment of the present application, a first proportion of sample videos in all sample videos can be used as first sample videos, and a second proportion of sample videos in the sample videos can be used as second sample videos, wherein the sum of the first proportion and the second proportion is 100%. For example, the first proportion can be 95%, and the second proportion can be 5%. The number of state bits of one-hot encoding is determined as the number of first sample videos + 1, the first sample videos are encoded by the state bits of the number of first sample videos before one-hot encoding, and all videos except the first sample videos are encoded by the last state bit.

[0084] For example, if the sample videos include 10 videos, and the first proportion is 80%, the number of first sample videos is 8, and the number of state bits of one-hot encoding is 8 + 1 = 9 bits. The one-hot encoding of the 8 first sample videos can be set as "100000000", "010000000",..., "000000010", and any remaining video is set as "000000001". The "000000001" can be referred to as a supplementary one-hot encoding, that is, the valid bit of the supplementary one-hot encoding is the last bit. In this way, for the target video that is the same as the first sample video, a unique corresponding one-hot encoding can be determined, and for the target video that is different from the first sample video, the supplementary one-hot encoding can be used as the target one-hot encoding thereof.

[0085] In step 2022, based on the position of the valid bit in the target one-hot encoding, a target two-dimensional feature matrix corresponding to the target video is extracted from the three-dimensional feature matrix. The three-dimensional feature matrix is composed of two-dimensional feature matrices corresponding to each encoding bit in the target one-hot encoding. The two-dimensional feature matrix is composed of a first preset number of columns of crowd characteristic elements, and each column of crowd characteristic elements represents a crowd category, and each column of crowd characteristic elements includes a second preset number of crowd category characteristic elements.

[0086] Based on the above one-hot encoding, in the embodiment of the present application, the crowd category feature set corresponding to each video can be stored in the form of a three-dimensional feature matrix. The three-dimensional feature matrix includes three dimensions, namely the state bit dimension, the crowd category dimension, and the crowd category feature dimension. Each two-dimensional feature matrix of the three-dimensional feature matrix corresponds to a state bit, and the number of two-dimensional feature matrices corresponds to the number of state bits in one-hot encoding. The two-dimensional feature matrix is composed of a first preset number of columns of crowd category characteristic elements. Each column of crowd category characteristic elements represents a crowd category, and each column of crowd category characteristic elements includes a second preset number of crowd category characteristic elements.

[0087] Since each state bit corresponds to a first preset number of crowd category elements in the three-dimensional feature matrix, each crowd category element corresponding to each state bit also corresponds to a second preset number of crowd category feature elements, each state bit corresponds to a two-dimensional feature matrix, which is composed of all crowd category crowd category feature elements corresponding to the state bit in the two-dimensional feature matrix.

[0088] After obtaining the target one-hot encoding corresponding to the target video, the position of the effective bit in the target one-hot encoding can be determined. For example, the target one-hot encoding is composed of 10 state bits, specifically "0000001000", and the effective bit in the target one-hot encoding is the 7th bit among all state bits, so the position of the effective bit in the target one-hot encoding is 7. Since each two-dimensional feature matrix in the three-dimensional feature matrix has a one-to-one correspondence with each state bit in the target one-hot encoding, the corresponding two-dimensional feature matrix can be determined from the three-dimensional feature matrix according to the position of the effective bit in the target one-hot encoding. In the above example, since the position of the effective bit in the target one-hot encoding is 7, the 7th two-dimensional feature matrix in the three-dimensional feature matrix can be determined as the target two-dimensional feature matrix corresponding to the target video.

[0089] Exemplarily, in the case that the one-hot encoding contains 10 state bits, the first preset number is 4, and the second preset number is 3, the three-dimensional feature matrix can be composed of 10 two-dimensional feature matrices, and each two-dimensional feature matrix corresponds to a state bit, wherein one two-dimensional feature matrix can refer to Table 1 as shown below:

[0090]

[0091] Table 1

[0092] In Table 1, the first column contains 3 crowd category feature elements A1, A2 and A3 of crowd category A, the second column contains 3 crowd category feature elements B1, B2 and B3 of crowd category B, and so on.

[0093] It should be noted that at the beginning of model training, the three-dimensional feature matrix can be constructed according to the number of bits of the one-hot encoding, the first preset number and the second preset number, and the initial value of each element in the three-dimensional feature matrix can be assigned. In the case that the number of state bits of the one-hot encoding is the number of first sample videos + 1, since the target one-hot encoding of the second sample video is a supplementary one-hot encoding, the target two-dimensional feature matrix of all target videos different from the first sample video is the two-dimensional feature matrix corresponding to the last state bit of the one-hot encoding in the three-dimensional feature matrix.

[0094] In substep 2023, the crowd class feature elements in the same column in the target two-dimensional feature matrix are combined into a crowd class feature, to obtain a crowd class feature set of the target video.

[0095] After obtaining the target two-dimensional feature matrix corresponding to the target video, all crowd class feature elements in the same column in the target two-dimensional feature matrix can be combined into a crowd class feature, and then a same number of crowd class features as the column number of the target two-dimensional feature matrix is obtained, which constitute the crowd class feature set of the target video. The combination manner can be splicing, operation, fusion, etc., which is not specifically limited in the embodiments of the present application.

[0096] After obtaining the video user feature and the crowd class feature set of the target video, the video user feature can be matched with each crowd class feature in the crowd class feature set respectively, to determine the crowd matching degree corresponding to each crowd class feature. The greater the crowd matching degree is, the higher the matching degree of the crowd class feature corresponding to the crowd matching degree with the user is, and the more the user conforms to the crowd class feature. Therefore, the crowd class feature with the maximum crowd matching degree in the crowd class feature set can be determined as the target crowd class feature.

[0097] Since the features are usually represented by vectors, in the case that the crowd class feature and the video user feature are both vectors, the cosine distance between the two can be calculated, and the crowd matching degree can be determined based on the cosine distance (for example, the reciprocal of the cosine distance can be taken as the crowd matching degree). The smaller the cosine distance is, the more similar the two features are, and the higher the feature matching degree is. Other ways can also be used to determine the crowd matching degree, which is not specifically limited in the embodiments of the present application.

[0098] In the embodiments of the present application, the video user feature can be input into the crowd class feature module, and the video user feature is matched with the crowd class feature set by the crowd class feature module, to output the target crowd class feature of the user for the target video.

[0099] In step 203, the video information, the user information and the barrage information are input into a third sub-network model of the user barrage matching model, to obtain the video barrage feature output by the third sub-network model.

[0100] The third sub-network model is used to extract the video user feature based on the user information, extract the video feature based on the video information, extract the barrage information corresponding to each video barrage in the video barrage set of the target video, to obtain the barrage feature of each video barrage, and fuse the user feature, the video feature and the barrage feature, to obtain the video barrage feature corresponding to the video barrage.

[0101] Optionally, step 203 can include:

[0102] In substep 2031, the video information, the user information and the barrage information are input into the second feature concatenation layer to obtain second concatenation features output by the second feature concatenation layer; wherein the second feature concatenation layer is configured to determine video features according to the video information, determine user features according to the user information, determine barrage features according to the barrage information, and concatenate the video features, the user features and the barrage features.

[0103] In the embodiments of the present application, the video barrage features of the video barrage can be determined based on the user information corresponding to the user who is watching the video, the video information of the video to which the video barrage belongs and the barrage information of the video barrage itself. Based on this, the user information can be feature-extracted to obtain user features, and the video information can be feature-extracted to obtain video features. For example, the user information can include user identification information and various user account information, the user identification information can be feature-extracted to obtain user identification information features, and each user account information can be feature-extracted to obtain user account information features corresponding to each user account information, thereby obtaining user features composed of user identification information features and user account information features. The video information can include video identification information, and the video identification information can be feature-extracted to obtain video features.

[0104] It should be noted that the above feature extraction process can be based on a feature extraction model, wherein the feature extraction model can include a sparse feature extraction model, a dense feature extraction model, a semantic feature extraction model, a feature encoding model (such as an embedding model), etc. For example, the embedding model can be used to feature-extract the above-mentioned user identification information and video identification information to obtain user identification information features and video identification information features represented in an encoding form, and the sparse feature extraction model can be used to feature-extract the gender information of the user account.

[0105] In the embodiments of the present application, since the barrage information of the video barrage can include not only the barrage content but also the barrage attributes of the video barrage, such as the font of the video barrage, the color of the video barrage, the number of likes of the video barrage and the number of dislikes of the video barrage, etc., the sparse feature extraction can be performed on the barrage attributes with less features (such as the font of the video barrage), the semantic feature extraction can be performed on the barrage attributes with more features (such as the number of likes of the video barrage), and the semantic feature extraction can be performed on the barrage content of the video barrage. That is, the barrage features of the video barrage can include the semantic features of the barrage content, the sparse features of the barrage attributes and the dense features of the barrage attributes.

[0106] The user features, the video features and the barrage features are concatenated to obtain second concatenation features corresponding to each video barrage.

[0107] The user features, the video features and the barrage features can be spliced first to obtain a second spliced feature corresponding to each video barrage. For example, the features are usually described in the form of a vector, and the vector usually appears in the form of an array. The concat function can be used for splicing operation.

[0108] In the embodiment of the present application, the third sub-network model can be used to obtain the second spliced feature, as shown in the formula (3), the third sub-network model can include a second feature splicing layer and a second full connection layer. The second feature splicing layer is used to receive the input video information, user information and barrage information, and output the second spliced feature. The second normalization layer is used to normalize the second spliced feature to obtain the video barrage feature. The third sub-network model can receive the input video information, user information and barrage information, and perform feature extraction, feature fusion, feature normalization and other operations on the video information, user information and barrage information, and output the video barrage feature. Figure 2

[0109] Specifically, the second feature splicing layer can be composed of a second feature extraction layer and a second splicing layer. The second feature extraction layer can further include a plurality of feature extraction modules and / or feature encoding modules. For example, as shown in the formula (4), the second feature extraction layer can include a video identification information encoding module, a user identification information encoding module and a barrage information feature extraction module. The video identification information encoding module is used to perform feature encoding on the video identification information in the video information. The user identification information encoding module is used to perform feature encoding on the user identification information in the user information. The barrage information feature extraction module is used to extract the barrage features of the barrage information. The above encoding modules can perform embedding encoding. The above barrage features can include semantic features of the barrage content of the barrage, and dense features and sparse features of the barrage attributes of the barrage. Figure 2 The second splicing layer can be constructed based on the concat function, and is used to combine the features output by each feature extraction module and / or feature encoding module in the second feature extraction layer to obtain the second spliced feature.

[0110]

[0111] ​​Furthermore, in this embodiment, the user information input to the first feature extraction layer and the second feature extraction layer can be different. For example, the user information input to the first feature extraction layer can include user identification information and user account information, while the user information input to the second feature extraction layer can only include user identification information. In addition, since some of the information input to the first and second feature extraction layers is the same, for the same information, the first and second feature extraction layers can share the same feature extraction module and / or feature encoding module to avoid redundant processing of the same information. For example, if both the first and second feature extraction layers need to extract features from video identification information and user identification information, then the first and second feature extraction layers can share the same video identification information encoding module and the same user identification information encoding module.

[0112] Sub-step 2032: Input the second splicing feature into the second normalization layer to obtain the video bullet screen feature output by the second normalization layer.

[0113] In this embodiment of the application, a second normalization layer may also be included. The second normalization layer is used to normalize the second splicing features to obtain the video bullet screen features corresponding to each video bullet screen.

[0114] It should be noted that, in order to improve the accuracy of video bullet screen features, a second fully connected layer can be set at the input of the second normalization layer. The second splicing features can be input into the second fully connected layer first, and then the output of the second fully connected layer can be used as the input of the second normalization layer to finally obtain the video bullet screen features output by the second normalization layer.

[0115] In this embodiment, the second normalization layer can normalize the second splicing features to obtain the final video bullet screen features. Specifically, the second normalization layer may include an nn.LayerNorm function layer, which can perform normalization operations using the nn.LayerNorm function.

[0116] Furthermore, the second normalization layer may also include an L2 regularization sub-layer. The L2 regularization sub-layer can be set at the end of the second normalization layer. The L2 regularization sub-layer can add an additional penalty term to the model to prevent overfitting.

[0117] like Figure 2 As shown, a second fully connected layer can also be set between the second feature splicing layer and the second normalization layer. The second splicing feature can be input into the second fully connected layer first, and then the output of the second fully connected layer can be used as the input of the second normalization layer to obtain the video bullet screen feature output by the second normalization layer.

[0118] Step 204, input the target crowd category feature and the video barrage feature into a feature matching sub-model of the user barrage matching model, to obtain a barrage matching degree output by the feature matching sub-model.

[0119] It should be noted that the barrage matching degree can also be determined by the initial neural network model, for example, as shown in the following formula: Figure 2 As shown in the formula, the initial neural network model can also include a feature matching sub-model, which is used to determine the barrage matching degree of the target crowd category feature and the video barrage feature.

[0120] Step 205, screening target barrages from the video barrage set based on the barrage matching degree.

[0121] Optionally, step 205 can include:

[0122] Sub-step 2051, arranging each video barrage in the video barrage set in descending order of the corresponding barrage matching degree to obtain a video barrage sequence.

[0123] After obtaining the barrage matching degree corresponding to each video barrage in the video barrage set, each video barrage can be arranged according to the size of the barrage matching degree to obtain a video barrage sequence. In the video barrage sequence, the video barrage with a larger barrage matching degree is ranked closer to the sequence header.

[0124] Sub-step 2052, determining the video barrages located at the sequence header of the video barrage sequence as target barrages in a fixed screening number or a fixed screening proportion.

[0125] In an embodiment, a fixed screening number of video barrages close to the sequence header in the video barrage sequence can be determined as target barrages. In another embodiment, a fixed screening proportion of video barrages close to the sequence header in the video barrage sequence can be determined as target barrages.

[0126] It should be noted that in the embodiments of the present application, the way of screening target barrages by using barrage matching degrees is not specifically limited, and the skilled in the art can determine a suitable screening method according to actual needs, for example, the video barrage in the video barrage set with a barrage matching degree greater than or equal to a preset matching degree can also be directly determined as a target barrage.

[0127] Step 206, outputting the target barrage.

[0128] This step can be referred to step 105, and the embodiments of the present application will not be repeated.

[0129] In summary, in another method for processing bullet screen provided by the embodiments of the present application, when a user plays a target video, a target crowd category feature can be determined from a crowd category feature set corresponding to the target video based on user information of the user and video information of the target video; feature extraction is performed on bullet screen information corresponding to each video bullet screen in a video bullet screen set of the target video to obtain video bullet screen features corresponding to the video bullet screen; a bullet screen matching degree is determined according to the target crowd category feature and the video bullet screen feature; a target bullet screen is filtered from the video bullet screen set based on the bullet screen matching degree; and the target bullet screen is output. The user bullet screen matching model obtained by training the initial neural network can determine the bullet screen matching degree of each bullet screen of the target video for the user who is watching the target video, so that the efficiency and accuracy of filtering the target bullet screen based on the bullet screen matching degree can be improved.

[0130] Figure 2 is a step flowchart of a model training method provided by the embodiments of the present application. As shown in Figure 2 , the method can include:

[0131] Step 301: obtaining sample video information of a sample video, sample bullet screen information of the sample video, sample user information of a sample user, and sample preference information of the sample user for the sample bullet screen.

[0132] Specifically, a training sample can be constructed, which can include sample video information of a sample video, sample bullet screen information of the sample video, sample user information of a sample user, and sample preference information of the sample user for the sample bullet screen.

[0133] The sample preference information can be determined according to historical operations such as historical likes and historical dislikes of the sample user for the sample bullet screen, for example, the sample preference information can include positive samples and negative samples, wherein if the sample preference information of the sample bullet screen is a positive sample, it indicates that the user is interested in the bullet screen, and if the sample preference information of the sample bullet screen is a negative sample, it indicates that the user is not interested in the bullet screen.

[0134] The sample preference information of the sample bullet screen can also be determined according to other bullet screen processing rules, for example, the sample preference information can be determined according to a bullet screen shielding rule set by the user, the sample bullet screen shielded by the bullet screen shielding rule is determined as a negative sample, and the sample bullet screen not shielded by the bullet screen shielding rule is determined as a positive sample.

[0135] Step 302: inputting the sample video information, the sample bullet screen information, and the sample user information into an initial neural network model to obtain a sample bullet screen matching degree between the sample user and the sample bullet screen output by a feature matching sub-model of the initial neural network model.

[0136] In the embodiment of the present application, an initial neural network model can be constructed, and the initial neural network model is trained by using the training samples to obtain a user barrage matching model. The user barrage matching model can output a feature matching degree between a user and a barrage according to input video information, user information and barrage information, so as to realize the purpose of screening a target barrage through the feature matching degree.

[0137] Specifically, the sample video information, the sample barrage information and the sample user information can be input into the initial neural network model to obtain a sample barrage matching degree between a sample user and a sample barrage output by a feature matching sub-model of the initial neural network model. A loss value between the sample barrage matching degree and sample preference information is calculated by using a loss function, and parameters in the initial neural network model are adjusted according to the loss value until the loss value converges, and the user barrage matching model is obtained.

[0138] In step 303, the initial neural network model is trained by using the sample barrage matching degree and the sample preference information to obtain the user barrage matching model.

[0139] The user barrage matching model is used to output a barrage matching degree according to input video information, barrage information and user information.

[0140] In the embodiment of the present application, after obtaining the barrage matching degree of a sample user for a sample barrage in a sample video, a loss value between sample preference information of the sample user for the sample barrage and the barrage matching degree can be determined, and various parameters in the initial neural network model are adjusted based on the loss value until the loss value converges, and the user barrage matching model is obtained. The parameters to be adjusted can include element values of each population category feature element in the three-dimensional feature matrix, so that the target population category feature determined by the three-dimensional feature matrix can be more accurate.

[0141] In sub-step 3031, a model loss is determined according to the sample barrage matching degree and the sample preference information.

[0142] In the embodiment of the present application, the model loss can be determined by one or more of a log-likelihood loss function, an ordinary least squares function, an Adaboost function, a mean squared error function (MSE), a mean absolute error function (MAE) and a cross entropy loss function. The skilled person can select the required loss function according to actual needs, and the present application does not make specific limitations in this regard.

[0143] Further, since the element values of the crowd class feature elements in the three-dimensional feature matrix need to be adjusted according to the model loss in the present application, in order to provide better adjustment effect, the focal loss function (Focal Loss) can be preferably used to calculate the model loss. The focal loss function can be shown in the following formula 1:

[0144] Formula 1

[0145] wherein, represents the model loss, , and are adjustable factors of the loss function, which can adjust the output features of the function, such as sensitivity, amplitude, etc. and is based on the value of the sample preference information, and the value of a is 1 when the sample preference information indicates that the sample barrage is a positive sample, and the value of a is 0 when the sample preference information indicates that the sample barrage is a negative sample. represents the barrage matching degree.

[0146] Sub-step 3032, adjusting the crowd class feature elements in the three-dimensional feature matrix based on the model loss to obtain the user barrage matching model.

[0147] In the training process, since the barrage matching degree output by the initial neural network model is determined based on the target two-dimensional feature matrix corresponding to the sample video, the crowd class feature elements in the two-dimensional feature matrix corresponding to the sample video in the three-dimensional feature matrix are adjusted according to the model loss, so as to determine a more accurate barrage matching degree through the target two-dimensional feature matrix. In the present application, the process of adjusting the crowd class feature elements through the model loss can be realized based on the gradient back propagation algorithm (back propagation, BP).

[0148] Referring to Figure 5 , Figure 5 is an architecture diagram of a barrage processing system provided by the present application. As shown in Figure 5 , the barrage processing system can be composed of two parts, an offline subsystem and an online subsystem.

[0149] ​The offline subsystem can include a sample collection unit, a feature engineering unit, and an offline training unit, wherein the sample collection unit is configured to obtain sample videos; the feature engineering unit is configured to generate video barrage features corresponding to barrages of the sample videos, and store the video barrage features; and the offline training unit is configured to train an initial neural network model, and generate a user barrage matching model, and can perform optimization training on the user barrage matching model based on newly obtained sample videos. The offline subsystem can write the video barrage features corresponding to the sample videos into a barrage feature storage unit, write a three-dimensional feature matrix generated during model training into a feature matrix database, and can deploy the trained user barrage matching model in the online subsystem.

[0150] The online subsystem can include an inference unit, a barrage processing unit, and a feature storage unit, wherein the inference unit can store and run the user barrage matching model trained by the offline training unit; the feature storage unit stores video barrage features extracted by the feature engineering unit in advance, and a three-dimensional feature matrix generated during model training by the offline training unit; and the barrage processing unit provides barrage processing services for users.

[0151] When a user watches a target video, the barrage processing unit can match the target video with sample videos used when training the user barrage matching model, and can provide video information and user information to the inference unit when it is detected that the target video belongs to the sample videos, call only part of the functions in the user barrage matching model, obtain video user features of the user for the target video, determine target crowd category features based on the video user features and the three-dimensional feature matrix stored in the feature storage unit, obtain video barrage features of the target video from the feature storage unit, calculate inner products between the video barrage features and the target crowd category features, arrange the inner products from high to low, select video barrages corresponding to the highest inner products of a preset number or a preset proportion as target barrages, and output the target barrages.

[0152] Figure 6 is a block diagram of a barrage processing device provided by an embodiment of the present application. The device can include:

[0153] The category feature module 501 is configured to determine target crowd category features from a crowd category feature set corresponding to a target video based on user information of a user and video information of the target video when the user plays the target video.

[0154] The barrage feature module 502 is configured to perform feature extraction on barrage information of each video barrage in a video barrage set of the target video, and obtain video barrage features corresponding to the video barrages.

[0155] The barrage matching degree module 503 is configured to determine a barrage matching degree according to the target crowd category feature and the video barrage feature.

[0156] The target barrage module 504 is configured to filter target barrage from the video barrage set based on the barrage matching degree.

[0157] The barrage output module 505 is configured to output the target barrage.

[0158] Optionally, the category feature module comprises:

[0159] The video user feature submodule is configured to input the video information and the user information into a first subnetwork model of a user barrage matching model to obtain the video user feature output by the first subnetwork model.

[0160] The crowd category feature submodule is configured to input the video user feature and the video information into a second subnetwork model of the user barrage matching model to obtain the target crowd category feature output by the second subnetwork model.

[0161] Optionally, the first subnetwork model is configured to extract a user feature based on the user information, extract a video feature based on the video information, and fuse the user feature and the video feature to obtain the video user feature.

[0162] Optionally, the first subnetwork model comprises a first feature splicing layer and a first normalization layer.

[0163] The video information and the video user feature submodule comprises:

[0164] The first spliced feature submodule is configured to input the video information and the user information into the first feature splicing layer to obtain a first spliced feature output by the first feature splicing layer; wherein the first feature splicing layer is configured to determine the video feature according to the video information, determine the user feature according to the user information, and splice the video feature and the user feature into the first spliced feature.

[0165] The first normalization submodule is configured to input the first spliced feature into the first normalization layer to obtain the video user feature output by the first normalization layer.

[0166] Optionally, the second subnetwork model is configured to determine a crowd matching degree between each crowd category feature in a crowd category feature set of the target video and the video user feature, and determine a crowd category feature with a maximum crowd matching degree in the crowd category feature set as the target crowd category feature.

[0167] Optionally, the user barrage matching model further comprises a three-dimensional feature matrix and a one-hot encoding module, and the device further comprises:

[0168] a one-hot encoding module, configured to input the video information into the one-hot encoding module to obtain a target one-hot encoding output by the one-hot encoding module;

[0169] a two-dimensional feature matrix module, configured to extract a target two-dimensional feature matrix corresponding to the target video from the three-dimensional feature matrix based on positions of valid bits in the target one-hot encoding; wherein the three-dimensional feature matrix is composed of two-dimensional feature matrices corresponding to respective encoding bits in the target one-hot encoding, and each two-dimensional feature matrix is composed of a first preset number of columns of crowd feature elements, and each column of crowd feature elements is composed of a second preset number of crowd category features.

[0170] a crowd category feature set module, configured to combine crowd category features in the same column in the target two-dimensional feature matrix into a crowd category feature to obtain a crowd category feature set of the target video.

[0171] Optionally, the barrage feature module is further configured to input the video information, the user information, and the barrage information into a third sub-network model of the user barrage matching model to obtain the video barrage feature output by the third sub-network model; wherein the third sub-network model is configured to extract a video user feature based on the user information, extract a video feature based on the video information, extract a barrage feature of each video barrage in a video barrage set of the target video based on barrage information corresponding to the video barrage, and fuse the user feature, the video feature, and the barrage feature to obtain a video barrage feature corresponding to the video barrage.

[0172] Optionally, the third sub-network model comprises a second feature splicing layer and a second normalization layer, and the barrage feature module comprises:

[0173] a second spliced feature submodule, configured to input the video information, the user information, and the barrage information into the second feature splicing layer to obtain a second spliced feature output by the second feature splicing layer; wherein the second feature splicing layer is configured to determine a video feature based on the video information, determine a user feature based on the user information, determine a barrage feature based on the barrage information, and splice the video feature, the user feature, and the barrage feature;

[0174] a second normalization submodule, configured to input the second spliced feature into the second normalization layer to obtain a video barrage feature output by the second normalization layer.

[0175] Optionally, the barrage matching degree module is further configured to input the target crowd category feature and the video barrage feature into a feature matching submodel of the user barrage matching model to obtain a barrage matching degree output by the feature matching submodel.

[0176] Optionally, the target barrage module comprises:

[0177] a video barrage sequence submodule configured to arrange each video barrage in the video barrage set in a descending order of the corresponding barrage matching degree to obtain a video barrage sequence;

[0178] a target barrage submodule configured to determine video barrages located at a head of the video barrage sequence as target barrages, the video barrages being of a fixed screening number or a fixed screening proportion.

[0179] Optionally, the method comprises:

[0180] an information acquisition module configured to acquire sample video information of a sample video, sample barrage information of the sample video, sample user information of a sample user, and sample preference information of the sample user for sample barrages;

[0181] a sample barrage matching degree module configured to input the sample video information, the sample barrage information, and the sample user information into an initial neural network model to obtain a sample barrage matching degree between a sample user and a sample barrage output by a feature matching submodel of the initial neural network model;

[0182] a training module configured to train the initial neural network model by using the sample barrage matching degree and the sample preference information to obtain a user barrage matching model, wherein the user barrage matching model is configured to output the barrage matching degree according to input of the video information, the barrage information, and the user information.

[0183] Optionally, the training module comprises:

[0184] a model loss submodule configured to determine a model loss according to the sample barrage matching degree and the sample preference information;

[0185] a training submodule configured to adjust crowd category feature elements in a three-dimensional feature matrix based on the model loss to obtain the user barrage matching model.

[0186] In summary, the barrage processing device provided by the embodiment of the present application comprises: a category feature module, configured to determine a target crowd category feature from a crowd category feature set corresponding to a target video based on user information of a user and video information of the target video when the user plays the target video; a barrage feature module, configured to perform feature extraction on barrage information corresponding to each video barrage in a video barrage set of the target video to obtain video barrage features corresponding to the video barrage; a barrage matching degree module, configured to determine a barrage matching degree according to the target crowd category feature and the video barrage features; a target barrage module, configured to filter a target barrage from the video barrage set based on the barrage matching degree; and a barrage output module, configured to output the target barrage. The crowd category feature to which the user is directed to the target video can be determined, and the video barrage is filtered according to the crowd category feature, so that the same user can correspond to different crowd category features for different target videos. The target barrage is matched from the barrage of the target video and displayed to the user through the crowd category feature with pertinence, the accuracy in filtering the video barrage is effectively improved, and the target barrage with a higher matching degree to the user can be effectively filtered.

[0187] For the above device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts refer to the part of the method embodiment.

[0188] Preferably, the embodiment of the present application further provides a terminal, comprising a processor, a memory, a computer program stored in the memory and executable on the processor, which implements each process of the above barrage processing method and model training method embodiment when executed by the processor, and achieves the same technical effect. To avoid repetition, it will not be repeated here.

[0189] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores computer instructions, which are executed by the processor to implement each process of the above barrage processing method and model training method embodiment, and achieve the same technical effect. To avoid repetition, it will not be repeated here. The computer readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0190] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same and similar parts of each embodiment can be referred to.

[0191] It is readily apparent to those skilled in the art that any combination of the various embodiments described above is possible, and thus any combination of the various embodiments described above is an embodiment of the present application, but due to the length of the specification, not all combinations are described herein.

[0192] The barrage processing method and model training method provided herein are not inherently related to any particular computer, virtual system, or other device. Various general purpose systems can also be used with the teachings herein. The structure for a system having the solution of the present application as described above is readily apparent from the above description. Moreover, the present application is not intended to be limited to any particular programming language. It will be appreciated that a variety of programming languages can be used to implement the teachings of the present application described herein, and any references below to specific languages are provided for disclosure of enablement of the best mode of the present application.

[0193] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been described in detail in order not to obscure the understanding of this description.

[0194] Similarly, it is to be understood that the mechanical features of the application sometimes are grouped together in a single embodiment, figure or description of related embodiments in the above description of illustrative embodiments of the application for clarity only. However, one or more features of an embodiment could be typically equally combined with some other embodiments of the application. Notably, as the above description of illustrative embodiments of the application sometimes describes features in terms of processes, procedures, or steps, it is to be understood that none of the disclosed aspects are limited to these specific processes, procedures, or steps, which are illustrative only. Rather, each of the disclosed aspects can be applied to any other illustrative embodiment where such processes, procedures, or steps provide beneficial results. Thus, the foregoing description of illustrative embodiments of the application applies to any embodiment or combination of embodiments of the application or parts thereof and the overall scope of the application includes all alternatives, modifications and equivalents thereof.

[0195] Those skilled in the art will appreciate that the modules in the apparatuses in the embodiments can be adapted and placed in one or more apparatuses other than the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and further can be divided into more sub-modules or sub-units or sub-components. Any combination of all the features disclosed in the specification (including the accompanying claims, abstract and drawings), and any method or apparatus so disclosed, can be used in any combination, except that at least some of such features and / or processes or units are mutually exclusive, unless otherwise explicitly stated. Each feature disclosed in the specification (including the accompanying claims, abstract and drawings), can be replaced by alternative features serving the same, equivalent or similar purpose, unless otherwise expressly stated.

[0196] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.

[0197] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the model training method according to the embodiments of this application. This application can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such an implementation of this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0198] It should be noted that the above embodiments are illustrative of this application and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

Claims

1. A method for processing a barrage, characterized in that, The method comprises: when a user plays a target video, determining a target crowd category feature from a crowd category feature set corresponding to the target video based on user information of the user and video information of the target video; extracting features of each video comment information in a video comment set of the target video to obtain video comment features of the video comment; determining a comment matching degree according to the target crowd category feature and the video comment features; screening a target comment from the video comment set based on the comment matching degree; outputting the target comment; wherein the determining of the target crowd category feature from the crowd category feature set corresponding to the target video based on the user information of the user and the video information of the target video comprises: determining video user features according to the user information and the video information; matching the video user features with each crowd category feature in the crowd category feature set; determining the target crowd category feature from the crowd category feature set according to the matching result; wherein the crowd category feature is obtained by classifying audiences of the target video and extracting features of personal information and / or behavior information of people in different categories of crowds.

2. The method of claim 1, wherein, The determining of the target crowd category feature from the crowd category feature set corresponding to the target video based on the user information of the user and the video information of the target video comprises: inputting the video information and the user information into a first sub-network model of a user comment matching model to obtain the video user features output by the first sub-network model; inputting the video user features and the video information into a second sub-network model of the user comment matching model to obtain the target crowd category features output by the second sub-network model.

3. The method of claim 2, wherein, The first sub-network model is used to extract user features based on the user information, extract video features based on the video information, and fuse the user features and the video features to obtain video user features.

4. The method according to claim 2 or 3, characterized in that, The first sub-network model comprises a first feature splicing layer and a first normalization layer. The inputting of the video information and the user information into the first sub-network model of the user comment matching model to obtain the video user features output by the first sub-network model comprises: inputting the video information and the user information into the first feature splicing layer to obtain first splicing features output by the first feature splicing layer; wherein the first feature splicing layer is used to determine video features according to the video information, determine user features according to the user information, and splice the video features and the user features into the first splicing features; inputting the first splicing features into the first normalization layer to obtain the video user features output by the first normalization layer.

5. The method of claim 2, wherein, The second sub-network model is used to determine a crowd matching degree between each crowd category feature in a crowd category feature set of the target video and the video user features, and determine a target crowd category feature as a crowd category feature in the crowd category feature set with the maximum crowd matching degree.

6. The method of claim 2, wherein, The user barrage matching model further comprises a three-dimensional feature matrix and a one-hot encoding module, and before the target crowd category feature is determined from the crowd category feature set corresponding to the target video, the method further comprises: inputting the video information into the one-hot encoding module to obtain a target one-hot encoding output by the one-hot encoding module; extracting a target two-dimensional feature matrix corresponding to the target video from the three-dimensional feature matrix based on the position of the effective bit in the target one-hot encoding; wherein the three-dimensional feature matrix is composed of two-dimensional feature matrices corresponding to each encoding bit in the target one-hot encoding, the two-dimensional feature matrix is composed of a first preset number of crowd feature element columns, and the crowd feature element column is composed of a second preset number of crowd category feature elements; combining crowd category feature elements in the same column in the target two-dimensional feature matrix into a crowd category feature to obtain a crowd category feature set of the target video.

7. The method of claim 1, wherein, The feature extraction of the barrage information corresponding to each video barrage in the video barrage set of the target video to obtain the video barrage feature corresponding to the video barrage comprises: inputting the video information, the user information and the barrage information into a third sub-network model of the user barrage matching model to obtain the video barrage feature output by the third sub-network model; wherein the third sub-network model is used to extract video user features based on the user information, extract video features based on the video information, extract features of the barrage information corresponding to each video barrage in the video barrage set of the target video to obtain the barrage features of the video barrage, and fuse the user features, the video features and the barrage features to obtain the video barrage feature corresponding to the video barrage.

8. The method of claim 7, wherein, The third sub-network model comprises a second feature splicing layer and a second normalization layer, and the inputting of the video information, the user information and the barrage information into the third sub-network model of the user barrage matching model to obtain the video barrage feature output by the third sub-network model comprises: inputting the video information, the user information and the barrage information into the second feature splicing layer to obtain a second splicing feature output by the second feature splicing layer; wherein the second feature splicing layer is used to determine video features according to the video information, determine user features according to the user information, determine barrage features according to the barrage information, and splice the video features, the user features and the barrage features; inputting the second splicing feature into the second normalization layer to obtain a video barrage feature output by the second normalization layer.

9. The method of claim 1, wherein, The determination of the barrage matching degree according to the target crowd category feature and the video barrage feature comprises: inputting the target crowd category feature and the video barrage feature into a feature matching sub-model of the user barrage matching model to obtain a barrage matching degree output by the feature matching sub-model.

10. The method of claim 1, wherein, The screening of the target barrage from the video barrage set based on the barrage matching degree comprises: arranging each video barrage in the video barrage set in descending order of the corresponding barrage matching degree to obtain a video barrage sequence; The video barrage located at the head of the video barrage sequence is determined as the target barrage.

11. The method of claim 2, wherein, The method comprises: obtaining sample video information of a sample video, sample barrage information of the sample video, sample user information of a sample user, and sample preference information of the sample user on sample barrage; inputting the sample video information, sample barrage information and sample user information into an initial neural network model to obtain a sample barrage matching degree between the sample user and the sample barrage output by a feature matching sub-model of the initial neural network model; training the initial neural network model using the sample barrage matching degree and the sample preference information to obtain a user barrage matching model; wherein the user barrage matching model is used to output the barrage matching degree according to inputted video information, barrage information and user information.

12. The method of claim 11, wherein, The training of the initial neural network model using the sample barrage matching degree and the sample preference information to obtain a user barrage matching model comprises: determining a model loss according to the sample barrage matching degree and the sample preference information; adjusting a crowd category feature element in a three-dimensional feature matrix based on the model loss to obtain the user barrage matching model.

13. A barrage processing apparatus characterized by comprising: The device comprises: a category feature module configured to determine a target crowd category feature from a crowd category feature set corresponding to a target video based on user information of a user and video information of the target video when the user plays the target video; a barrage feature module configured to extract features from barrage information corresponding to each video barrage in a video barrage set of the target video to obtain video barrage features corresponding to the video barrage; a barrage matching degree module configured to determine a barrage matching degree according to the target crowd category feature and the video barrage features; a target barrage module configured to screen a target barrage from the video barrage set based on the barrage matching degree; a barrage output module configured to output the target barrage; wherein the determination of the target crowd category feature from the crowd category feature set corresponding to the target video based on the user information of the user and the video information of the target video comprises: determining a video user feature according to the user information and the video information; matching the video user feature with each crowd category feature in the crowd category feature set; determining the target crowd category feature from the crowd category feature set according to a matching result; wherein the crowd category feature is obtained by classifying audiences of the target video and extracting features from personal information and / or behavior information of different categories of crowds.

14. An electronic device, comprising: comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1-12.

15. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to execute the method of any one of claims 1-12. The computer instructions are used to enable the computer to execute the method of any one of claims 1-12.

16. A computer program product, characterised in that, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Bullet screen pushing method and bullet screen system

    CN115700475A