Video recommendation method and device, electronic equipment and storage medium

By determining the user's behavioral activity and matching the historical preference videos of high-activity users, the problem of poor recommendations for newly registered users in video applications is solved, and the accuracy and user experience of recommendations are improved.

CN120343305APending Publication Date: 2025-07-18BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510479143.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing video applications are difficult to accurately recommend content that new registered users or users with low behavioral activity, resulting in poor recommendation results.

Method used

Improve recommendation accuracy by determining user behavioral activity, obtaining metadata, and matching the historical preference videos of highly active users with similar high activity to recommend.

Benefits of technology

Improve the video recommendation effect for newly registered users or users with low behavioral activity and enhance the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343305A_ABST
    Figure CN120343305A_ABST
Patent Text Reader

Abstract

The invention provides a video recommendation method and device, electronic equipment and a storage medium. The method comprises the following steps: determining a behavior type of a first user in a video application; under the condition that the behavior type of the first user is a first behavior type, metadata of the first user is obtained, and the first behavior type is used for representing that the behavior activeness of the user in the video application is smaller than a first preset threshold value; matching N second users similar to the metadata of the first user, the behavior type of the second user being a second behavior type, the second behavior type being used for representing that the behavior activeness of the user in the video application is greater than a second preset threshold, the second preset threshold being greater than the first preset threshold, and N being a positive integer; and performing video recommendation on the first user based on P target videos historically preferred by the N second users in the video application, P being a positive integer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and particularly to a video recommendation method, apparatus, electronic device, and storage medium. Background Art

[0002] With the rapid development of video processing technology, video applications have been widely used and are deeply loved by users. Users can watch content they are interested in through video applications.

[0003] Currently, in order to improve the user experience, video applications usually recommend videos to users based on the users' own video playback behaviors in the video applications. For example, if user A often watches martial arts TV dramas in a video application, when user A starts the video application, high-rated martial arts TV dramas are usually displayed on the interface.

[0004] However, if the registration time of a user in a video application is relatively short, or the user's video playback behaviors in the video application are few, it is difficult for the video application to directly predict the content that the user is interested in, resulting in relatively poor video recommendation effects. Summary of the Invention

[0005] Embodiments of the present invention provide a video recommendation method, apparatus, electronic device, and storage medium to solve the technical problem of relatively poor video recommendation effects in related technologies. The specific technical solutions are as follows:

[0006] In the first aspect of the embodiments of the present invention, a video recommendation method is first provided. The method includes:

[0007] Determine the behavior type of a first user in a video application;

[0008] When the behavior type of the first user is a first behavior type, obtain the metadata of the first user. The first behavior type is used to indicate that the user's behavior activity in the video application is less than a first preset threshold;

[0009] Match N second users whose metadata is similar to that of the first user. The behavior type of the second users is a second behavior type. The second behavior type is used to indicate that the user's behavior activity in the video application is greater than a second preset threshold. The second preset threshold is greater than the first preset threshold, and N is a positive integer;

[0010] Based on P target videos that N second users have historically preferred in the video application, recommend videos to the first user. P is a positive integer.

[0011] In the second aspect of the embodiments of the present invention, a video recommendation apparatus is further provided. The apparatus includes:

[0012] A determination module, configured to determine the behavior type of a first user in a video application;

[0013] An acquisition module, configured to acquire metadata of the first user when the behavior type of the first user is a first behavior type, where the first behavior type is used to indicate that the behavior activity of the user in the video application is less than a first preset threshold;

[0014] A matching module, configured to match N second users similar to the metadata of the first user, where the behavior type of the second users is a second behavior type, the second behavior type is used to indicate that the behavior activity of the users in the video application is greater than a second preset threshold, the second preset threshold is greater than the first preset threshold, and N is a positive integer;

[0015] A recommendation module, configured to perform video recommendation for the first user based on P target videos that the N second users have historically preferred in the video application, where P is a positive integer.

[0016] In a third aspect of the embodiments of the present invention, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used to store a computer program; the processor is configured to implement the video recommendation method described in any one of the above when executing the program stored on the memory.

[0017] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium, and when the instructions run on a computer, the computer is enabled to execute the video recommendation method described in any one of the above.

[0018] In a fifth aspect of the embodiments of the present invention, a computer program product including instructions is further provided. When the computer program product runs on a computer, the computer is enabled to execute the video recommendation method described in any one of the above.

[0019] In the embodiments of the present invention, by determining the behavior activity of a first user in a video application; acquiring the metadata of the first user when the behavior activity of the first user in the video application is less than a first preset threshold; then matching N second users similar to the metadata of the first user, where the behavior activity of the second users in the video application is greater than a second preset threshold; and performing video recommendation for the first user based on P target videos that the N second users have historically preferred in the video application. Since the behavior activity of the second users in the video application is relatively high, the video content they prefer is easy to capture, and second users similar to the first user are matched through the metadata of the users, and the video content preferred by the second users is pushed to the first user, so that video content that the first user may potentially prefer can be recommended to the first user, thereby improving the effect of video recommendation. Brief Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0021] Figure 1 It is a flowchart of the video recommendation method in the embodiments of the present invention;

[0022] Figure 2 It is a schematic flowchart of the video recommendation method of a specific example

[0023] Figure 3 It is a schematic structural diagram of a video recommendation device in the embodiments of the present invention;

[0024] Figure 4 It is a schematic structural diagram of an electronic device in the embodiments of the present invention. Detailed Embodiments

[0025] The following will describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention.

[0026] Please refer to Figure 1 , Figure 1 It is a flowchart of the video recommendation method in the embodiments of the present invention. The video recommendation method includes:

[0027] S101, determining the behavior type of the first user in the video application.

[0028] The video recommendation method provided by the embodiments of the present invention is applied to a video application. The video application can be deployed at both ends, namely the client and the server. At the client, the user can interact with the video application to trigger the playback of videos in the video application. At the server, video recommendations can be made to the user through video recommendation strategies, and the video recommendation method in this embodiment is applied to the server.

[0029] The first user can be a user registered in the video application, which can be a newly registered user or a user previously registered in the video application.

[0030] The behavior type of the user in the video application is divided according to the behavior activity of the first user in the video application. According to the behavior activity of the user in the video application, the behavior type of the user in the video application can be divided into at least two categories.

[0031] In some embodiments, according to the behavior activity of the user in the video application, the behavior type of the user in the video application can be divided into three categories, namely the first behavior type, the second behavior type, and the third behavior type.

[0032] Among them, the first behavior type is used to indicate that the user's behavior activity in the video application is less than the first preset threshold. The first preset threshold is usually relatively small, indicating that the user's behavior in the video application is inactive. This user group can be classified as shallow-behavior users. Among them, the user's behavior in the video application can refer to the user's video playback behavior in the video application.

[0033] The second behavior type is used to characterize that the user's behavior activity in the video application is greater than the second preset threshold. The second preset threshold is usually relatively large, indicating that the user's behavior in the video application is relatively active. This user group can be classified as deep-behavior users.

[0034] The user's behavior activity in the video application indicated by the third behavior type is between the first preset threshold and the second preset threshold. This user group can be classified as middle-behavior users.

[0035] In some embodiments, the behavior activity can be normalized to between 0 and 1. The closer it is to 0, the less active the user's behavior in the video application is. For example, the first preset threshold can be close to 0. The closer it is to 1, the more active the user's behavior in the video application is. For example, the second preset threshold can be close to 1.

[0036] In the related art, the behavior of shallow-behavior users in the video application is relatively less, and it is usually difficult to accurately capture their specific interests and needs. The behavior of deep-behavior users in the video application is more, and the video recommendation algorithm can relatively easily capture the content that the user is interested in. The interests of middle-behavior users are more than those of shallow-behavior users and weaker than those of deep-behavior users. The video recommendation algorithm also has a certain ability to push the video content that they prefer. Among the above three user groups, the behavior of shallow-behavior users in the video application is less, so the effect of video recommendation for them is relatively poor. To solve this problem, in this step, the behavior type of the first user in the video application can be determined according to the behavior activity of the first user in the video application. And when it is determined that the behavior type of the first user is the first behavior type, it is determined to adopt a similarity-driven personalized recommendation strategy to find deep-behavior users similar to the shallow-behavior users, and based on the video content preferred by the deep-behavior users, perform video recommendation for the shallow-behavior users.

[0037] In some embodiments, the behavior type of the first user in the video application can be determined according to the number of interactions between the first user and the video application. Among them, each operation such as clicking, swiping, gesturing, and voice in the interface of the video application can be called an interaction of the user in the video application.

[0038] In some embodiments, the behavior type of the first user in the video application can be determined based on at least one of the number of videos triggered by the first user in the video application and the time of the most recent video trigger by the first user in the video application. Herein, a user triggering a video in the video application refers to the behavior of the user clicking on a video in the video application for playback.

[0039] In some embodiments, the behavior type of the first user in the video application can be determined comprehensively based on the number of videos triggered by the first user in the video application and the time of the most recent video trigger by the first user in the video application. This can improve the accuracy of determining the behavior type, thereby enabling a more accurate determination of the user's behavior activity level in the video application.

[0040] For example, if the number of videos triggered by the first user in the video application is relatively large and the time of the most recent video trigger is within one week, the behavior type of the first user in the video application can be determined as the second behavior type. Also, for example, if the number of videos triggered by the first user in the video application is relatively small and the most recent trigger time is more than one month ago, the behavior type of the first user in the video application can be determined as the first behavior type.

[0041] Step S102, when the behavior type of the first user is the first behavior type, obtain the metadata of the first user, where the first behavior type is used to indicate that the user's behavior activity level in the video application is less than the first preset threshold.

[0042] In this step, when it is determined that the behavior type of the first user is the first behavior type, the metadata of the first user can be obtained to perform user similarity matching using the metadata of the first user, thereby improving the effect of video recommendation.

[0043] The metadata of the first user is the descriptive information of the first user, which is used to describe the characteristics of the first user. For example, the metadata of the first user may include gender, age, age group, channel preference, channel preference time, permanent city, education level, mobile phone system, mobile phone brand, favorite tags, star preference, and watched TV series, etc.

[0044] For example, the metadata of the first user may be: Name A, male gender, age between 31 and 35 years old, age group is young men, channel preferences are: TV dramas, variety shows, movies, children's programs, life, etc., preference times are: 11 o'clock, 12 o'clock, 18 o'clock, 19 o'clock, permanent cities are Shanghai and Suzhou, education level is undergraduate, mobile phone system is Android system, mobile phone brand is Brand A, favorite video tags are: competition shows, comedies, talk shows, suspense, action, wuxia, fantasy, star preference is Star A, and watched TV series is TV Series A.

[0045] The metadata of the first user can be obtained from the registration information of the first user in the video application. For example, the basic information of the first user can be obtained from the registration information to get the metadata of the first user. Alternatively, the metadata of the first user can be obtained by mining the behavior of the first user in the video application. For example, by mining the videos triggered by the user, the TV series watched by the user, channel preferences, star preferences, and preference time can be obtained.

[0046] Step S103: Match N second users whose metadata is similar to that of the first user. The behavior type of the second user is the second behavior type, and the second behavior type is used to indicate that the behavior activity of the user in the video application is greater than a second preset threshold, and the second preset threshold is greater than the first preset threshold. N is a positive integer.

[0047] The second user can be a deep behavior user.

[0048] In some embodiments, the metadata of the second user stored in advance can be obtained, and the metadata of the first user is compared with the metadata of the second user to determine the metadata similar to the metadata of the first user, so as to match and obtain the second user similar to the first user.

[0049] In some embodiments, the vector representation of the metadata of the first user can be obtained, and the vector representation of the metadata of the second user stored in advance can be obtained. The similarity matching of the two vector representations is performed to determine the vector representation of the metadata of the second user whose distance from the vector representation of the metadata of the first user is close, so as to match and obtain the second user similar to the first user.

[0050] Step S104: Based on the P target videos that the N second users have historically preferred in the video application, perform video recommendation for the first user. P is a positive integer.

[0051] The P target videos that the second user has historically preferred in the video application can be the videos that the second user has watched many times in the video application, or the videos with a high viewing completion rate of the second user in the video application. For example, for video A, its viewing completion rate is 100%, so it is a video with a high viewing completion rate of the second user in the video application, that is, it is a target video.

[0052] Among them, the number of target videos that a second user has historically preferred in the video application can be at least one. In some embodiments, all P target videos can be recommended to the first user, or some target videos can be randomly selected from the P target videos and recommended to the first user. Alternatively, the recommendation weights of the target videos can be determined based on the preference degree of the second user for the target videos and the similarity degree between the second user and the first user. Then, based on the recommendation weights, some target videos are selected from the P target videos and video recommendations are made to the first user.

[0053] In this embodiment, the behavioral activity of the first user in the video application is determined; when the behavioral activity of the first user in the video application is less than the first preset threshold, the metadata of the first user is obtained; then, N second users whose metadata is similar to that of the first user are matched, and the behavioral activity of the second users in the video application is greater than the second preset threshold; and video recommendations are made to the first user based on P target videos that the N second users have historically preferred in the video application. Since the behavioral activity of the second users in the video application is relatively high, the video content they prefer is easy to capture. Moreover, by matching the second users similar to the first user through the user's metadata and pushing the video content preferred by the second users to the first user, potential preferred video content can be recommended to the first user, thereby improving the effect of video recommendations.

[0054] In some embodiments, the behavioral activity of the user in the video application is determined by the behavioral data of the user in the video application, and step 101 specifically includes:

[0055] Obtain the behavioral data of the first user in the video application, where the behavioral data includes: the number of videos triggered by the first user in the video application, and the time of the most recent video trigger.

[0056] When it is determined that the behavioral data meets the first preset condition, determine that the behavioral type of the first user is the first behavioral type; the first preset condition is: the number of videos triggered by the first user in the video application is less than or equal to the third preset threshold, and the time of the most recent video trigger is before the first preset time period; or, the first preset condition is: the number of videos triggered by the first user in the video application is zero; where, when it is determined that the behavioral data meets the first preset condition, determine that the behavioral activity of the first user in the video application is less than the first preset threshold.

[0057] In some embodiments, step 101 further includes:

[0058] When it is determined that the behavior data meets the second preset condition, determine that the behavior type of the first user is the second behavior type; the second preset condition is that the number of videos triggered by the first user in the video application is greater than the fourth preset threshold, and the time of the most recent video trigger is within the second preset time length, the fourth preset threshold is greater than the third preset threshold, and the first preset time length is greater than the second preset time length; wherein, when it is determined that the behavior data meets the first preset condition, determine that the behavior activity of the first user in the video application is less than the first preset threshold.

[0059] According to the user's behavior habits, two points are mainly considered, namely the number of videos the user has watched historically and the time interval between the user's most recent viewing and the current time. Based on these two points, users can be divided into three categories: shallow behavior users, middle behavior users, and deep behavior users. That is, the population can be classified according to the number of videos the user has watched and the time of the last video viewing. The main reason for this classification is that video recommendation mainly relies on the video content triggered by the user's history for collaborative recall. The more videos are triggered, the more interesting video content can be recalled, and the more accurately the user's interest can be characterized. On the contrary, the fewer videos are triggered, the less interesting video content can be recalled, and it is more difficult to accurately characterize the user's interest.

[0060] The third preset threshold and the fourth preset threshold can be set according to the actual situation. In some embodiments, the third preset threshold can be set to 3, and the fourth preset threshold can be set to 20. The first preset time length and the second preset time length can also be set according to the actual situation. In some embodiments, the first preset time length can be set to 30, and the second preset time length can be set to 7. Table 1 is a user statistics table for each trigger behavior in an example. In this example, deep behavior users can be those with the number of triggered videos > 20 and the behavior of the last video trigger within 7 days, that is, within one week. The user statistical ratios are 21.6% (the behavior of the last video trigger is within the same day) and 8.35% (the behavior of the last video trigger is within this week). Shallow behavior users are those with the number of triggered videos within 3 and the behavior of the last video trigger 30 days ago, and the user statistical ratio is 7.13%, or users with the number of triggered videos being 0, and the user statistical ratio is 19.08%. The remaining users can be classified as middle behavior users.

[0061] Table 1

[0062]

[0063] In this embodiment, by splitting the user behavior types into different behavior types according to the user's behavior, and based on the number of videos the user has watched historically and the time interval between the user's most recent viewing and the current time, the users are divided into three categories: shallow behavior users, medium behavior users, and deep behavior users. This can improve the accuracy of user classification, thereby improving the accuracy of user classification, and further improving the accuracy of video recommendation.

[0064] In some embodiments, step 103 specifically includes:

[0065] Obtain a first vector, where the first vector is a vector representation of the metadata of the first user;

[0066] Match N second vectors with a distance close to the first vector from the vectors of the user metadata stored in advance, to obtain N second users whose metadata is similar to that of the first user, where the second vector is a vector representation of the metadata of the second user.

[0067] In this embodiment, by obtaining the vector representation of the metadata of the first user and obtaining the vectors of the user metadata stored in advance, and performing similarity matching on the two vector representations to determine the vector representation of the metadata of the second user that is close to the vector representation of the metadata of the first user, so as to match and obtain the second user similar to the first user. This can improve the accuracy of user matching. Among them, the vector representations of the metadata of the users registered in the video application can be stored in advance. In this case, when the vector representation of the metadata of the first user is obtained, the second vector with a distance close to the first vector can be matched from the vectors of the user metadata stored in advance through vector matching.

[0068] In some embodiments, the first vector can be obtained through a deep learning model or a large model. The metadata of the first user can be input into a pre-trained deep learning model or large model, and the first vector is output.

[0069] In some embodiments, the metadata of the user includes multiple feature description information. The vector representation of each feature description information or the aggregated vector of the vector representations of multiple feature description information can be stored respectively. Then, the first vector is obtained by aggregating the vectors of the fine-grained feature description information. That is, the vector representations of the fine-grained feature description information of each user can be obtained first. Optionally, the metadata includes M feature description information, where M is a positive integer greater than 1. Before step 103, the method further includes:

[0070] Use a target model to perform vector representation on each of the feature description information, to obtain L third vectors of the M feature description information, where L is a positive integer less than or equal to M;

[0071] The obtaining of the first vector includes:

[0072] Aggregating the L third vectors to obtain the first vector.

[0073] In this embodiment, the target model may include a deep learning model and / or a large model. The feature description information of each user can be input into the target model to perform vector representation on the feature description information through the target model, and L third vectors of M feature description information are obtained.

[0074] In some embodiments, each feature description information of the first user can be vector-represented by an online and pre-trained target model to obtain L third vectors of M feature description information.

[0075] In some embodiments, an offline target model can be used to perform vector representation on the feature description information of each user, and the feature description information of each user and its vector representation are stored in an associated manner. Among them, the feature description information of each user may include the feature description information of the first user. In this way, through the M feature description information of the first user, L third vectors associated therewith can be obtained, and then the L third vectors are aggregated to obtain the first vector.

[0076] Among them, the third vector can be a vector representation of one feature description information or an aggregated vector of vector representations of multiple feature description information.

[0077] In this embodiment, by aggregating the vector representations of the fine-grained M feature description information, the vector representation of the user's metadata is obtained. In this way, the acquisition accuracy of the vector representation of the user's metadata can be improved, thereby improving the accuracy of user matching, and further improving the accuracy of video recommendation to the user.

[0078] In some embodiments, the M feature description information includes: K first feature description information and J second feature description information. The K first feature description information is related to the feature information of the first user, and the J second feature description information is related to Q video episodes triggered by the first user in the video application. The sum of K and J is equal to M, and Q is a positive integer;

[0079] The target model includes a first model and a second model. The step of using the target model to perform vector representation on each feature description information to obtain L third vectors of M feature description information includes:

[0080] Inputting the K first feature description information into the first model for unsupervised training, and obtaining K third vectors of the K first feature description information when the training is completed;

[0081] Input the J second feature description information into a second model for natural language processing to obtain Q third vectors corresponding to Q video series;

[0082] The obtaining of the first vector includes:

[0083] Aggregate the K third vectors of the K first feature description information to obtain a fourth vector; and aggregate the Q third vectors corresponding to the Q video series to obtain a fifth vector;

[0084] Concatenate the fourth vector and the fifth vector to obtain the first vector.

[0085] In this embodiment, the user's metadata can be vectorized in two ways, and different ways can be adopted according to different user feature description information. In some scenarios, the user's first feature description information is relatively concise and difficult to expand, and it is related to the user's feature information. For example, for relatively concise information such as name, gender, age, circle, favorite channel, preferred time, permanent city, education level, mobile phone system, mobile phone brand, favorite star label, etc., the first feature description information of each user can be input into a first model for unsupervised training to obtain the third vector of each first feature description information at a fine-grained level. The vector representation of the first user's metadata can be aggregated according to the vector representations of the first feature description information at each fine-grained level, and the vector representations of the first feature description information at each fine-grained level can be aggregated using an average fusion method to obtain a fourth vector. Among them, the first model can be a neural network model. For example, the first model can be a word vector representation model based on a neural network (word2vec model).

[0086] In some scenarios, the user's second feature description information is easy to expand, and it is related to the video series triggered by the user in the video application. At least one second feature description information of each video series can be input into a second model for natural language processing, and a third vector corresponding to the video series is output. The second model can be a large model, so that the information learned by the large model can be borrowed, and the generalization ability can be improved to a certain extent. In some embodiments, for a video series watched by a user, the video series can be described by information such as the name of the video, content summary, director, lead actor, release time, etc. The expanded second feature description information is input to the large model, and the large model can output an aggregated vector of the vector representations of the second feature description information, that is, output a third vector. Among them, the large model can be a natural language processing model based on the Transformer architecture. For example, the large model can be the bge-large-zh model.

[0087] In some embodiments, the video series triggered by the first user in the video application may be multiple. In this scenario, for one video series, the large model can output a third vector, and correspondingly, multiple third vectors can be output.

[0088] After that, aggregate the third vectors of each first feature description information to obtain a fourth vector; and aggregate the third vectors corresponding to each video series to obtain a fifth vector, and splice the fourth vector and the fifth vector to obtain a first vector. The metadata of each deep-behavior user in the video application can be vector-represented and stored in the above manner, and a vector index can be constructed for the deep-behavior users for convenient real-time retrieval to facilitate subsequent user matching.

[0089] Not only can the vector representations of the metadata of each user be stored, but in some embodiments, the vector representations of the feature description information at a fine-grained level can also be stored separately to facilitate obtaining the vector representations of the metadata of the shallow-behavior users.

[0090] In the case of obtaining multiple third vectors through the first model and the second model respectively, each third vector can be stored for online use. Its storage form is in the key-value (KV) format, such as: male gender: vector representation 1, female gender: vector representation 2, Android: vector representation 3, series 1: vector representation 4, series 2: vector representation 5, series 3: vector representation 6. Among them, for different vectors, their vector representation identifiers are different. For example, these vectors such as vector representation 1, vector representation 2, and vector representation 3 are all different.

[0091] In some embodiments, when the first user goes online, the feature description information of the first user can be obtained, and from the pre-stored vector representations, the vector representations associated with the feature description information of the first user can be obtained and aggregated to obtain a first vector.

[0092] In this embodiment, the user is characterized by using metadata. According to whether the feature description information in the metadata has the characteristic of extensibility, different methods are used for vector representation, and the vector representations of the two types of feature description information are spliced to vector-represent the metadata of the user. Combining the coarse and the fine can ensure the accuracy and generalization of the vector representation of the metadata of the user.

[0093] In some embodiments, step 104 specifically includes:

[0094] For each of the target videos, multiply the first weight corresponding to the second user who prefers the target video by the second weight corresponding to the target video to obtain the recommendation weight for recommending the target video to the first user. The first weight is used to represent the similarity between the second user and the first user, and the second weight is used to represent the preference degree of the second user for the target video;

[0095] Based on the recommendation weights of the P target videos, screen out the videos recommended to the first user from the P target videos.

[0096] When matching users, in the case of matching N second users, each second user has a certain degree of similarity with the first user. In some embodiments, the distance between the first vector and each vector of the pre-stored user metadata can be calculated to find N deep-behavior users similar to the first user. The closer the distance between the first vector and the pre-stored vector, the more similar the first user is to the second user corresponding to the vector. The similar second users will have corresponding similarity weights, that is, the first weight. The closer the distance between the first vector and the pre-stored vector, the greater the first weight. For example, for the first user, the similarity weight between the deep-behavior user 1 and the first user is 0.9, the similarity weight between the deep-behavior user 2 and the first user is 0.88, and the similarity weight between the deep-behavior user 3 and the first user is 0.7.

[0097] In some embodiments, for each shallow-behavior user, at most 10 deep-behavior users most similar to it can be found, that is, N is less than or equal to 10.

[0098] For the similar deep-behavior users, the target videos they historically preferred are at least one, and the user also has a corresponding preference weight for each target video, that is, the second weight. For each target video, the first weight corresponding to the second user who prefers the target video can be multiplied by the second weight corresponding to the target video to obtain the recommendation weight for recommending the target video to the first user. Among them, in some embodiments, when recommending videos for similar deep-behavior users, the preference weight of the deep-behavior user for the target video, that is, the second weight, can be determined based on relevant parameters such as the number of video clicks and the number of video views. For example, the more times the second user has historically watched the target video, the greater the second weight.

[0099] In some embodiments, the target videos with the top-ranked recommendation weights can be recommended to the first user. In this way, content that the shallow-behavior user may be interested in can be recommended, reducing the exploration cost of video recommendations, quickly capturing the interests of this part of users, and improving their user experience.

[0100] To more detailedly elaborate the video recommendation method of this embodiment, the following specific example is used for detailed description.Figure 2 is a schematic flowchart of a video recommendation method for a specific example. As Figure 2 shown, when the user goes online, the user requests to obtain a video in real time, and the user's metadata can be obtained. According to the user's metadata, the vector representation of the first feature description information is obtained from the storage, aggregated into emb1, that is, the fourth vector, and the vector representation of the second feature description information is obtained from the storage, aggregated into emb2, that is, the fifth vector. Concatenate emb1 and emb2 to form the vector representation of the user, that is, the first vector. By calculating the distances between the first vector and the vectors of the user metadata pre-stored, search for the second vector with a distance close to the first vector from the vectors of the user metadata pre-stored, so as to find the deep behavior users similar to this user, and multiply the preference weight of the deep behavior users for the target video, that is, the second weight, by the similarity weight between the deep behavior users and this user, that is, the first weight, to obtain the recommendation weight of the target video, and screen out the target videos with the top-ranked recommendation weights and push them to the user.

[0101] As Figure 3 shown, an embodiment of the present invention also provides a video recommendation device 300, including:

[0102] A determination module 301, configured to determine the behavior type of the first user in the video application;

[0103] An acquisition module 302, configured to obtain the metadata of the first user when the behavior type of the first user is the first behavior type, where the first behavior type is used to indicate that the behavior activity of the user in the video application is less than a first preset threshold;

[0104] A matching module 303, configured to match N second users similar to the metadata of the first user, where the behavior type of the second users is the second behavior type, and the second behavior type is used to indicate that the behavior activity of the user in the video application is greater than a second preset threshold, the second preset threshold is greater than the first preset threshold, and N is a positive integer;

[0105] A recommendation module 304, configured to perform video recommendation for the first user based on P target videos that N second users have historically preferred in the video application, where P is a positive integer.

[0106] Optionally, the behavior activity of the user in the video application is determined by the behavior data of the user in the video application. The determination module 301 is specifically configured to:

[0107] Obtain the behavior data of the first user in the video application, where the behavior data includes: the number of videos triggered by the first user in the video application, and the time of the most recent video trigger;

[0108] When it is determined that the behavior data meets the first preset condition, determine that the behavior type of the first user is the first behavior type; the first preset condition is: the number of videos triggered by the first user in the video application is less than or equal to the third preset threshold, and the time of the most recent video trigger is before the first preset time length; or, the first preset condition is: the number of videos triggered by the first user in the video application is zero; wherein, when it is determined that the behavior data meets the first preset condition, determine that the behavior activity of the first user in the video application is less than the first preset threshold.

[0109] Optionally, the determining module 301 is further configured to:

[0110] When it is determined that the behavior data meets the second preset condition, determine that the behavior type of the first user is the second behavior type; the second preset condition is: the number of videos triggered by the first user in the video application is greater than the fourth preset threshold, and the time of the most recent video trigger is within the second preset time length, the fourth preset threshold is greater than the third preset threshold, and the first preset time length is greater than the second preset time length; wherein, when it is determined that the behavior data meets the first preset condition, determine that the behavior activity of the first user in the video application is less than the first preset threshold.

[0111] Optionally, the matching module 303 is specifically configured to:

[0112] Obtain a first vector, where the first vector is a vector representation of the metadata of the first user;

[0113] Match N second vectors with distances close to the first vector from the vectors of the user metadata stored in advance, and obtain N second users whose metadata is similar to that of the first user, where the second vector is a vector representation of the metadata of the second user.

[0114] Optionally, the metadata includes M feature description information, M is a positive integer greater than 1, and the device further includes:

[0115] A vector representation module, configured to perform vector representation on each of the feature description information by using a target model, and obtain L third vectors of the M feature description information, where L is a positive integer less than or equal to M;

[0116] The matching module 303 is specifically configured to aggregate the L third vectors to obtain the first vector.

[0117] Optionally, the M feature description information includes: K first feature description information and J second feature description information. The K first feature description information is related to the feature information of the first user, and the J second feature description information is related to Q video episodes triggered by the first user in the video application. The sum of K and J is equal to M, and Q is a positive integer;

[0118] The target model includes a first model and a second model. The vector representation module is specifically configured to:

[0119] Input the K first feature description information into the first model for unsupervised training, and obtain K third vectors of the K first feature description information when the training is completed;

[0120] Input the J second feature description information into the second model for natural language processing to obtain Q third vectors corresponding to the Q video episodes;

[0121] The matching module 303 is specifically configured to:

[0122] Aggregate the K third vectors of the K first feature description information to obtain a fourth vector; and aggregate the Q third vectors corresponding to the Q video episodes to obtain a fifth vector;

[0123] Concatenate the fourth vector and the fifth vector to obtain the first vector.

[0124] Optionally, the recommendation module 304 is specifically configured to:

[0125] For each of the target videos, multiply the first weight corresponding to the second user who prefers the target video by the second weight corresponding to the target video to obtain a recommendation weight for recommending the target video to the first user. The first weight is used to characterize the similarity between the second user and the first user, and the second weight is used to characterize the preference degree of the second user for the target video;

[0126] Based on the recommendation weights of the P target videos, screen out the videos recommended to the first user from the P target videos.

[0127] An embodiment of the present invention further provides an electronic device, as Figure 4 shown, including a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete communication with each other through the communication bus 404.

[0128] The memory 403 is used to store a computer program;

[0129] The processor 401, when executing the program stored in the memory 403, implements the following steps:

[0130] Determine the behavior type of the first user in the video application;

[0131] When the behavior type of the first user is the first behavior type, obtain the metadata of the first user, where the first behavior type is used to indicate that the behavior activity of the user in the video application is less than the first preset threshold;

[0132] Match N second users whose metadata is similar to that of the first user. The behavior type of the second users is the second behavior type, where the second behavior type is used to indicate that the behavior activity of the users in the video application is greater than the second preset threshold, the second preset threshold is greater than the first preset threshold, and N is a positive integer;

[0133] Based on P target videos that N second users have historically preferred in the video application, perform video recommendations for the first user, where P is a positive integer.

[0134] Optionally, the behavior activity of the user in the video application is determined by the behavior data of the user in the video application. When the computer program is executed by the processor 401, it is also used for:

[0135] Obtain the behavior data of the first user in the video application, where the behavior data includes: the number of videos triggered by the first user in the video application, and the time of the most recent video trigger;

[0136] When it is determined that the behavior data meets the first preset condition, determine that the behavior type of the first user is the first behavior type; the first preset condition is: the number of videos triggered by the first user in the video application is less than or equal to the third preset threshold, and the time of the most recent video trigger is before the first preset time length; or, the first preset condition is: the number of videos triggered by the first user in the video application is zero; where, when it is determined that the behavior data meets the first preset condition, it is determined that the behavior activity of the first user in the video application is less than the first preset threshold.

[0137] Optionally, when the computer program is executed by the processor 401, it is also used for:

[0138] When it is determined that the behavior data meets the second preset condition, determine that the behavior type of the first user is the second behavior type; the second preset condition is that the number of videos triggered by the first user in the video application is greater than the fourth preset threshold, and the time of the most recent video trigger is within the second preset time length, the fourth preset threshold is greater than the third preset threshold, and the first preset time length is greater than the second preset time length; wherein, when it is determined that the behavior data meets the second preset condition, determine that the behavior activity of the first user in the video application is greater than the second preset threshold.

[0139] Optionally, when the computer program is executed by the processor 401, it is further configured to:

[0140] Obtain a first vector, where the first vector is a vector representation of the metadata of the first user;

[0141] Match N second vectors with distances close to the first vector from the vectors of the user metadata stored in advance, and obtain N second users whose metadata is similar to that of the first user, where the second vector is a vector representation of the metadata of the second user.

[0142] Optionally, the metadata includes M feature description information, where M is a positive integer greater than 1. When the computer program is executed by the processor 401, it is further configured to:

[0143] Use a target model to perform vector representation on each of the feature description information, and obtain L third vectors of the M feature description information, where L is a positive integer less than or equal to M;

[0144] Aggregate the L third vectors to obtain the first vector.

[0145] Optionally, the M feature description information includes: K first feature description information and J second feature description information. The K first feature description information is related to the feature information of the first user, and the J second feature description information is related to Q video episodes triggered by the first user in the video application. The sum of K and J is equal to M, and Q is a positive integer. The target model includes a first model and a second model. When the computer program is executed by the processor 401, it is further configured to:

[0146] Input the K first feature description information into the first model for unsupervised training, and when the training is completed, obtain K third vectors of the K first feature description information;

[0147] Input the J second feature description information into the second model for natural language processing, and obtain Q third vectors corresponding to the Q video episodes;

[0148] Aggregate the K third vectors of the K first feature description information to obtain a fourth vector; and aggregate the Q third vectors corresponding to the Q video episodes to obtain a fifth vector;

[0149] Concatenate the fourth vector and the fifth vector to obtain the first vector.

[0150] Optionally, when the computer program is executed by the processor 401, it is further configured to:

[0151] For each of the target videos, multiply the first weight corresponding to the second user who prefers the target video by the second weight corresponding to the target video to obtain a recommendation weight for recommending the target video to the first user, where the first weight is used to characterize the similarity between the second user and the first user, and the second weight is used to characterize the preference degree of the second user for the target video;

[0152] Based on the recommendation weights of the P target videos, screen out the videos recommended to the first user from the P target videos.

[0153] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0154] The communication interface is used for communication between the above electronic device and other devices.

[0155] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0156] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it may also be a Digital Signal Processor (DSP for short), an Application Specific Integrated Circuit (ASIC for short), a Field-Programmable Gate Array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0157] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the video recommendation method described in any one of the above embodiments.

[0158] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute the video recommendation method described in any one of the above embodiments.

[0159] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a Solid State Disk (SSD)).

[0160] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0161] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.

[0162] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A video recommendation method, characterized in that, The method includes: Determining the behavior type of a first user in a video application; When the behavior type of the first user is a first behavior type, obtaining metadata of the first user, where the first behavior type is used to indicate that the behavior activity of the user in the video application is less than a first preset threshold; Matching N second users whose metadata is similar to that of the first user, where the behavior type of the second users is a second behavior type, the second behavior type is used to indicate that the behavior activity of the user in the video application is greater than a second preset threshold, the second preset threshold is greater than the first preset threshold, and N is a positive integer; Based on P target videos that the N second users have historically preferred in the video application, performing video recommendation for the first user, where P is a positive integer.

2. The method according to claim 1, characterized in that, The behavior activity of the user in the video application is determined by the behavior data of the user in the video application. Determining the behavior type of the first user in the video application includes: Obtaining the behavior data of the first user in the video application, where the behavior data includes: the number of videos triggered by the first user in the video application, and the time of the most recent video trigger; When it is determined that the behavior data meets a first preset condition, determining that the behavior type of the first user is a first behavior type; the first preset condition is: the number of videos triggered by the first user in the video application is less than or equal to a third preset threshold, and the time of the most recent video trigger is before a first preset time length; or, the first preset condition is: the number of videos triggered by the first user in the video application is zero; where, when it is determined that the behavior data meets the first preset condition, it is determined that the behavior activity of the first user in the video application is less than the first preset threshold.

3. The method according to claim 2, wherein Determining the behavior type of the first user in the video application further includes: When it is determined that the behavior data meets a second preset condition, determining that the behavior type of the first user is a second behavior type; the second preset condition is: the number of videos triggered by the first user in the video application is greater than a fourth preset threshold, and the time of the most recent video trigger is within a second preset time length, the fourth preset threshold is greater than the third preset threshold, and the first preset time length is greater than the second preset time length; where, when it is determined that the behavior data meets the second preset condition, it is determined that the behavior activity of the first user in the video application is greater than the second preset threshold.

4. The method according to any one of claims 1 to 3, characterized in that Matching N second users whose metadata is similar to that of the first user includes: Obtaining a first vector, where the first vector is a vector representation of the metadata of the first user; Matching N second vectors that are close in distance to the first vector from the vectors of the user metadata stored in advance, to obtain N second users whose metadata is similar to that of the first user, where the second vector is a vector representation of the metadata of the second user.

5. The method according to claim 4, wherein The metadata includes M feature description information, where M is a positive integer greater than 1. Before matching N second users whose metadata is similar to that of the first user, the method further includes: Use the target model to perform vector representation on each of the feature description information, obtaining L third vectors of the M feature description information, where L is a positive integer less than or equal to M; The obtaining of the first vector includes: Aggregate the L third vectors to obtain the first vector.

6. The method according to claim 5, wherein The M feature description information includes: K first feature description information and J second feature description information. The K first feature description information is related to the feature information of the first user, and the J second feature description information is related to Q video episodes triggered by the first user in the video application. The sum of K and J is equal to M, and Q is a positive integer; The target model includes a first model and a second model. The using the target model to perform vector representation on each of the feature description information, obtaining L third vectors of the M feature description information, includes: Input the K first feature description information into the first model for unsupervised training, and when the training is completed, obtain K third vectors of the K first feature description information; Input the J second feature description information into the second model for natural language processing to obtain Q third vectors corresponding to the Q video episodes; The obtaining of the first vector includes: Aggregate the K third vectors of the K first feature description information to obtain a fourth vector; and aggregate the Q third vectors corresponding to the Q video episodes to obtain a fifth vector; Concatenate the fourth vector and the fifth vector to obtain the first vector.

7. The method according to claim 1, wherein The video recommendation for the first user based on P target videos of the historical preferences of N second users in the video application includes: For each of the target videos, multiply the first weight corresponding to the second user who prefers the target video by the second weight corresponding to the target video to obtain the recommendation weight for recommending the target video to the first user. The first weight is used to represent the similarity degree between the second user and the first user, and the second weight is used to represent the preference degree of the second user for the target video; Based on the recommendation weights of the P target videos, screen out the videos recommended to the first user from the P target videos.

8. A video recommendation device, characterized in that, The device includes: A determination module, configured to determine the behavior type of the first user in the video application; An obtaining module, configured to obtain the metadata of the first user when the behavior type of the first user is the first behavior type, where the first behavior type is used to indicate that the behavior activity of the user in the video application is less than a first preset threshold; A matching module, configured to match N second users similar to the metadata of the first user, where the behavior type of the second user is the second behavior type, and the second behavior type is used to indicate that the behavior activity of the user in the video application is greater than a second preset threshold, and the second preset threshold is greater than the first preset threshold, and N is a positive integer; A recommendation module, configured to perform video recommendation for the first user based on P target videos of the historical preferences of N second users in the video application, where P is a positive integer.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the method described in any one of claims 1-7 when executing the program stored on the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1-7.

11. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by the processor, it implements the method described in any one of claims 1-7.