Video recommendation method, apparatus, device, and storage medium

HK40083063BActive Publication Date: 2026-07-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Patents
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2023-04-27
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing video platforms' collaborative filtering-based video recommendation methods are not accurate enough when new users or new videos are added, and the lack of interactive behavior information leads to poor recommendation results.

Method used

By simulating the input of sample data during cold start through a pre-trained vector reconstruction model, the embedding vectors and feature vectors of users and videos are mapped to the reconstructed embedding vectors to determine the interaction probability between users and candidate videos, and then recommend target videos.

Benefits of technology

It accurately recommends videos during cold starts, improving recommendation effectiveness, and saves storage space and improves recommendation efficiency during non-cold starts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This application provides a video recommendation method, apparatus, device, and storage medium, relating to the field of communication technology. The method includes: upon receiving a video acquisition request, acquiring a candidate video list; acquiring the interaction probability between the user and each candidate video in the candidate video list, the interaction probability being determined based on the user's reconstructed embedding vector and the candidate video's reconstructed embedding vector; the user's reconstructed embedding vector being determined based on the user's embedding vector, the user's feature vector, and a pre-trained vector reconstruction model; the candidate video's reconstructed embedding vector being determined based on the candidate video's embedding vector, the candidate video's feature vector, and the vector reconstruction model; simulating the input of sample data during cold start when training the vector reconstruction model; and determining the target video to recommend to the user from the candidate video list based on the interaction probability between the user and each candidate video in the candidate video list. This method can accurately recommend videos to users during cold start, improving the video recommendation effect during cold start.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a video recommendation method, apparatus, device, and storage medium. Background Technology

[0002] With the development of internet technology, more and more users are watching videos through video platforms (such as video websites or video applications). Because video platforms offer a vast variety and quantity of videos, and new videos and users are constantly joining each platform, accurately recommending videos that users prefer from this massive amount of content is particularly important.

[0003] Currently, video platforms generally use video recommendation methods based on collaborative filtering algorithms. This method predicts the videos that current users prefer based on the historical interaction behavior information of existing user groups with existing videos. However, new users and new videos only have initial features at the time of registration and upload, respectively, and lack interaction behavior information. Therefore, the videos recommended by this method to users during cold starts (i.e., when new users or new videos are added) are not accurate. Summary of the Invention

[0004] This application provides a video recommendation method, apparatus, device, and storage medium to improve the accuracy of recommended videos during cold starts.

[0005] Firstly, this application provides a video recommendation method, including:

[0006] Upon receiving a video retrieval request, obtain a list of candidate videos;

[0007] The interaction probability between a user and each candidate video in the candidate video list is obtained. The interaction probability is determined based on the reconstructed embedding vector of the user and the reconstructed embedding vector of the candidate video. The reconstructed embedding vector of the user is determined based on the user's embedding vector, the user's feature vector, and a pre-trained vector reconstruction model. The reconstructed embedding vector of the candidate video is determined based on the candidate video's embedding vector, the candidate video's feature vector, and the vector reconstruction model. The input of sample data during cold start is simulated when training the vector reconstruction model.

[0008] Based on the interaction probability between the user and each candidate video in the candidate video list, a target video is determined from the candidate video list to be recommended to the user.

[0009] Secondly, this application provides a video recommendation device, comprising:

[0010] The first acquisition module is used to acquire a list of candidate videos after receiving a video acquisition request;

[0011] The second acquisition module is used to acquire the interaction probability between the user and each candidate video in the candidate video list. The interaction probability is determined based on the reconstructed embedding vector of the user and the reconstructed embedding vector of the candidate video. The reconstructed embedding vector of the user is determined based on the user's embedding vector, the user's feature vector, and a pre-trained vector reconstruction model. The reconstructed embedding vector of the candidate video is determined based on the candidate video's embedding vector, the candidate video's feature vector, and the vector reconstruction model. The input of sample data during cold start is simulated when training the vector reconstruction model.

[0012] The determining module is used to determine the target video to recommend to the user from the candidate video list based on the interaction probability between the user and each candidate video in the candidate video list.

[0013] Thirdly, this application provides a terminal device, including: a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the method of the first aspect.

[0014] Fourthly, this application provides a computer-readable storage medium for storing a computer program that causes a computer to perform the method of the first aspect.

[0015] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method of the first aspect.

[0016] In summary, this application, upon receiving a video acquisition request, obtains a list of candidate videos. By simulating the input of sample data during cold start when pre-training the vector reconstruction model, for each user or video, the embedding vector and feature vector are mapped to the same vector to obtain a reconstructed embedding vector. Therefore, for new users or new videos, when the embedding vector is missing, the embedding vector can be reconstructed from the feature vector. Furthermore, based on the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video, the interaction probability between the user and each candidate video can be determined. Finally, based on the interaction probability between the user and each candidate video in the candidate video list, the target video recommended to the user is determined from the candidate video list. This allows for accurate video recommendation to the user during cold start, improving the video recommendation effect during cold start.

[0017] Furthermore, in this application, videos can be accurately recommended to users even during non-cold start conditions. Moreover, since the user's reconstructed embedding vector can uniquely represent the user's interaction behavior information with the video and the user's feature information, and the video's reconstructed embedding vector can uniquely represent the video's interaction behavior information with the user and the video's feature information, storage space can be saved and recommendation efficiency can be improved. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram illustrating an application scenario of the video recommendation method provided in the embodiments of this application;

[0020] Figure 2 A flowchart illustrating a video recommendation method provided in this application embodiment;

[0021] Figure 3 A flowchart illustrating the training process of a vector reconstruction model in a video recommendation method provided in this application embodiment;

[0022] Figure 4 A flowchart illustrating a video recommendation method provided in this application embodiment;

[0023] Figure 5 This is a schematic diagram of the structure of a video recommendation device provided in an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0027] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0028] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0029] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0030] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence.

[0031] Deep learning is the core of machine learning, and it typically includes techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching-based learning.

[0032] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0033] The solutions provided in this application involve technologies such as natural language processing and deep learning in artificial intelligence, and are specifically illustrated through the following embodiments.

[0034] Before introducing the technical solution of this application, the following is a brief introduction to relevant knowledge about this application:

[0035] 1. Embedded vector: A vector extracted from the user-video interaction network to represent a single user or a single video. That is, it describes the relationship between the user or video in the user-video interaction network in the form of a vector. In the embodiments of this application, the embedded vectors of the user and the embedded vectors of the video are involved. The embedded vector of the user can characterize the interactive behavior information between the user and the video, and the embedded vector of the video can characterize the interactive behavior information between the video and the user.

[0036] 2. A user's feature vector is a vector representing the user's characteristic information, which may include gender, age, region, and the model of the device used. The relationship between users and feature vectors is one-to-many, meaning a user can be represented by multiple feature vectors {a}. u i i = 1, 2, ..., N u} jointly describe, where u is the user ID, i represents the i-th feature information of user u, and N u This represents the total number of user characteristic information items.

[0037] 3. Video feature vectors are vectors used to represent the feature information of a video. This feature information can include length, tags, and multimedia features, such as background music, clip templates, and objects in the video (e.g., people or objects). The relationship between videos and feature vectors is one-to-many, meaning a video can have multiple feature vectors {a}. v i i = 1, 2, ..., N v} jointly describe, where v is the video ID, i represents the i-th feature information of the video, and N v This represents the total number of feature information items in the video.

[0038] 4. The reconstructed embedding vector of a user refers to the vector obtained by reconstructing the embedding vector from the information in the user's feature vector when the user's embedding vector is missing (new users do not have embedding vectors).

[0039] 5. The reconstructed embedding vector of a video refers to the vector obtained by reconstructing the embedding vector from the feature vector of the video when the embedding vector of the video is missing (the new video has no embedding vector).

[0040] In related technologies, new users and new videos only possess initial features from registration and upload, lacking interaction behavior information. Therefore, video recommendations to users during cold starts (i.e., when a new user or video is added) are inaccurate. To address this issue, since new users or videos lack interaction behavior information and thus cannot obtain embedding vectors, this application simulates the input of sample data during cold starts when pre-training a vector reconstruction model. For users or videos, the embedding vector and feature vector are mapped to the same vector to obtain a reconstructed embedding vector. Therefore, for new users or videos, when the embedding vector is missing, the embedding vector can be reconstructed from the feature vector. Furthermore, based on the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video, the interaction probability between the user and each candidate video can be determined. Finally, based on the interaction probability between the user and each candidate video in the candidate video list, the target video recommended to the user is determined from the candidate video list. This allows for accurate video recommendations to users during cold starts, improving the video recommendation effect during cold starts.

[0041] The method of this embodiment can accurately recommend videos to users even during non-cold start scenarios. Furthermore, since the user's reconstructed embedding vector uniquely represents the user's interaction with the video and the user's feature information, and the video's reconstructed embedding vector uniquely represents the video's interaction with the user and the video's feature information, storage space can be saved and recommendation efficiency improved. The technical solution provided in this application will now be described in detail with reference to the accompanying drawings.

[0042] Next, examples of application scenarios involved in the embodiments of this application will be provided.

[0043] The video recommendation method provided in this application embodiment can be applied to at least the following application scenarios, which will be described below in conjunction with the accompanying drawings.

[0044] For example, Figure 1 This is a schematic diagram illustrating an application scenario of the video recommendation method provided in the embodiments of this application, such as... Figure 1As shown, the application scenario in this embodiment involves terminal device 1 and server 2. Terminal device 1 can be a portable device such as a mobile phone, tablet, or laptop, or a personal computer. Server 2 can be any device capable of providing internet services. Terminal device 1 and server 2 can communicate via a network. Terminal device 1 has an application with video playback functionality installed and running. The user can send a video request to server 2 through the application running on terminal device 1. After receiving the video request, server 2 executes the video recommendation method provided in this embodiment to recommend a target video to the user. The user can then play the target video through the application running on terminal device 1. The target video is a video that the user is interested in or prefers.

[0045] For example, the video recommendation method provided in this application embodiment can be applied to scenarios where applications with video playback functions have a large number of new users and new videos.

[0046] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0047] Figure 2 This is a flowchart illustrating a video recommendation method provided in an embodiment of this application. This video recommendation method can be executed by a video recommendation device, which can be implemented through software and / or hardware. The video recommendation device can be a server. Figure 2 As shown, the method in this embodiment may include:

[0048] S101. After receiving the video acquisition request, obtain the candidate video list.

[0049] Specifically, in one feasible approach, after responding to a user-triggered action to retrieve recommended videos, the terminal device sends a video retrieval request to the server. This request may carry the user's identifier. The user's action to retrieve recommended videos can be triggered by the user opening an application with video playback functionality installed on the terminal device, or by the user clicking a prompt button on the current interface or swiping the current screen while watching a video through an application with video playback functionality.

[0050] Optionally, the candidate video list can be retrieved from memory. The candidate video list can include the identifiers (such as IDs or numbers) of the candidate videos. The candidate videos can be all videos in an application with video playback functionality that are up to the current time. Since videos also have a time limit, the candidate videos can also be all videos within a preset time period before the current time. The preset time period can be, for example, one month, three months, six months, or one year, etc.

[0051] S102. Obtain the interaction probability between the user and each candidate video in the candidate video list. The interaction probability is determined based on the user's reconstructed embedding vector and the candidate video's reconstructed embedding vector. The user's reconstructed embedding vector is determined based on the user's embedding vector, the user's feature vector, and the pre-trained vector reconstruction model. The candidate video's reconstructed embedding vector is determined based on the candidate video's embedding vector, the candidate video's feature vector, and the vector reconstruction model. Simulate the input of sample data during cold start when training the vector reconstruction model.

[0052] Specifically, the interaction probability between a user and a candidate video is determined based on the user's reconstructed embedding vector and the candidate video's reconstructed embedding vector. Optionally, the interaction probability can be the cosine distance, inner product, or Euclidean distance between the user's and candidate video's reconstructed embedding vectors. The interaction probability is directly proportional to the cosine distance and inner product, and inversely proportional to the Euclidean distance. Taking the interaction probability as the cosine distance as an example, the user's reconstructed embedding vector e... u 'and the reconstructed embedding vector e of the candidate video v 'cosine distance The larger the cosine distance, the greater the probability of interaction; the smaller the cosine distance, the smaller the probability of interaction.

[0053] S103. Based on the interaction probability between the user and each candidate video in the candidate video list, determine the target video to recommend to the user from the candidate video list.

[0054] In this embodiment, the reconstructed embedding vector of a user is determined based on the user's embedding vector, the user's feature vector, and a pre-trained vector reconstruction model. The reconstructed embedding vector of a candidate video is determined based on the candidate video's embedding vector, the candidate video's feature vector, and the vector reconstruction model. In this embodiment, the input of sample data during a cold start is simulated when training the vector reconstruction model. A cold start refers to the addition of a new user or a new video. For a new user, the new user's embedding vector is a zero vector, which can be determined based on the new user's feature vector and the pre-trained vector reconstruction model. For a new video, the new video's embedding vector is also a zero vector, which can be determined based on the new video's feature vector and the pre-trained vector reconstruction model. For historical users or historical videos, which have interaction behavior information, the embedding vectors of historical users and historical videos can be determined. These can be determined separately using the vector reconstruction model. Furthermore, the interaction probability of a candidate video can be determined based on the user's reconstructed embedding vector and the candidate video's reconstructed embedding vector. With the interaction probability between the user and each candidate video in the candidate video list, the target video recommended to the user can be determined from the candidate video list. In this embodiment, the user's embedding vector and the user's feature vector are mapped to the same vector through a vector reconstruction model to become the user's reconstructed embedding vector. The embedding vector and the feature vector of each candidate video are mapped to the same vector through a vector reconstruction model to become the reconstructed embedding vector of each candidate video. This saves storage space and improves recommendation efficiency when storing the user's reconstructed embedding vector and the candidate video's reconstructed embedding vector.

[0055] In this embodiment, as one implementable method, obtaining the interaction probability between the user and each candidate video in the candidate video list in S102 may include:

[0056] S1021. Obtain the user's reconstructed embedding vector.

[0057] Optionally, the reconstructed embedding vector of the user can be obtained, specifically as follows:

[0058] If the user is a new user, the 0 vector is determined as the user's embedding vector, and the user's feature vector is determined based on the user's feature information. The 0 vector and the user's feature vector are concatenated and input into the vector reconstruction model, and the reconstructed embedding vector of the user is output.

[0059] If the user is a historical user, the user's embedding vector is obtained based on the interaction network between the user and historical videos. The user's embedding vector is then concatenated with the user's feature vector and input into the vector reconstruction model, and the reconstructed embedding vector of the user is output.

[0060] Here, the user's embedding vector is a matrix, the user's feature vector is a matrix, and the concatenation of the two vectors is the concatenation of the two matrices.

[0061] S1022. Obtain the reconstructed embedding vector for each candidate video.

[0062] Optionally, the reconstructed embedding vector for each candidate video can be obtained, specifically as follows:

[0063] If the candidate video is a new video, the 0 vector is determined as the embedding vector of the candidate video, and the feature vector of the candidate video is determined according to the feature information of the candidate video. The 0 vector and the feature vector of the candidate video are concatenated and input into the vector reconstruction model, and the reconstructed embedding vector of the candidate video is output.

[0064] If the candidate video is a historical video, the embedding vector of the candidate video is obtained based on the interaction network between historical users and the candidate video. The embedding vector of the candidate video is concatenated with the feature vector of the candidate video and then input into the vector reconstruction model, and the reconstructed embedding vector of the candidate video is output.

[0065] S1023. Determine the interaction probability between the user and each candidate video based on the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video.

[0066] Optionally, in S1023, the interaction probability between the user and each candidate video is determined based on the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video. Specifically, this can be:

[0067] Calculate the cosine distance, inner product, or Euclidean distance between the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video, and determine the interaction probability of each candidate video based on the cosine distance, inner product, or Euclidean distance between the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video.

[0068] In this embodiment, as an implementable approach, the vector reconstruction model is trained on multiple sample data sets. Each sample data set includes the embedding vector of a sample user, the feature vector of a sample user, the embedding vector of a sample video, and the feature vector of a sample video. The sample users are those who have interacted with all the sample videos, and the sample videos are also those that have interacted with all the sample users.

[0069] When training the vector reconstruction model, the inputs to the vector reconstruction model are the embedding vectors of the sample users, the feature vectors of the sample users, the embedding vectors of the sample videos, and the feature vectors of the sample videos. The outputs of the vector reconstruction model are the reconstructed embedding vectors of the sample users and the reconstructed embedding vectors of the sample videos.

[0070] When training the vector reconstruction model, the input of sample data during cold start is simulated to minimize the value of the objective function. The vector reconstruction model is the vector reconstruction model corresponding to the minimum value of the objective function. The objective function represents the difference between the first interaction probability between the sample user and the sample video and the second interaction probability between the sample user and the sample video. The first interaction probability is determined based on the embedding vector of the sample user and the embedding vector of the sample video, and the second interaction probability is determined based on the reconstructed embedding vector of the sample user and the reconstructed embedding vector of the sample video.

[0071] During a cold start, the embedding vectors of new users and new videos are both zero vectors. When training the vector reconstruction model, the input of the sample data during the cold start is simulated. That is, the embedding vectors of some sample users are randomly set to zero vectors for multiple training sessions, and the embedding vectors of some sample videos are randomly set to zero vectors for multiple training sessions.

[0072] Optionally, in this embodiment, simulating the input of cold-start sample data when training the vector reconstruction model can be done by simulating the input of user cold-start sample data and video cold-start sample data. For example, the input of cold-start sample data can be simulated based on the Dropout strategy in deep learning.

[0073] The process of simulating the input of sample data during user cold start when training the vector reconstruction model may include: training the vector reconstruction model multiple times based on the sample data; randomly selecting at least one first sample user from the sample users each time the vector reconstruction model is trained; setting the embedding vector of each first sample user to a 0 vector and concatenating it with the feature vector of each first sample user before inputting it into the vector reconstruction model.

[0074] The process of simulating the input of sample data during video cold start when training the vector reconstruction model may include: training the vector reconstruction model multiple times based on the sample data; in each training of the vector reconstruction model, randomly selecting at least one first sample video from the sample videos, setting the embedding vector of each first sample video to a 0 vector and concatenating it with the feature vector of each first sample video before inputting it into the vector reconstruction model.

[0075] As an feasible approach, the vector reconstruction model includes a user-side multilayer perceptron and a video-side multilayer perceptron. The input of the user-side multilayer perceptron is the first vector obtained by concatenating the embedding vector of the sample user with the feature vector of the sample user, and the output of the user-side multilayer perceptron is the reconstructed embedding vector of the sample user. The input of the video-side multilayer perceptron is the second vector obtained by concatenating the embedding vector of the sample video with the feature vector of the sample video, and the output of the video-side multilayer perceptron is the reconstructed embedding vector of the sample video.

[0076] Optionally, the objective function L is:

[0077] L=||(e u e v T -f U (e u ||a u )f V (e v ||a v ) T )||2

[0078] Among them, e u e is the embedding vector of the sample user. v Let a be the embedding vector of the sample video. u Let a be the feature vector of the sample users. v f is the feature vector of the sample video. U For the user-side multilayer perceptron, f V For video-side multilayer perceptron, e u ||a u Let e ​​be the first vector. v ||a v Let f be the second vector, ||X||2 be the L2 norm of X, and f U (e u ||a u f is the reconstructed embedding vector of the sample user. V (e v ||a v ) represents the reconstructed embedding vector of the sample video.

[0079] When the objective function is as shown above, the input of sample data during cold start is simulated when training the vector reconstruction model to minimize the value of the objective function L. The vector reconstruction model is the vector reconstruction model corresponding to the minimum value of the objective function L, that is, the f corresponding to the minimum value of the objective function L. U and f V .

[0080] When simulating the input of sample data during user cold start when reconstructing the model using training vectors, at least one first sample user is randomly selected from the sample users. The embedding vector of each first sample user is set to a zero vector, and the resulting objective function L is: L=||(e u e v T -f U (0||a u )f V (e v ||a v ) T )2.

[0081] When simulating the input of sample data during video cold start when reconstructing the model using training vectors, at least one first sample video is randomly selected from the sample videos, and the embedding vector of each first sample video is set to a zero vector. The resulting objective function L is: L=||(e u e v T -f U (e u ||a u )f V (0||a v ) T )||2.

[0082] In the above embodiments, S103 can be:

[0083] The system selects a preset number of candidate videos as target videos from the candidate video list, in descending order of interaction probability. The preset number could be, for example, 10 or 5. Taking 10 as an example, this means selecting the 10 candidate videos with the highest interaction probability from the candidate video list as target videos.

[0084] It should be noted that, in one feasible implementation, the method of this embodiment can also be executed by a terminal device. For example, when the terminal device has strong storage and processing capabilities, the method of this embodiment can also be executed by the terminal device. After responding to the user's triggered operation to obtain recommended videos, the terminal device executes the method provided in this application embodiment to recommend target videos to the user. The target videos can be videos that the user is interested in.

[0085] The video recommendation method provided in this embodiment, upon receiving a video acquisition request, obtains a list of candidate videos. By simulating the input of sample data during cold start when pre-training the vector reconstruction model, for each user or video, the embedding vector and feature vector are mapped to the same vector to obtain a reconstructed embedding vector. Therefore, for new users or new videos, when the embedding vector is missing, the embedding vector can be reconstructed from the feature vector. Then, based on the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video, the interaction probability between the user and each candidate video can be determined. Finally, based on the interaction probability between the user and each candidate video in the candidate video list, the target video recommended to the user is determined from the candidate video list. This allows for accurate video recommendation to the user during cold start, improving the video recommendation effect during cold start.

[0086] Furthermore, the video recommendation method provided in this embodiment can accurately recommend videos to users even during non-cold start conditions. Moreover, since the user's reconstructed embedding vector can uniquely represent the user's interaction behavior information with the video and the user's feature information, and the video's reconstructed embedding vector can uniquely represent the video's interaction behavior information with the user and the video's feature information, storage space can be saved and recommendation efficiency can be improved.

[0087] The following details a specific implementation method for obtaining the user's reconstructed embedding vector in S1021, where the user is a historical user. This method involves obtaining the user's embedding vector based on the interaction network between the user and historical videos. Specifically, it may include:

[0088] First, construct the interaction matrix R∈R between the user and all videos in the candidate videos. M×N Let R be the number of users (M) and the number of videos (N). Each element in R represents the strength of the interaction between a user and a video. User-video interaction behaviors include video exposure (i.e., being seen by the user), video completion (i.e., watching the entire video), user liking, user sharing, and user saving. Specifically, when constructing the user-video interaction matrix, based on the historical user-video interaction behaviors, all users and videos on the network are connected into an interaction network through interaction relationships. The edges between users and videos and their weights are determined by the strength of the interaction behavior. Specifically, if a pair of user-video interactions is strong, the connection between that pair of user-videos has a high weight; if the interaction is weak, the connection between that pair of user-videos has a low weight; if there is no interaction between users and videos, there is no connection between them. The strength of user-video interaction behaviors can be measured by the weighted sum of the interaction behaviors. Optionally, the strength of the user-video interaction relationship... Where b i w represents the number of the i-th interaction behavior between user u and video v. i Let represent the weight of the i-th interaction behavior between user u and video v. In this embodiment, the weights of all interaction behaviors are normalized, therefore we have K represents the total number of interaction behaviors.

[0089] Next, based on the node embedding method, the user-video interaction matrix described above, and the user-video interaction network constructed above, the embedding vectors for each user and each video are extracted.

[0090] The node embedding methods mentioned above can include matrix factorization (such as GRMF), network embedding (such as Node2Vec), and graph embedding (such as NGCF). Taking the Node2Vec embedding method as an example, based on the user-video interaction network constructed above, each node in the interaction network (each user or video is considered a node) is used as a starting point, and multiple trajectories are randomly walked (for example, 100 trajectories are walked for each node). Since the connection weights of different user-video pairs in the interaction network may be different, the node adjacency can be probability-weighted during the random walk. Before the random walk begins, the interaction matrix R∈R needs to be expanded. M×N The adjacency matrix, after expansion, is shown in the following equation:

[0091]

[0092] Next, all the trajectories obtained from the random walk are used as corpus input into the Word2Vec word vector embedding algorithm for training. The training result yields the embedding vector e for each user. u and the embedding vector e of each video v .

[0093] The following details the specific implementation of determining the user's feature vector based on the user's feature information in S1021 and the specific implementation of determining the candidate video's feature vector based on the candidate video's feature information in S1022.

[0094] Specifically, user characteristics can include gender, age, region, and the model of the device used. The embedding vector for user u is e. u The length of the embedding vector is d e The feature vector set of user u is A u ={a u i i = 1, 2, ..., N u}, where N u The number of user feature information, where each feature vector has a length of d. a For the embedding vector e of user u u Multiply by a trainable parameter matrix Get the query vector q u =W q T ·e u For A u eigenvector a in u i Multiply by two trainable parameter matrices W in sequence. k and Obtain the key vector k u i =Wk T ·a u i Sum vector v u i =W k T ·a u i d h Let q be the length of the query vector, key vector, and value vector. Thus, we obtain the query vector q for each user. u Key vector k u i Sum vector v u i .

[0095] For user u, using query vector q u The key vectors {k} of all the user's features in sequence u i i = 1, 2, ..., N u Perform inner product operations to obtain the attention score {att} for each feature. u i |att u i =q u ·k u i i = 1, 2, ..., N u Then, the attention score is normalized using Softmax to obtain the normalized attention score:

[0096]

[0097] Using the normalized attention score as the attention weight, all features of user u are integrated:

[0098]

[0099] At this point, the feature fusion based on the self-attention mechanism for user u is complete, resulting in the fused feature vector a for the user. u The feature fusion process for video v based on the self-attention mechanism is similar to the process described above, resulting in the fused feature vector a of the video. v '.

[0100] The above process obtains the user's feature vector and the video's feature vector through a single-head attention mechanism. It is also possible to obtain the user's feature vector and the video's feature vector through a multi-head attention mechanism. The user's feature vector and the video's feature vector obtained through the multi-head attention mechanism are more accurate than those obtained through the single-head attention mechanism.

[0101] In a multi-head attention mechanism, multiple attention modules will be constructed, meaning the aforementioned trainable parameter matrix will be represented as W. q h W k h W v h Here, h represents the number of the attention module, and there are a total of H attention modules. Correspondingly, the query vector q of user u can be obtained through the trainable parameter matrix under different attention modules. u h Key vector k u ih Sum vector v u ih Then, the fusion feature a is calculated. u h '. After obtaining the fusion features under all attention modules {a u h After concatenating the features of the user (i = 1, 2, ..., H), we obtain the final feature vector a. u =a u 1 '||a u 2 '||...||a u H The length of the eigenvector is H×d. h Similarly, the final feature vector a of the video can be obtained. v =a v 1 '||a v 2 '||...||a v H '.

[0102] The video recommendation method provided in this application will be described in detail below with reference to a specific embodiment.

[0103] Figure 3 A flowchart illustrating the training process of a vector reconstruction model in a video recommendation method provided in this application embodiment is shown below. Figure 3 As shown, the method in this embodiment may include:

[0104] S201. Obtain the embedding vectors of the sample users and the embedding vectors of the sample videos.

[0105] Specifically, the sample users are all users who have interacted with all the sample videos, and the sample videos are also all videos that have interacted with all the sample users.

[0106] As an implementable approach, S201 specifically includes:

[0107] S2011. Construct the interaction matrix R∈R for all sample users and all sample videos. M×N The number of sample users is M, and the number of sample videos is N.

[0108] Each element in R represents the strength of the interaction between a sample user and a sample video. User-video interactions include video exposure (i.e., being seen by the user), video completion (i.e., watching the entire video), user liking, user sharing, and user saving. Specifically, when constructing the interaction matrix between sample users and sample videos, sample users and sample videos are connected into an interaction network based on their historical interaction behaviors. The edges between sample users and sample videos and their weights are determined by the strength of the interaction. Specifically, if a pair of sample users and sample videos has a strong interaction, the connection between them has a high weight; if the interaction is weak, the connection has a low weight; if there is no interaction, there is no connection between them. The strength of the interaction between sample users and sample videos can be determined by a weighted sum of the interaction behaviors. Optionally, the strength of the interaction between sample users and sample videos... Where b i w represents the number of the i-th interaction behavior between sample user u and sample video v. i Let represent the weight of the i-th interaction behavior between sample user u and sample video v. In this embodiment, the weights of all interaction behaviors are normalized, therefore we have K represents the total number of interaction behaviors.

[0109] S2012. Based on the node embedding method, the interaction matrix between sample users and sample videos, and the interaction network between sample users and sample videos constructed above, extract the embedding vector of each sample user and the embedding vector of each sample video.

[0110] Taking the Node2Vec embedding method as an example, based on the interaction network of sample users and sample videos constructed above, each node in the interaction network (each sample user or sample video is considered a node) is used as a starting point, and multiple trajectories are randomly walked (e.g., 100 trajectories for each node). Since the connection weights of different sample user-sample video pairs in the interaction network may be different, the node adjacency can be probability-weighted during the random walk. Before the random walk begins, the interaction matrix R∈R needs to be expanded. M×N The adjacency matrix, after expansion, is shown in the following equation:

[0111]

[0112] Next, all the trajectories obtained from the random walk are used as corpus input into the Word2Vec word embedding algorithm for training. The training result yields the embedding vector e for each sample user. u and the embedding vector e of each sample video v .

[0113] S202. Obtain the feature vectors of the sample users and the feature vectors of the sample videos.

[0114] Specifically, the process of obtaining the feature vectors of sample users and the feature vectors of sample videos can be found in the specific implementations of S1021 and S1022 described in the above embodiments.

[0115] S203. Reconstruct the model based on the embedding vectors of sample users, the embedding vectors of sample videos, the feature vectors of sample users, the feature vectors of sample videos, and the training vectors of the objective function.

[0116] In this embodiment, the vector reconstruction model includes a user-side multilayer perceptron and a video-side multilayer perceptron, and the objective function L is:

[0117] L=||(e u e v T -f U (e u ||a u )f V (e v ||a v ) T )||2

[0118] Among them, e u e is the embedding vector of the sample user. v Let a be the embedding vector of the sample video. u Let a be the feature vector of the sample users. v f is the feature vector of the sample video. U For the user-side multilayer perceptron, f V For a video-side multilayer perceptron, ||X||2 is the L2 norm of X, f U (e u ||a u f is the reconstructed embedding vector of the sample user. V (e v ||a v ) represents the reconstructed embedding vector of the sample video.

[0119] When training the vector to reconstruct the model, the input of sample data during user cold start and the input of sample data during video cold start are simulated. Optionally, the number of user cold start samples, the number of video cold start samples, and the number of non-cold start samples in the sample data can be preset, and then training is performed according to the preset sample data.

[0120] Specifically, when training the vector reconstruction model, the input of sample data during user cold start can be simulated. This can be achieved by training the vector reconstruction model multiple times based on the sample data. During each training iteration, at least one first sample user is randomly selected from the sample users. The embedding vector of each first sample user is set to a zero vector and concatenated with its feature vector before being input into the vector reconstruction model. That is, the objective function is: L=||(e u e v T -f U (0||a u )f V (e v ||a v ) T )||2. Simulating the input of sample data during video cold start when training the vector reconstruction model can be achieved by: training the vector reconstruction model multiple times based on the sample data. During each training iteration, at least one first sample video is randomly selected from the sample videos. The embedding vector of each first sample video is set to a zero vector and concatenated with its feature vector before being input into the vector reconstruction model. That is, the objective function is: L=||(e u e v T -f U (e u ||a u )f V (0||a v ) T )||2. Following the above process, train the vector reconstruction model multiple times until the objective function is minimized. The vector reconstruction model corresponding to the minimum objective function value is the training result, and f can be obtained. U and f V .

[0121] Through the above process, the vector reconstruction model f is trained. U and f V .

[0122] Figure 4 This is a flowchart illustrating a video recommendation method provided in an embodiment of this application. This video recommendation method can be executed by a video recommendation device; this embodiment uses the video recommendation device as a server as an example for explanation. Figure 4As shown, the method in this embodiment may include:

[0123] S301. After receiving the video acquisition request, obtain the candidate video list.

[0124] For details of the process, please refer to the description in S101, which will not be repeated here.

[0125] S302. Obtain the user's reconstructed embedding vector.

[0126] Specifically, if the user is a new user, the 0 vector is determined as the user's embedding vector, and the user's feature vector is determined based on the user's feature information. The 0 vector and the user's feature vector are concatenated and input into the vector reconstruction model, which outputs the user's reconstructed embedding vector. For example, the pre-trained vector reconstruction model includes f U and f V The user's feature vector is a u The user's reconstructed embedding vector e u '=f U (0||a u The acquisition of the user's feature vector can be found in the description of the specific implementation of S1021, which will not be repeated here.

[0127] If the user is a historical user, the user's embedding vector is obtained based on the user's interaction network with historical videos. This embedding vector is then concatenated with the user's feature vector and input into the vector reconstruction model, which outputs the user's reconstructed embedding vector. For example, the pre-trained vector reconstruction model includes f U and f V The obtained user embedding vector is e u The user's feature vector is a u Then the user's reconstructed embedding vector e u '= f U (e u ||a u ).

[0128] S303. Obtain the reconstructed embedding vector for each candidate video.

[0129] Specifically, for each candidate video, if the candidate video is new, the 0 vector is determined as the embedding vector of the candidate video, and the feature vector of the candidate video is determined based on the feature information of the candidate video. The 0 vector and the feature vector of the candidate video are concatenated and input into the vector reconstruction model, and the reconstructed embedding vector of the candidate video is output. For example, the pre-trained vector reconstruction model includes f U and f V The feature vector of the candidate video is a v The reconstructed embedding vector e of the candidate video v '=f v (0||av The acquisition of the feature vectors of the candidate videos can be found in the description of the specific implementation of S1022, which will not be repeated here.

[0130] If the candidate video is a historical video, its embedding vector is obtained based on the historical user interaction network. This embedding vector is then concatenated with the candidate video's feature vector and input into the vector reconstruction model, which outputs the reconstructed embedding vector of the candidate video. For example, the pre-trained vector reconstruction model includes f... U and f V The embedded vector of the obtained candidate video is e v The feature vector of the candidate video is a v Then the user's reconstructed embedding vector e v '=f V (e v ||a v ).

[0131] S304. Based on the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video, determine the interaction probability between the user and each candidate video.

[0132] Taking the interaction probability between a user and a candidate video as the cosine distance between the user's reconstructed embedding vector and the candidate video's reconstructed embedding vector as an example, the user's reconstructed embedding vector e is calculated respectively. u 'and the reconstructed embedding vector e of each candidate video u 'cosine distance This gives the probability of user interaction with each candidate video {ρ}. uv ,v=1,2,...,N}, where N represents the total number of candidate videos.

[0133] S305. Based on the interaction probability between the user and each candidate video in the candidate video list, determine the target video to recommend to the user from the candidate video list.

[0134] Optionally, a preset number of candidate videos can be selected as target videos from the candidate video list in descending order of interaction probability.

[0135] The video recommendation method provided in this embodiment can accurately recommend videos to users during cold starts, improving the video recommendation effect during cold starts. It can also accurately recommend videos to users during non-cold starts. Moreover, since the user's reconstructed embedding vector can uniquely represent the user's interaction behavior information with the video and the user's feature information, and the video's reconstructed embedding vector can uniquely represent the video's interaction behavior information with the user and the video's feature information, it can save storage space and improve recommendation efficiency.

[0136] The following are embodiments of the apparatus described in this application, which can be used to execute the method embodiments described above. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments described above.

[0137] Figure 5 This is a schematic diagram of the structure of a video recommendation device provided in an embodiment of this application, as shown below. Figure 5 As shown, the apparatus of this embodiment may include: a first acquisition module 11, a second acquisition module 12, and a determination module 13, wherein,

[0138] The first acquisition module 11 is used to acquire a list of candidate videos after receiving a video acquisition request;

[0139] The second acquisition module 12 is used to acquire the interaction probability between the user and each candidate video in the candidate video list. The interaction probability is determined based on the user's reconstructed embedding vector and the candidate video's reconstructed embedding vector. The user's reconstructed embedding vector is determined based on the user's embedding vector, the user's feature vector, and a pre-trained vector reconstruction model. The candidate video's reconstructed embedding vector is determined based on the candidate video's embedding vector, the candidate video's feature vector, and the vector reconstruction model. The input of sample data during cold start is simulated when training the vector reconstruction model.

[0140] The determining module 13 is used to determine the target video to recommend to the user from the candidate video list based on the interaction probability between the user and each candidate video in the candidate video list.

[0141] Optionally, the second acquisition module 12 is used to acquire the user's reconstructed embedding vector;

[0142] Obtain the reconstructed embedding vector for each candidate video;

[0143] The interaction probability between the user and each candidate video is determined based on the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video.

[0144] Optionally, the second acquisition module 12 is specifically used to: if the user is a new user, determine the 0 vector as the embedding vector of the user, and determine the feature vector of the user based on the feature information of the user;

[0145] The zero vector is concatenated with the user's feature vector and then input into the vector reconstruction model to output the user's reconstructed embedding vector.

[0146] If the user is a historical user, obtain the user's embedding vector based on the interaction network between the user and historical videos;

[0147] The user's embedding vector and the user's feature vector are concatenated and then input into the vector reconstruction model, which outputs the user's reconstructed embedding vector.

[0148] Optionally, the second acquisition module 12 is specifically used to: if the candidate video is a new video, determine the 0 vector as the embedding vector of the candidate video, and determine the feature vector of the candidate video according to the feature information of the candidate video;

[0149] The zero vector is concatenated with the feature vector of the candidate video and then input into the vector reconstruction model to output the reconstructed embedding vector of the candidate video.

[0150] If the candidate video is a historical video, the embedding vector of the candidate video is obtained based on the interaction network between historical users and the candidate video;

[0151] The embedding vector of the candidate video is concatenated with the feature vector of the candidate video and then input into the vector reconstruction model to output the reconstructed embedding vector of the candidate video.

[0152] Optionally, the second acquisition module 12 is specifically used to: calculate the cosine distance, inner product, or Euclidean distance between the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video;

[0153] The cosine distance, inner product, or Euclidean distance between the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video is determined as the interaction probability of each candidate video.

[0154] Optionally, the vector reconstruction model is trained on multiple sample data, each sample data including the embedding vector of the sample user, the feature vector of the sample user, the embedding vector of the sample video, and the feature vector of the sample video;

[0155] When training the vector reconstruction model, the inputs to the vector reconstruction model are the embedding vector of the sample user, the feature vector of the sample user, the embedding vector of the sample video, and the feature vector of the sample video, and the outputs of the vector reconstruction model are the reconstructed embedding vector of the sample user and the reconstructed embedding vector of the sample video.

[0156] When training the vector reconstruction model, the input of sample data during cold start is simulated to minimize the value of the objective function. The vector reconstruction model is the vector reconstruction model corresponding to the minimum value of the objective function. The objective function represents the difference between the first interaction probability between the sample user and the sample video and the second interaction probability between the sample user and the sample video. The first interaction probability is determined based on the embedding vector of the sample user and the embedding vector of the sample video, and the second interaction probability is determined based on the reconstructed embedding vector of the sample user and the reconstructed embedding vector of the sample video.

[0157] Optionally, the process of simulating the input of sample data during cold start when training the vector reconstruction model includes:

[0158] The vector reconstruction model is trained by simulating the input of sample data during user cold start and the input of sample data during video cold start.

[0159] Optionally, the process of simulating the input of sample data during a user's cold start when training the vector reconstruction model includes:

[0160] The vector reconstruction model is trained multiple times based on the sample data. During each training of the vector reconstruction model, at least one first sample user is randomly selected from the sample users. The embedding vector of each first sample user is set to a 0 vector and concatenated with the feature vector of each first sample user before being input into the vector reconstruction model.

[0161] The process of simulating the input of sample data during video cold start when training the vector reconstruction model includes:

[0162] The vector reconstruction model is trained multiple times based on the sample data. During each training of the vector reconstruction model, at least one first sample video is randomly selected from the sample videos. The embedding vector of each first sample video is set to a 0 vector and concatenated with the feature vector of each first sample video before being input into the vector reconstruction model.

[0163] Optionally, the vector reconstruction model includes a user-side multilayer perceptron and a video-side multilayer perceptron. The input of the user-side multilayer perceptron is a first vector obtained by concatenating the embedding vector of the sample user with the feature vector of the sample user, and the output of the user-side multilayer perceptron is the reconstructed embedding vector of the sample user. The input of the video-side multilayer perceptron is a second vector obtained by concatenating the embedding vector of the sample video with the feature vector of the sample video, and the output of the video-side multilayer perceptron is the reconstructed embedding vector of the sample video.

[0164] Optionally, the objective function L is:

[0165] L=||(e u e v T -f U (e u ||a u )f V (e v ||a v ) T )||2

[0166] Among them, e u Let e ​​be the embedding vector of the sample user. v Let a be the embedding vector of the sample video. u Let a be the feature vector of the sample user. v Let f be the feature vector of the sample video. U For the user-side multilayer sensor, f V For the video-side multilayer perceptron, e u ||a u Let e ​​be the first vector. v ||a v Let f be the second vector, ||X||2 be the L2 norm of X, and f U (e u ||a u f is the reconstructed embedding vector of the sample user. V (e v ||a v ) is the reconstructed embedding vector of the sample video.

[0167] Optionally, the determining module 13 is used to select a preset number of candidate videos from the candidate video list as the target videos in descending order of the interaction probability.

[0168] The apparatus provided in this application embodiment can execute the above method embodiment. Its specific implementation principle and technical effect can be found in the above method embodiment, and will not be repeated here.

[0169] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, a processing module can be a separate processing element, or it can be integrated into a chip within the device. Alternatively, it can be stored as program code in the device's memory, and its functions can be called and executed by a processing element. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0170] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together to implement a system-on-a-chip (SOC).

[0171] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0172] Figure 6 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application, such as... Figure 6 As shown, the terminal device in this embodiment may include a processor 21 and a memory 22.

[0173] The memory 22 is used to store the executable instructions of the processor 21.

[0174] The processor 21 is configured to execute the video recommendation method in the above method embodiments by executing executable instructions.

[0175] Alternatively, the memory 22 can be either standalone or integrated with the processor 21.

[0176] When the memory 22 is a device independent of the processor 21, the terminal device in this embodiment may further include:

[0177] Bus 23 is used to connect memory 22 and processor 21.

[0178] Optionally, the terminal device in this embodiment may further include a communication interface 24, which can be connected to the processor 21 via a bus 23.

[0179] This application also provides a computer-readable storage medium storing computer-executable instructions that, when run on a computer, cause the computer to perform the video recommendation method as described in the above embodiments.

[0180] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the video recommendation method as described in the above embodiments.

[0181] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0182] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A video recommendation method, characterized in that, include: Upon receiving a video retrieval request, obtain a list of candidate videos; The interaction probability between a user and each candidate video in the candidate video list is obtained. The interaction probability is determined based on the reconstructed embedding vector of the user and the reconstructed embedding vector of the candidate video. The reconstructed embedding vector of the user is determined based on the user's embedding vector, the user's feature vector, and a pre-trained vector reconstruction model. The reconstructed embedding vector of the candidate video is determined based on the candidate video's embedding vector, the candidate video's feature vector, and the vector reconstruction model. The input of sample data during cold start is simulated when training the vector reconstruction model. Based on the interaction probability between the user and each candidate video in the candidate video list, a target video to be recommended to the user is determined from the candidate video list; When training the vector reconstruction model, the inputs to the vector reconstruction model are the embedding vector of the sample user, the feature vector of the sample user, the embedding vector of the sample video, and the feature vector of the sample video, and the outputs of the vector reconstruction model are the reconstructed embedding vector of the sample user and the reconstructed embedding vector of the sample video. The vector reconstruction model is trained by simulating the input of sample data during a cold start, so as to minimize the value of the objective function. The vector reconstruction model is the vector reconstruction model corresponding to the minimum value of the objective function.

2. The method according to claim 1, characterized in that, The step of obtaining the interaction probability between the user and each candidate video in the candidate video list includes: Obtain the reconstructed embedding vector of the user; Obtain the reconstructed embedding vector for each candidate video; The interaction probability between the user and each candidate video is determined based on the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video.

3. The method according to claim 2, characterized in that, The step of obtaining the user's reconstructed embedding vector includes: If the user is a new user, then the 0 vector is determined as the embedding vector of the user, and the feature vector of the user is determined according to the feature information of the user; The zero vector is concatenated with the user's feature vector and then input into the vector reconstruction model to output the user's reconstructed embedding vector. If the user is a historical user, obtain the user's embedding vector based on the interaction network between the user and historical videos; The user's embedding vector and the user's feature vector are concatenated and then input into the vector reconstruction model, which outputs the user's reconstructed embedding vector.

4. The method according to claim 2, characterized in that, The step of obtaining the reconstructed embedding vector for each candidate video includes: If the candidate video is a new video, then the 0 vector is determined as the embedding vector of the candidate video, and the feature vector of the candidate video is determined according to the feature information of the candidate video. The zero vector is concatenated with the feature vector of the candidate video and then input into the vector reconstruction model to output the reconstructed embedding vector of the candidate video. If the candidate video is a historical video, the embedding vector of the candidate video is obtained based on the interaction network between historical users and the candidate video; The embedding vector of the candidate video is concatenated with the feature vector of the candidate video and then input into the vector reconstruction model to output the reconstructed embedding vector of the candidate video.

5. The method according to claim 2, characterized in that, The step of determining the interaction probability between the user and each candidate video based on the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video includes: Calculate the cosine distance, inner product, or Euclidean distance between the reconstructed embedding vector of the user and the reconstructed embedding vector of each candidate video; The cosine distance, inner product, or Euclidean distance between the user's reconstructed embedding vector and the reconstructed embedding vector of each candidate video is determined as the interaction probability of each candidate video.

6. The method according to any one of claims 1-5, characterized in that, The objective function represents the difference between the first interaction probability of the sample user and the sample video and the second interaction probability of the sample user and the sample video. The first interaction probability is determined based on the embedding vector of the sample user and the embedding vector of the sample video, and the second interaction probability is determined based on the reconstructed embedding vector of the sample user and the reconstructed embedding vector of the sample video.

7. The method according to claim 6, characterized in that, The process of simulating the input of sample data during cold start when training the vector reconstruction model includes: The vector reconstruction model is trained by simulating the input of sample data during user cold start and the input of sample data during video cold start.

8. The method according to claim 7, characterized in that, The process of simulating the input of sample data during a user's cold start when training the vector reconstruction model includes: The vector reconstruction model is trained multiple times based on the sample data. During each training of the vector reconstruction model, at least one first sample user is randomly selected from the sample users. The embedding vector of each first sample user is set to a 0 vector and concatenated with the feature vector of each first sample user before being input into the vector reconstruction model. The process of simulating the input of sample data during video cold start when training the vector reconstruction model includes: The vector reconstruction model is trained multiple times based on the sample data. During each training of the vector reconstruction model, at least one first sample video is randomly selected from the sample videos. The embedding vector of each first sample video is set to a 0 vector and concatenated with the feature vector of each first sample video before being input into the vector reconstruction model.

9. The method according to claim 6, characterized in that, The vector reconstruction model includes a user-side multilayer perceptron and a video-side multilayer perceptron. The input of the user-side multilayer perceptron is a first vector obtained by concatenating the embedding vector of the sample user with the feature vector of the sample user. The output of the user-side multilayer perceptron is the reconstructed embedding vector of the sample user. The input of the video-side multilayer perceptron is a second vector obtained by concatenating the embedding vector of the sample video with the feature vector of the sample video. The output of the video-side multilayer perceptron is the reconstructed embedding vector of the sample video.

10. The method according to claim 9, characterized in that, The objective function L is: L=||(and u And v T -f U (And u ||a u )f V (And v ||a v ) T )||2 Among them, e u Let e ​​be the embedding vector of the sample user. v Let a be the embedding vector of the sample video. u Let a be the feature vector of the sample user. v Let f be the feature vector of the sample video. U For the user-side multilayer sensor, f V For the video-side multilayer perceptron, e u ||a u Let e ​​be the first vector. v ||a v Let f be the second vector, ||X||2 be the L2 norm of X, and f U (e u ||a u f is the reconstructed embedding vector of the sample user. V (e v ||a v ) is the reconstructed embedding vector of the sample video.

11. The method according to claim 1, characterized in that, The step of determining the target video to recommend to the user from the candidate video list based on the interaction probability between the user and each candidate video in the candidate video list includes: According to the order of interaction probability from largest to smallest, a preset number of candidate videos are selected from the candidate video list as the target videos.

12. A video recommendation device, characterized in that, include: The first acquisition module is used to acquire a list of candidate videos after receiving a video acquisition request; The second acquisition module is used to acquire the interaction probability between the user and each candidate video in the candidate video list. The interaction probability is determined based on the reconstructed embedding vector of the user and the reconstructed embedding vector of the candidate video. The reconstructed embedding vector of the user is determined based on the user's embedding vector, the user's feature vector, and a pre-trained vector reconstruction model. The reconstructed embedding vector of the candidate video is determined based on the candidate video's embedding vector, the candidate video's feature vector, and the vector reconstruction model. The input of sample data during cold start is simulated when training the vector reconstruction model. The determining module is used to determine the target video to recommend to the user from the candidate video list based on the interaction probability between the user and each candidate video in the candidate video list; When training the vector reconstruction model, the inputs to the vector reconstruction model are the embedding vector of the sample user, the feature vector of the sample user, the embedding vector of the sample video, and the feature vector of the sample video, and the outputs of the vector reconstruction model are the reconstructed embedding vector of the sample user and the reconstructed embedding vector of the sample video. The vector reconstruction model is trained by simulating the input of sample data during a cold start, so as to minimize the value of the objective function. The vector reconstruction model is the vector reconstruction model corresponding to the minimum value of the objective function.

13. A terminal device, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to invoke and run the computer program stored in the memory to perform the method of any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 1 to 11.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.