Video recommendation method and device, equipment and storage medium
By constructing a multi-dimensional tagging system and multi-sequence behavior, combined with a video recommendation model, the problem of inaccurate user interest representation in existing technologies is solved, achieving higher recommendation accuracy and diversity, and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-07
AI Technical Summary
Existing video recommendation models rely on single-level labels as sparse features, lacking a deep and refined semantic expression of video content. This results in inaccurate representation of user interests, weak relevance of recommendation results, insufficient ranking accuracy, and negatively impacts user experience.
By constructing a multi-dimensional tagging system and combining it with multi-sequence behaviors, a multi-dimensional tagging behavior sequence is generated through the establishment of a multi-dimensional tagging system and multi-granular user interest expression. This sequence is then used as input for the video recommendation model, thereby improving the accuracy and diversity of video recommendations.
By combining a multi-dimensional labeling system with multi-sequence behavior, the accuracy of user interest representation is significantly improved, the relevance and ranking accuracy of recommendation results are enhanced, and the user experience is improved.
Smart Images

Figure CN121808100A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular, to a video recommendation method and device, equipment and storage medium. BACKGROUND
[0002] In the video recommendation technology, the existing video recommendation model usually relies on a single level of label as a sparse feature input, lacks deep and fine semantic expression of video content, and thus leads to inaccurate representation of user interest, weak relevance of recommendation results, and insufficient sorting accuracy, which seriously affects user experience. SUMMARY
[0003] To solve the above technical problems, the present application provides a video recommendation method, device, equipment and storage medium.
[0004] In a first aspect, the present application provides a video recommendation method, comprising: obtaining historical behavior data of a target user, wherein the historical behavior data refers to a plurality of interaction data generated by the target user in the process of watching historical videos; establishing a multi-dimensional label system for structurally describing video content of the historical videos, wherein each dimension label in the multi-dimensional label system is used to represent an independent semantic classification; determining the association relationship between the historical behavior data and each dimension label in the multi-dimensional label system, and integrating the association relationship to generate a multi-dimensional label behavior sequence for representing video interest preferences of the target user; taking the multi-dimensional label behavior sequence and multi-dimensional label features of at least one candidate video generated in advance as inputs of a pre-trained video recommendation model, and outputting at least one target video to be recommended to the target user.
[0005] In a second aspect, the present application provides a video recommendation device, comprising: an obtaining unit configured to obtain historical behavior data of a target user, wherein the historical behavior data refers to a plurality of interaction data generated by the target user in the process of watching historical videos; an establishing unit configured to establish a multi-dimensional label system for structurally describing video content of the historical videos, wherein each dimension label in the multi-dimensional label system is used to represent an independent semantic classification; a generating unit configured to determine the association relationship between the historical behavior data and each dimension label in the multi-dimensional label system, and integrate the association relationship to generate a multi-dimensional label behavior sequence for representing video interest preferences of the target user; An output unit is configured to input the multi-dimensional label behavior sequence and the multi-dimensional label features of the at least one candidate video pre-generated as inputs of the pre-trained video recommendation model, and output at least one target video to be recommended to the target user.
[0006] In a third aspect, the embodiments of the present disclosure provide an electronic device, comprising: a memory; a processor; and a computer program; The computer program is stored in the memory and configured to be executed by the processor to implement the method of the first aspect.
[0007] In a fourth aspect, the embodiments of the present disclosure provide a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method of the first aspect are implemented.
[0008] The video recommendation method provided by the present disclosure comprises: obtaining historical behavior data of a target user, wherein the historical behavior data refers to a plurality of interaction data generated by the target user during the process of watching historical videos; establishing a multi-dimensional label system for structurally describing the video content of the historical videos, wherein each dimension label in the multi-dimensional label system is used to represent an independent semantic classification; determining the association between the historical behavior data and each dimension label in the multi-dimensional label system, and generating a multi-dimensional label behavior sequence for representing the video interest preferences of the target user; and outputting a target video to be recommended to the user by a video recommendation model according to the multi-dimensional label behavior sequence and the multi-dimensional label features of the candidate video. The method provided by the present disclosure improves the accuracy of user interest representation, strengthens the relevance of the recommendation results, and further improves the sorting precision and user experience. BRIEF DESCRIPTION OF DRAWINGS
[0009] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0011] Figure 1 A flowchart of the video recommendation method provided by the embodiments of the present disclosure is shown in the figure; Figure 2 A detailed flowchart of S104 in the video recommendation method is shown in the figure; Figure 1 A detailed flowchart of S104 in the video recommendation method is shown in the figure; Figure 3 A structural schematic diagram of a video recommendation device provided by an embodiment of the present disclosure is shown in the following figure. Figure 4 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown in the following figure. DETAILED DESCRIPTION
[0012] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0013] In the following description, many specific details are set forth in order to provide a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments described in the specification are only some of the embodiments of the present disclosure, not all the embodiments.
[0014] Specifically, current fine ranking models mostly rely on sparse features of single-level labels, and are difficult to capture rich semantic information of videos, and also fail to fully model user behavior sequences, resulting in poor recommendation effect of recommended videos, serious cold start problem, and difficulty in meeting the recommendation accuracy and diversity requirements.
[0015] To solve the above technical problems, an embodiment of the present disclosure provides a video recommendation method, which combines multi-dimensional labels and multi-sequence behaviors, constructs hierarchical video semantic features and multi-granularity user interest expressions, and realizes accurate matching in a fine ranking model to improve the accuracy and diversity of video recommendation. One or more of the following embodiments are described in detail.
[0016] The video recommendation method provided by the embodiment of the present disclosure can be applied to a video recommendation scene, such as a short video recommendation scene. The method can be executed by a video recommendation device, which can be implemented by software and / or hardware, and the device can be integrated in an electronic device. The electronic device can include but is not limited to mobile terminals such as smartphones, notebook computers, digital broadcast receivers, personal digital assistants (PDA), tablet computers (Tablet PC), PMP (portable multimedia player), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, and fixed terminals such as digital televisions, desktop computers, and smart home devices.
[0017] Figure 1 A flowchart of the video recommendation method provided by an embodiment of the present disclosure is shown in the following figure, which specifically includes the following steps as shown in the figure: Figure 1 S101, obtaining historical behavior data of a target user.
[0018] Historical behavioral data refers to various interactive data generated by target users while watching historical videos.
[0019] Understandably, target users refer to users of the videos to be recommended. Historical behavioral data refers to records of various types of interactive behaviors generated by users while watching videos, including but not limited to video exposure, clicks to play, viewing duration, playback progress, likes, favorites, comments, sharing, and following the author. Each interactive behavior record is associated with video information, behavior type, behavior time, and behavior intensity.
[0020] S102. Establish a multi-dimensional tagging system for the structured description of historical video content.
[0021] In this multi-dimensional labeling system, each dimension label is used to represent an independent semantic classification.
[0022] Understandably, based on the above S101, videos watched that interact with the target user are recorded as historical videos. Multi-dimensional tags are established to structurally describe the content of these historical videos. When a user has watched a large number of videos, multi-dimensional tags can be established for at least a portion of the watched videos, or the user can select certain watched videos and then establish multi-dimensional tags for those videos. Understandably, multi-dimensional tags can be established for each historical video. These multi-dimensional tags include a set of tags, specifically using a structured multi-level tagging system organized hierarchically by an intelligent model.
[0023] In one embodiment, a multi-dimensional tagging system is a tagging system with a tree-like or graph-like hierarchical structure. It can be understood as a multi-level tagging system used to represent the semantic information of video content from coarse-grained to fine-grained, and from abstract to concrete. For example, hierarchical relationships such as technology, artificial intelligence, deep learning, and neural networks can be used to achieve layered descriptions of video themes, categories, and attributes. This allows multi-dimensional tags to better support interest generalization, semantic reasoning, and accurate recommendations, improving the depth of content understanding and the ranking accuracy of the recommendation system.
[0024] In another embodiment, the multi-dimensional label system is a label system without hierarchy, which specifically refers to labeling video content from multiple perspectives to comprehensively capture and describe various attributes of the video. Each dimension represents a specific type of feature or perspective, which is used to provide a richer and more detailed understanding of the video content. Without hierarchy means that all labels are at the same level, and there may be intersections between labels but no hierarchical relationship, such as technology, education, entertainment, sports, news, etc. are all at the same level. A video is a sports news video, which conforms to both the sports label and the news label, so the video can be labeled with both the sports and news labels.
[0025] Optionally, the multi-dimensional label system is composed of a label set generated based on multiple mutually independent semantic classification dimensions, wherein the semantic classification dimensions include at least two of video content, video theme, video style, video audience, video production mode, and video emotional tone.
[0026] In one embodiment, the multi-dimensional label includes a first-level label, a second-level label, and a third-level label. The first-level label is determined based on video content, the second-level label is determined based on video theme, and the third-level label is determined based on video style. The first-level label, the second-level label, and the third-level label belong to the same level and have no hierarchical relationship. The first-level label (which can also be understood as the first type of label or the first dimension label) is determined based on video content, for example, the first-level label includes labels such as plot, comedy, news, variety, and game. The second-level label is determined based on video theme and is used to describe the basic theme or field of the video, for example, the second-level label includes labels such as urban, costume, action, etc. The third-level label is determined based on video style and is used to describe the visual or narrative style characteristics of the video, for example, the third-level label includes labels such as warm, stimulating, and relaxed. For example, video A is a commentary video of a certain criminal investigation TV series, the first-level label of the video A is "plot", the second-level label is "action", and the third-level label is "stimulating".
[0027] In another embodiment, the multi-dimensional label further includes a fourth-level label, a fifth-level label, and a sixth-level label. The fourth-level label is determined based on entities and is used to identify specific characters, locations, brands, or other entities involved in the video, for example, labels such as A area and electronic device. The fifth-level label is determined based on events and is used to describe the main activities or event types occurring in the video, for example, labels such as speech, cooking, travel, and concert. The sixth-level label is determined based on video scenes and is used to represent the physical environment or background in which the video is shot, for example, labels such as indoor, outdoor, office, and beach. For example, the sixth-level label of the above-mentioned video A is "outdoor". It can be understood that the number of levels included in the multi-dimensional label is not limited.
[0028] S103, determine the association relationship between the historical behavior data and each dimension label in the multi-dimension label system, and integrate the association relationship to generate a multi-dimension label behavior sequence for representing the video interest preference of the target user.
[0029] Understandably, on the basis of S102, for the historical behavior record of each watched video in the historical behavior data, the associated multi-dimension label is determined, and the multi-dimension label behavior sequence is generated according to the association relationship with the multi-dimension label. Specifically, the label path of the historical video on different semantic levels is extracted, and each historical behavior is associated with the corresponding label path in time sequence to form an ordered sequence composed of video behaviors and multi-level semantic labels, and each level label corresponds to a behavior sequence, and then a multi-dimension label sequence for representing the video interest preference of the target user is composed.
[0030] In an embodiment, the user has watched 10 videos, and the labels involved in the 10 videos include love drama, costume drama, urban drama and suspense drama. The first 9 videos in the 10 videos are labeled with love drama, among which the first 5 videos are also labeled with costume drama (i.e. costume love drama), the 6th and 7th videos are labeled with urban drama (i.e. urban love drama), the 8th and 9th videos are labeled with suspense drama (i.e. love suspense drama), and the 10th video is labeled with suspense drama and urban drama (i.e. urban suspense drama). In this case, the love drama level label is associated with the historical behavior data of the first 9 videos, the costume drama level label is associated with the historical behavior data of the first 5 videos, the suspense drama level label is associated with the historical behavior data of the 8th, 9th and 10th videos, the urban drama level label is associated with the historical behavior data of the 6th, 7th and 10th videos, and the multi-level label behavior sequence is obtained accordingly. On the basis of the example of the user watching 10 videos, the user may be more interested in costume love drama.
[0031] Optionally, the association relationship between the historical behavior data and each dimension label in the multi-dimension label system is determined, and the multi-dimension label behavior sequence for representing the video interest preference of the target user is integrated and generated. Specifically, it can be realized by the following steps: The historical behavior data is constructed into a plurality of behavior sequences according to a plurality of behavior types, wherein the plurality of behavior sequences are different behavior sequences extracted from different stages of the interaction between the target user and the historical video; the plurality of behavior sequences are mapped to the multi-dimension label system to establish the association relationship between different behaviors and dimension labels; and the association relationship is integrated to generate a multi-dimension label behavior sequence for representing the video interest preference of the target user, wherein each dimension label corresponds to a behavior sequence.
[0032] It can be understood that, for the collected historical behavior data, a plurality of behavior sequences are constructed according to behavior types, wherein the plurality of behavior sequences are different behavior sequences extracted from different stages of the target user interacting with the historical video, for example, behavior data belonging to the same behavior type is taken as a behavior sequence. The behavior types include at least two behaviors of exposure behavior, viewing behavior and positive feedback behavior, and the exposure, viewing and positive feedback are all user interaction behaviors with the video, wherein the behavior sequence corresponding to the exposure behavior is recorded as a browsing sequence or an exposure behavior sequence, the behavior sequence corresponding to the viewing behavior is recorded as a complete playback sequence or a viewing behavior sequence, and the behavior sequence corresponding to the positive feedback behavior is recorded as a like and / or collection sequence or a positive feedback behavior sequence. Then, each behavior sequence is mapped to a corresponding multi-dimensional label system to construct a multi-dimensional label behavior sequence. For example, a user watches a science fiction movie, the movie has a "science fiction" theme label, a "movie" type label, an "actor" actor label, etc. Thus, the association between user behavior and multiple dimension labels is established, and then the different dimension labels are sorted into sequences respectively, that is, each dimension label corresponds to a behavior sequence, and these sequences collectively represent the user's interest preferences. Subsequently, according to the multi-dimensional label behavior sequence, it can be directly seen that "the user has 2 viewing behaviors in science fiction, 1 in comedy, 2 in movies, and 1 in variety shows", and it can be analyzed that the user strengthens the interest in science fiction and the preference for movies is also stable, and the recommendation weight of science fiction movies can be increased.
[0033] The plurality of behavior types include at least two of the exposure behavior, the viewing behavior and the positive feedback behavior; and the plurality of behavior sequences include at least two of the exposure behavior sequence, the viewing behavior sequence and the positive feedback behavior sequence of the historical video; the exposure behavior sequence is used to record the number of times of displaying the historical video to the target user, the viewing behavior sequence is used to record the playing time of the historical video, and the positive feedback behavior sequence is used to record the trigger operation of the target user expressing the preference for the historical video.
[0034] It can be understood that the exposure behavior sequence refers to a list of videos displayed (recommended) by the system to the user, whether the user clicks or watches or not, which reflects the content that the system has tried to guide the user to pay attention to. The viewing behavior sequence refers to a list of videos actually clicked and played by the user, which usually requires a certain playing progress (such as playing time > 3 seconds or playing ratio > 10%), representing the content actively watched by the user, which can better reflect the real interest preference of the user, and the viewing behavior sequence includes at least one of the viewing start time, the viewing time, the fast forward / fast backward times, the playing mode and the playing completion rate. The positive feedback behavior sequence refers to a list of behaviors of the user expressing explicit love or recognition for the video content, which is a strong signal behavior and can better reflect the "preference intensity" of the user than the viewing behavior, and can more accurately determine the video of interest to the user.
[0035] S104, taking the multi-dimensional label behavior sequence and the multi-dimensional label feature of the at least one candidate video pre-generated as inputs of the pre-trained video recommendation model, and outputting at least one target video to be recommended to the target user.
[0036] It can be understood that, on the basis of S103, the multi-dimensional label feature of the at least one candidate video is obtained, the multi-dimensional label feature is obtained by feature extraction on the multi-dimensional label system of the candidate video, and the multi-dimensional label system of the candidate video and the multi-dimensional label system of the historical video are determined based on the same basic label system. Taking the multi-dimensional label behavior sequence and the multi-dimensional label feature of the candidate video as inputs of the video recommendation model, the video recommendation model can be understood as a fine ranking model, which is a key component for fine ranking of the candidate video set screened in the recall stage. Mapping the historical behavior data to the multi-dimensional label system and constructing the corresponding behavior sequence can significantly improve the depth and accuracy of the understanding of user interest by the fine ranking model, and thus improve the relevance and ranking accuracy of the recommendation result. The way of determining the ranking of each candidate video in the fine ranking model is not described herein.
[0037] Optionally, the multi-dimensional label behavior sequence and the multi-dimensional label feature of the at least one candidate video pre-generated are taken as inputs of the pre-trained video recommendation model, and at least one target video to be recommended to the target user is output, which can be realized by the following steps: Obtaining video preference information and / or historical viewing information of the target user; using the pre-trained video recommendation model to determine at least one target video to be recommended to the target user according to the multi-dimensional label behavior sequence and the multi-dimensional label feature of the at least one candidate video pre-generated, and the video preference information and / or historical viewing information; wherein the video recommendation model takes the user behavior index of the target user as the optimization target in the training process, and the user behavior index includes the click rate and / or the complete play rate of the video.
[0038] It can be understood that before determining the target video to be recommended to the user based on the video recommendation model, the video preference information, user portrait information and / or historical viewing information of the user are obtained, wherein the video preference information refers to the user's tendency expression information for specific video content, for example, the user prefers more cooking and / or eating videos. The user portrait information describes the structured information of the static or semi-static attributes of the user, which is used to depict the basic characteristics and long-term attributes of the user. The user portrait information includes basic attributes (age, gender, region, etc.) and behavior portrait information, for example, animation video clips are recommended for children. The historical viewing information refers to the list of videos actually viewed by the user in the past period of time, for example, the user has watched a large number of variety shows in the recent period of time, and the user may prefer variety show videos in this period of time. Subsequently, the multi-dimensional label behavior sequence and the multi-dimensional label features of the candidate video, and the video preference information, user portrait information and / or historical viewing information are used as inputs of the video recommendation model to determine the target video to be recommended to the user. The specific implementation steps of the video recommendation model are not described in detail.
[0039] It can be understood that the fine ranking model can be trained by using a deep neural network (such as Deep Neural Network, DNN), and the optimization target is a user behavior index such as Click-Through Rate (CTR) or Conversion Rate (CVR), wherein the Click-Through Rate refers to the ratio of the number of times that the user clicks the video content to the number of times that the video content is displayed (exposure times), and the Conversion Rate refers to the proportion of users who complete video playback.
[0040] It can be understood that the interest vector, the label features of the candidate video and other auxiliary features are collectively used as inputs and sent to the fine ranking model for end-to-end training. In the training process, the user behavior index of the target user is used as the optimization target, and the user behavior index includes the Click-Through Rate and / or the Conversion Rate of the video, wherein the optimization target refers to the improvement direction in the model training process, so that the user is more willing to click (higher Click-Through Rate) or more willing to watch (higher Conversion Rate) the video recommended by the model. In this way, the fine ranking model can more accurately capture the matching relationship between the user interest and the candidate content by jointly optimizing the interest modeling and the ranking target, improve the personalization degree and the ranking accuracy of the recommendation system, and enhance the prediction ability of the user behavior, thereby effectively improving the key business indexes such as the Click-Through Rate and the Conversion Rate.
[0041] The embodiments of the present disclosure aim at the problems of shallow semantic expression of video content and insufficient user interest modeling in related video recommendation, and propose a fine-grained recommendation method based on a multi-dimensional label system. By constructing a multi-dimensional label system, the video content is represented in multiple granularity / dimension semantics from macro to micro, and the semantic information contained in the video is fully represented. Further, the historical behavior data of the user for the historical video is obtained, and the multi-sequence interest modeling technology is used to capture the behavior sequence preference of the user under each level of label, and generate a multi-level fine-grained interest vector. In addition, in the fine arrangement stage, the hierarchical attention mechanism is introduced to realize the semantic matching and alignment between the user interest vector and the multi-dimensional label features of the candidate video. This method significantly enhances the depth of the model's understanding of the video content and the user's interest expression ability, effectively improves the relevance, ranking accuracy and diversity of the recommendation results, and especially shows stronger modeling ability and system robustness in the long-tail video and user interest evolution scenarios.
[0042] On the basis of the above embodiments, Figure 2 For Figure 1 The fine process schematic diagram of S104 in the video recommendation method shown is optional, and the multi-dimensional label behavior sequence and the multi-dimensional label features of the at least one candidate video are taken as the input of the pre-trained video recommendation model, and the at least one target video to be recommended to the target user is output, which specifically includes the following steps as shown in Figure 2 The fine process schematic diagram of S104 in the video recommendation method shown is optional, and the multi-dimensional label behavior sequence and the multi-dimensional label features of the at least one candidate video are taken as the input of the pre-trained video recommendation model, and the at least one target video to be recommended to the target user is output, which specifically includes the following steps as shown in S1041, calculating the similarity of the multi-dimensional label features of the at least one candidate video and the multi-dimensional label behavior sequence to obtain multi-dimensional label interest features.
[0043] The multi-dimensional label interest features are used to represent the matching degree between the candidate video and the video interested by the target user.
[0044] It can be understood that the interaction features of different dimensional label behavior sequences (such as first-level labels, second-level labels, third-level labels, etc.) and user interests are extracted respectively, and are fused through a unified attention mechanism model to construct a unified multi-dimensional label interest representation (i.e., multi-dimensional label interest features). Specifically, the label features of each level and the user historical behavior sequence can be input into a deep interest network (DIN) or a Transformer module to mine the interest intensity and preference pattern of the user for the level label.
[0045] S1042, taking the multi-dimensional label interest features and the multi-dimensional label features as the input of the pre-trained video recommendation model, and outputting at least one target video to be recommended to the target user.
[0046] It can be understood that, on the basis of S1041, the multi-dimensional label interest features extracted from the historical video, the historical behavior data, and the candidate video and the multi-dimensional label features of the candidate video are input into the pre-trained video recommendation model. The video recommendation model calculates the matching degree of each candidate video and the user interest, automatically scores and sorts, and finally selects one or more videos that best match the user interest as the recommendation result (i.e., the target video).
[0047] Optionally, the similarity between the multi-dimensional label features of the at least one candidate video and the multi-dimensional label behavior sequence is calculated to obtain the multi-dimensional label interest features, which can be achieved by the following steps: The multi-dimensional label behavior sequence is feature-encoded to obtain multi-dimensional label behavior features; the similarity between the dimensional label features and the dimensional label behavior features belonging to the same dimension is calculated through an attention mechanism to obtain label interest features of each dimension; wherein the multi-dimensional label features include the dimensional label features, and the multi-dimensional label behavior features include the dimensional label behavior features; the label interest features of each dimension are weighted and fused or spliced to obtain unified multi-dimensional label interest features.
[0048] It can be understood that the multi-dimensional label behavior sequence is feature-encoded to obtain multi-dimensional label behavior features, and the label ID can be mapped to a dense vector representation using an Embedding layer. Subsequently, the multi-dimensional label behavior features are input into a DIN or a Transformer module. The module dynamically calculates the interest weight of the multi-dimensional label features of the current candidate video through an attention mechanism, mines the time sequence dependence and hierarchical association of the user's interest, and obtains the label interest features. It can be understood that each level of label hierarchy corresponds to a DIN or a Transformer module. For each level of label hierarchy, the user interest representation vector (i.e., the label interest features) of each level of label hierarchy can be obtained through the corresponding DIN or Transformer module, that is, the correlation / matching degree / similarity between the first level label features of the candidate video and the first level label behavior features of the historical video belonging to the same level is calculated. For example, on the basis of the above embodiment, the first level label corresponds to the first label interest features, the second level label corresponds to the second label interest features, and the third level label corresponds to the third label interest features. Subsequently, the interest representation vectors of each level (i.e., the label interest features of each level) are weighted and fused or spliced to construct a unified multi-dimensional label interest representation (i.e., the multi-dimensional label interest features).
[0049] It can be understood that the unified multi-dimensional label feature and the unified multi-dimensional label interest feature of the candidate video can be further interactively matched, and a unified interest similarity score can be calculated, wherein the unified multi-dimensional label feature of the candidate video is obtained by weighting and fusing or splicing the label features of each dimension. That is, after calculating and weighting and fusing the unified label interest features of the candidate video and the historical video, that is, considering the similarity from the perspective of each level, the interest similarity between the unified label feature and the unified label interest feature of the candidate video can be calculated, that is, the similarity is further considered from the overall perspective.
[0050] Optionally, according to the multi-dimensional label interest feature and the multi-dimensional label feature, at least one target video is determined in at least one candidate video by using a pre-trained video recommendation model, and the determination can be implemented through the following steps: According to the multi-dimensional label interest feature and the multi-dimensional label feature of the at least one candidate video, and the video preference information and / or the historical viewing information, at least one target video is determined in the at least one candidate video by using a pre-trained video recommendation model.
[0051] It can be understood that after the multi-dimensional label interest feature is obtained by encoding the multi-dimensional label behavior sequence, other auxiliary features (such as user portrait information, video preference information and / or historical viewing information) can be obtained. Subsequently, the multi-dimensional label interest feature, the multi-dimensional label feature and the other auxiliary features are used as inputs of the video recommendation model. For specific implementation manners, refer to the above embodiments, which will not be described here.
[0052] The video recommendation method provided by the embodiments of the present disclosure can mine the interest preferences of a user on coarse-grained to fine-grained labels layer by layer by respectively independently modeling the label behavior sequences of the user on different semantic levels. At each level, the interest weight or long sequence dependency relationship related to the content of the current candidate video in the historical behavior of the user is dynamically captured by using an attention mechanism, and the interest evolution path is captured. Subsequently, the interest vectors generated by each level are integrated by weighting and fusing or splicing to form a unified multi-level interest representation vector. The method can more comprehensively and hierarchically depict the complex interest structure of the user, and effectively improve the understanding accuracy of the user's intention. Through the layered modeling, the confusion of the semantic levels by the traditional single interest vector is avoided, and the expression ability of the model to the interest generalization and focusing is enhanced.
[0053] Figure 3 A structural schematic diagram of the video recommendation device provided by the embodiments of the present disclosure is shown. The video recommendation device provided by the embodiments of the present disclosure can execute the processing flow provided by the video recommendation method embodiments, as shown in Figure 3 The video recommendation device 300 includes: The acquisition unit 301 is configured to acquire historical behavior data of a target user, wherein the historical behavior data refers to a plurality of interaction data generated by the target user in a process of watching a historical video. The establishment unit 302 is configured to establish a multi-dimensional label system for structurally describing video content of the historical video, wherein each dimension label in the multi-dimensional label system is used to represent an independent semantic classification. The generation unit 303 is configured to determine an association relationship between the historical behavior data and each dimension label in the multi-dimensional label system, and integrate the association relationship to generate a multi-dimensional label behavior sequence for representing a video interest preference of the target user. The output unit 304 is configured to take the multi-dimensional label behavior sequence and multi-dimensional label features of at least one candidate video as inputs of a pre-trained video recommendation model, and output at least one target video to be recommended to the target user.
[0054] Optionally, the output unit 304 is configured to: calculate a similarity between the multi-dimensional label features of the at least one candidate video and the multi-dimensional label behavior sequence, to obtain multi-dimensional label interest features, wherein the multi-dimensional label interest features are used to represent a matching degree between the candidate video and a video interested by the target user; and take the multi-dimensional label interest features and the multi-dimensional label features as inputs of the pre-trained video recommendation model, and output at least one target video to be recommended to the target user.
[0055] Optionally, the output unit 304 is configured to: perform feature coding on the multi-dimensional label behavior sequence to obtain multi-dimensional label behavior features; calculate a similarity between the dimension label features and the dimension label behavior features belonging to the same dimension through an attention mechanism, to obtain label interest features of each dimension; wherein the multi-dimensional label features include the dimension label features, and the multi-dimensional label behavior features include the dimension label behavior features; and perform weighted fusion processing or splicing processing on the label interest features of each dimension, to obtain unified multi-dimensional label interest features.
[0056] Optionally, the generation unit 303 is configured to: construct a plurality of behavior sequences from the historical behavior data according to a plurality of behavior types, wherein the plurality of behavior sequences are different behavior sequences extracted from different stages of interaction between the target user and the historical video; map the plurality of behavior sequences to the multi-dimensional label system, to establish an association relationship between different behaviors and dimension labels; integrate the association relationship to generate a multi-dimensional label behavior sequence for representing a video interest preference of the target user, wherein each dimension label corresponds to a behavior sequence.
[0057] The plurality of behavior types include at least two of exposure behavior, viewing behavior, and positive feedback behavior; the plurality of behavior sequences include at least two of an exposure behavior sequence, a viewing behavior sequence, and a positive feedback behavior sequence of the historical video; the exposure behavior sequence is used to record a number of times of presentation of the historical video to the target user, the viewing behavior sequence is used to record a playing time length of the historical video, and the positive feedback behavior sequence is used to record a trigger operation of expressing a preference for the historical video by the target user.
[0058] Optionally, the output unit 304 is configured to: obtain video preference information and / or historical viewing information of the target user; determine at least one target video to be recommended to the target user according to the multi-dimensional label behavior sequence and multi-dimensional label features of the at least one candidate video generated in advance, and the video preference information and / or the historical viewing information, using a pre-trained video recommendation model; The video recommendation model takes a user behavior index of the target user as an optimization target in a training process, and the user behavior index includes a click rate and / or a complete play rate of a video.
[0059] The multi-dimensional label system is composed of a label set generated based on a plurality of mutually independent semantic classification dimensions, and the semantic classification dimensions include at least two of video content, video theme, video style, video audience, video production mode, and video emotional tone.
[0060] Figure 3 The video recommendation apparatus of the illustrated embodiment can be used to execute the technical solutions of the method embodiments described above, and has similar implementation principles and technical effects, which will not be described again here.
[0061] Figure 4 The structural schematic diagram of the electronic device provided by the embodiments of the present disclosure is shown in FIG. 4. Figure 4 FIG. 4 shows a structural schematic diagram of an electronic device 400 suitable for implementing the electronic device in the embodiments of the present disclosure. The electronic device 400 in the embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle terminal (for example, a vehicle navigation terminal), a wearable electronic device, and the like, and a fixed terminal such as a digital TV, a desktop computer, a smart home device, and the like. Figure 4 The electronic device shown is merely an example, and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0062] As shown in FIG. 4, the electronic device 400 can include a communication unit 401, a user input unit 402, a user output unit 403, a storage unit 404, and a processor 405. Figure 4As shown, the electronic device 400 can include a processing device 401 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes to implement the video recommendation method of embodiments as described in the present disclosure according to programs stored in a read-only memory (ROM) 402 or loaded into a random access memory (RAM) 403 from a storage device 408. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0063] In general, the following devices can be connected to the I / O interface 405: input devices 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 408 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 409. The communication devices 409 can allow the electronic device 400 to communicate wirelessly or wired with other devices to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that all of the illustrated devices are not required, and more or fewer devices can alternatively be implemented.
[0064] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts, thereby implementing the video recommendation method as described above. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 409, or installed from the storage devices 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-described functions defined in the methods of embodiments of the present disclosure are performed.
[0065] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the foregoing.
[0066] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0067] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled in the electronic device.
[0068] Optionally, when the one or more programs described above are executed by the electronic device, the electronic device can further perform other steps described in the embodiments described above.
[0069] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0070] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0071] The units described in the embodiments of the present disclosure can be implemented by hardware, software, or a combination of hardware and software. In some cases, the names of the units do not constitute a limitation on the units themselves.
[0072] The functions described in this specification can be implemented in part or in whole through one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0073] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
[0074] It has to be noted that, in the present document, the terms "first" and "second", etc., are used only to identify different entities or operations, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0075] The above description is that of current embodiments of the disclosure. Various modifications will be apparent to those skilled in the art from this disclosure, which is intended to be illustrative and not limiting. Changes can be made to the embodiments described without departing from the spirit or scope of the disclosure. Thus, the present disclosure is not to be limited to the described embodiments but is intended to embrace the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A video recommendation method, characterized in that, include: Acquire historical behavioral data of the target user, wherein the historical behavioral data refers to various interaction data generated by the target user during the viewing of historical videos; A multi-dimensional tag system is established to provide a structured description of the video content of the historical videos, wherein each dimension tag in the multi-dimensional tag system is used to represent an independent semantic classification; Determine the correlation between the historical behavior data and the labels in the multi-dimensional label system, and integrate the correlation to generate a multi-dimensional label behavior sequence that characterizes the target user's video interest preferences; The multi-dimensional label behavior sequence and the multi-dimensional label features of at least one pre-generated candidate video are used as input to a pre-trained video recommendation model, which outputs at least one target video to be recommended to the target user.
2. The method according to claim 1, characterized in that, The step of using the multi-dimensional label behavior sequence and the multi-dimensional label features of at least one pre-generated candidate video as input to a pre-trained video recommendation model, and outputting at least one target video to be recommended to the target user, includes: Calculate the similarity between the multi-dimensional label features of at least one candidate video and the multi-dimensional label behavior sequence to obtain multi-dimensional label interest features, wherein the multi-dimensional label interest features are used to characterize the matching degree between the candidate video and the video that the target user is interested in; The multi-dimensional tag interest features and the multi-dimensional tag features are used as input to a pre-trained video recommendation model, which outputs at least one target video to be recommended to the target user.
3. The method according to claim 2, characterized in that, The calculation of the similarity between the multi-dimensional label features of at least one candidate video and the multi-dimensional label behavior sequence yields multi-dimensional label interest features, including: The multi-dimensional label behavior sequence is feature-encoded to obtain multi-dimensional label behavior features; The similarity between dimensional label features and dimensional label behavior features belonging to the same dimension is calculated using an attention mechanism to obtain the label interest features of each dimension; wherein, the multi-dimensional label features include the dimensional label features, and the multi-dimensional label behavior features include the dimensional label behavior features. The tag interest features of each dimension are weighted and fused or spliced to obtain unified multi-dimensional tag interest features.
4. The method according to claim 1, characterized in that, The step of determining the association between the historical behavior data and the labels in the multi-dimensional label system, and integrating the association to generate a multi-dimensional label behavior sequence to characterize the target user's video interest preferences, includes: The historical behavior data is used to construct multiple behavior sequences according to multiple behavior types, wherein the multiple behavior sequences are different behavior sequences extracted from different stages of the target user's interaction with the historical video; The multiple behavior sequences are mapped to the multi-dimensional label system to establish the association between different behaviors and dimension labels; The relationships are integrated to generate a multi-dimensional tag behavior sequence that represents the target user's video interest preferences, wherein each dimension tag corresponds to a behavior sequence.
5. The method according to claim 4, characterized in that, The multiple behavior types include at least two of exposure behavior, viewing behavior, and positive feedback behavior; the multiple behavior sequences include at least two of the exposure behavior sequence, viewing behavior sequence, and positive feedback behavior sequence of the historical video; the exposure behavior sequence is used to record the number of times the historical video is shown to the target user, the viewing behavior sequence is used to record the playback duration of the historical video, and the positive feedback behavior sequence is used to record the triggering operation of the target user expressing preference for the historical video.
6. The method according to claim 1, characterized in that, The step of using the multi-dimensional label behavior sequence and the multi-dimensional label features of at least one pre-generated candidate video as input to a pre-trained video recommendation model, and outputting at least one target video to be recommended to the target user, includes: Obtain the target user's video preference information and / or viewing history information; Based on the multi-dimensional tag behavior sequence and the multi-dimensional tag features of at least one pre-generated candidate video, as well as the video preference information and / or the historical viewing information, a pre-trained video recommendation model is used to determine at least one target video to be recommended to the target user; The video recommendation model is trained with the target user's user behavior metrics as the optimization objective, including the video's click-through rate and / or completion rate.
7. The method according to claim 1, characterized in that, The multi-dimensional tag system consists of a set of tags generated based on multiple independent semantic classification dimensions, wherein the semantic classification dimensions include at least two of the following: video content, video theme, video style, video audience, video production mode, and video emotional tone.
8. A video recommendation device, characterized in that, include: The acquisition unit is used to acquire the historical behavior data of the target user, wherein the historical behavior data refers to various interaction data generated by the target user during the viewing of historical videos; A building unit is used to build a multi-dimensional tag system for structurally describing the video content of the historical videos, wherein each dimension tag in the multi-dimensional tag system is used to represent an independent semantic classification; The generation unit is used to determine the association between the historical behavior data and the labels of each dimension in the multi-dimensional label system, and to integrate the association to generate a multi-dimensional label behavior sequence that represents the video interest preferences of the target user. The output unit is used to take the multi-dimensional label behavior sequence and the multi-dimensional label features of at least one pre-generated candidate video as input to a pre-trained video recommendation model, and output at least one target video to be recommended to the target user.
9. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the video recommendation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the video recommendation method as described in any one of claims 1 to 7.