Video processing method and device, electronic equipment and storage medium
By acquiring video information, recommendation strategies, and user interaction contexts, and using a granular decision model to generate video summaries, this technology addresses the problem of insufficient and inflexible video understanding in existing video recommendation systems, achieving more user-interest-oriented and diversified recommendation effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-07
AI Technical Summary
Existing video recommendation systems suffer from insufficient and inflexible understanding of videos when generating video descriptions, resulting in repetitive and unoriginal recommended content, difficulty in handling differences between long and short videos, and unsatisfactory recommendation results.
By acquiring video information, recommendation strategies, and user interaction contexts, a granular decision model is used to determine the level of detail in video descriptions, generating candidate descriptions that match the granular information, and then optimizing recommendations by combining user interest models.
The generated video descriptions are more tailored to user preferences, enhancing the personalization and diversity of the recommendation system and increasing the likelihood that users will discover videos that meet their needs.
Smart Images

Figure CN121815018A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a video processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] Video recommendation systems play a crucial role on major video platforms, helping users find videos of interest from a vast amount of content, while also increasing viewing time and user engagement.
[0003] Existing video recommendation systems use artificial intelligence to automatically analyze video visuals and audio to understand video content and then generate a text summary of the video content.
[0004] However, current solutions based on manually generated video descriptions suffer from insufficient and inflexible understanding of the videos, leading to repetitive and unoriginal recommended content. Summary of the Invention
[0005] This application provides a video processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can solve the above-mentioned problems of the prior art. The technical solution is as follows: According to one aspect of the embodiments of this application, a video processing method is provided, the method comprising: Obtain the video interaction context of the target object, the video recommendation strategy, and the video information of the video; The video information, the video recommendation strategy, and the video interaction context are input into a pre-trained granular decision model to obtain the granular information of the video, which is used to represent the level of detail in the video description. Based on the video information and granularity information, at least one candidate introduction that matches the granularity information is generated.
[0006] According to another aspect of the embodiments of this application, a video processing apparatus is provided, the apparatus comprising: The information acquisition module is used to obtain video recommendation strategies, the video interaction context of target objects, and video information of the videos; The granularity determination module is used to input the video information, the video recommendation strategy, and the video interaction context into a pre-trained granularity decision model to obtain the granularity information of the video, wherein the granularity information is used to represent the level of detail of the video description; The description generation module is used to generate at least one candidate video description that matches the granularity information based on the video information and the granularity information.
[0007] According to another aspect of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-described video processing method.
[0008] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the above-described video processing method.
[0009] According to one aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described video processing method.
[0010] The beneficial effects of the technical solutions provided in this application are: By acquiring the video interaction context of the target audience, the video recommendation strategy, and the video information, the generation of video descriptions is deeply coupled with the video recommendation strategy of the recommendation system and the video interaction context of the target audience. Furthermore, instead of directly generating video descriptions, a granular decision model intelligently determines the granularity information of the video based on its own video information, video interaction context, and video recommendation strategy. Granularity information represents the level of detail in the video description, thus providing suggestions on the word count and focus of the description. Finally, at least one candidate video description matching the granularity information is generated based on the video information and granularity information. The generated video descriptions are more likely to be prioritized by the video recommendation system and are more aligned with the preferences of the target audience, effectively increasing the likelihood that the target audience will discover videos that meet their needs. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0012] Figure 1 This is a schematic diagram of the system architecture for implementing the video processing method provided in the embodiments of this application; Figure 2 A flowchart illustrating a video processing method provided in an embodiment of this application; Figure 3 A flowchart illustrating a video processing method provided in an embodiment of this application; Figure 4 A schematic diagram illustrating a process for determining a target video summary by an electronic device, as provided in an embodiment of this application; Figure 5 A schematic diagram illustrating a process for obtaining a target video summary by an electronic device, as provided in an embodiment of this application; Figure 6A schematic diagram illustrating a process for generating candidate feature video descriptions by an electronic device using a description generation model, as provided in an embodiment of this application; Figure 7 This is a schematic diagram illustrating a video recommendation process performed by an electronic device, as provided in an embodiment of this application. Figure 8 A flowchart illustrating the process of optimizing video recommendation based on video descriptions, provided for an embodiment of this application; Figure 9 This is a schematic diagram of the structure of a video processing apparatus provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0014] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.”
[0015] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0016] Currently, video recommendation systems play a very important role on major video platforms. They help users find videos that interest them from a vast amount of content, while also increasing the platform's viewing time and user stickiness.
[0017] Cold start: This mainly targets newly launched content, such as new videos or new users. Due to the lack of historical data, the system struggles to accurately determine its quality, popularity, and relevance to user interests, making recommendations more difficult. For example, with newly released videos, there is no data on user views, likes, or comments, making it difficult for the recommendation system to determine which users to recommend them to.
[0018] Warm start: This refers to the startup state when content or users re-enter the system, which already has a certain data base and user feedback. The system can then use accumulated data to make more accurate recommendations and provide personalized services. For example, a video that has been published for some time and has accumulated a certain number of views and likes, or a returning user whose interests the system has learned about.
[0019] Existing video recommendation systems typically use basic information to understand video content, such as the video title, tags, category, or manually edited descriptions. With the development of artificial intelligence technology, especially video understanding and natural language processing, some systems are beginning to explore the use of automatic video description generation technology. This technology can automatically analyze video visuals and audio, and then generate a text summary of the video content.
[0020] Specifically, in existing technologies, automatically generated video summaries typically operate as a standalone module. It uses a Transformer model to analyze each frame of the video's image and audio signals, identifying objects, actions, scenes, and dialogues, and then organizing them into sentences in natural language. These generated text descriptions are sometimes referred to as video summaries or captions.
[0021] When these descriptions are used in recommendation systems, the most common approach is to treat them as additional textual features of the video. That is, the recommendation system concatenates these automatically generated descriptions with the video's original title, tags, description, and other textual information to form a longer text string. Then, the recommendation system uses text analysis techniques (such as word vectors and text encoding) to extract features from this text, and combines this with data such as the user's viewing history and interests to calculate the similarity between videos or the degree of match between videos and the user's interests, thereby making recommendations.
[0022] However, this simple approach has some problems. First, most video description models are not designed with the specific needs of recommendation systems in mind. Their goal is to describe video content as accurately and fluently as possible, rather than to improve recommendation performance. For example, the generated descriptions have a relatively fixed granularity and are not flexibly adjusted based on whether the video is short or long, or which aspect the recommendation system needs to emphasize. Second, these descriptions are used as static features without deep integration and optimization with the goals of the recommendation system (such as recommendation diversity and user satisfaction). This may result in the recommendation system's understanding of the videos still being insufficient, sometimes recommending overly similar or uninteresting videos, especially when dealing with content with significant differences in length, such as short and long videos. In such cases, the results are often unsatisfactory, making users feel that the recommended content is not fresh or diverse enough.
[0023] The relevant technologies face several key challenges when processing video content and making recommendations: 1. Insufficient and Flexible Video Understanding: Current video description technologies, while capable of generating some text, fail to recognize that these descriptions are intended for recommendation purposes. Consequently, the descriptions generated by these technologies are rather rigid and simplistic, regardless of whether the video is long or short, or what information the recommendation system wants to emphasize. This results in a lack of nuanced understanding of the videos by the recommendation system, making it unable to grasp the truly engaging aspects of the video for users.
[0024] 2. Recommended content is prone to repetition and lacks originality: Because video descriptions are not flexible enough, the recommendation system can only treat these descriptions as ordinary text, unable to optimize based on users' specific preferences (such as whether users prefer exciting plots or lighthearted slice-of-life content), or based on goals such as "recommending more unseen and different types of content to users." As a result, users constantly see similar content, feeling that the recommendations lack surprise and diversity.
[0025] 3. Lack of differentiation in handling short and long videos: Current technology struggles to effectively handle videos of different lengths simultaneously. For short videos, overly long descriptions lack focus; for long videos, overly short descriptions fail to capture the essence. This makes recommendation systems inadequate when dealing with diverse content formats.
[0026] The video processing methods, apparatus, electronic devices, computer-readable storage media, and computer program products provided in this application are intended to solve the above-mentioned technical problems of the prior art.
[0027] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0028] Figure 1 This is a schematic diagram of the system architecture for implementing the video processing method provided in the embodiments of this application. The system architecture includes a server 101 and a client 102. The server 101 and the client 102 can communicate with each other. The communication method can be wired communication technology, such as communicating through a network cable or serial cable; or wireless communication technology, such as communicating through Bluetooth or wireless fidelity (WIFI). There is no specific limitation.
[0029] Client 102 generally refers to devices capable of displaying videos or video descriptions, such as terminal devices, third-party applications accessible by terminal devices, or web pages accessible by terminal devices. Server 101 generally refers to devices capable of generating video descriptions, such as terminal devices or servers. Terminal devices include, but are not limited to, mobile phones, computers, smart medical devices, smart home appliances, in-vehicle terminals, or aircraft. Servers include, but are not limited to, cloud servers, local servers, or associated third-party servers.
[0030] In some embodiments, both the client 101 and the server 102 can use cloud computing to reduce the consumption of local computing resources; similarly, cloud storage can also be used to reduce the consumption of local storage resources.
[0031] As one embodiment, the server 101 and the client 102 can be the same device, or they can be different devices, or they can be different devices with some modules shared, etc., and there are no specific restrictions.
[0032] This application provides a video processing method that can be executed by an electronic device, which can be at least one of a server 101 and a client 102. For example, the method can be executed by the client 102 alone, by the server 101 alone, or by the server 101 and the client 102 working together. Figure 2 As shown, the method includes: S101. Obtain the video recommendation strategy, the video interaction context of the target object, and the video information of the video.
[0033] Existing technologies often treat video description generation as simply text-to-text conversion without considering its potential for recommendation. Consequently, they typically only require video information, resulting in rigid description formats that fail to capture the essence of the video (whether long or short) or the information the recommendation system prioritizes. This leads to an inaccurate and incomplete understanding of the video description by the recommendation system, hindering the effective recommendation of videos based on their unique features and user preferences. To overcome these issues, this application not only acquires video information but also the recommendation system's video recommendation strategy and the target user's video traffic records before generating a description. By gathering information from the video, the recommendation system, and the user, a foundation is laid for obtaining video descriptions that meet both the recommendation strategy and user preferences.
[0034] The video interaction context of the target object is used to characterize the scene, environment, and interactive behavior of the target user when interacting with video content, and has significant technical value. Specifically, the scene referred to in the embodiments of this application can include the type of page in which the user browses video content. The page type includes, but is not limited to: a short video information stream page, which displays multiple short videos in a fast scrolling manner; a long video details page, which presents detailed information and a playback window for a single long video; and a search results page, which displays a list of videos related to the user's search keywords. At the same time, by recording and analyzing the target object's most recent N (N is a positive integer, such as 10) interaction behaviors, the target object's preference for content granularity or type can be mined. The interaction behaviors in the embodiments of this application include click operations, reflecting the target object's initial interest in a specific video; viewing behavior, reflecting the target object's actual level of attention to the video content; and dwell behavior, indicating the target object's willingness to browse the video content in depth. These video interaction contexts, after integration and processing, can be constructed into information based on the target object's behavior sequence modeling, providing key data support for subsequent technologies such as personalized recommendation.
[0035] It should be noted that the data collection and processing in this application embodiment strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0036] In some embodiments, video information includes video metadata, such as video duration, video category, existing tags and keywords, etc. Video duration supports subsequent determination of video type (long or short) based on duration and generation of a description matching the video type. Video categories cover various types, including but not limited to movies, TV series, variety shows, short dramas, and UGC short videos. Clearly defining the video's category helps build a classification model, enabling accurate classification, storage, and retrieval of videos, meeting users' needs for browsing videos by category. The embedding of existing tags and keywords is an important extension of this application's embodiments. By converting tags and keywords into vector form and utilizing a vector space model, the semantic similarity between tags / keywords can be quantified, thereby uncovering potential connections between video content and providing technical support for intelligent video recommendation, similar video retrieval, and other steps.
[0037] In some embodiments, video information may include, in addition to video metadata, keyframes and audio information in the video.
[0038] Video recommendation strategies, as the core logic of recommendation systems, play a crucial guiding role in system performance and user experience. The recommendation strategies in this application can encompass various optimization objectives. For example, a strategy emphasizing diversity aims to present users with rich and diverse content, avoiding highly homogenized recommendation results. This is achieved by introducing diversity metrics, such as content category distribution and creator diversity, and adding diversity constraints to the recommendation algorithm, prompting the system to filter content for users from different dimensions. A strategy emphasizing precise matching focuses on personalized user needs. Based on user historical behavior, interest preferences, and other data, it uses precise user profile modeling and similarity calculation methods to accurately push content highly relevant to the user's interests. A strategy that utilizes cold-start exploration, targeting new users or new content, employs specific exploration mechanisms, such as random exploration or context-based exploration, to help uncover the potential interests of target users, enrich system data, and improve the overall performance and adaptability of the recommendation system while ensuring a certain level of recommendation quality.
[0039] S102. Input the video information, video recommendation strategy, and video interaction context into the pre-trained granular decision model to obtain the granular information of the video.
[0040] The granular information in this embodiment represents the level of detail in the video description. Since the model input includes information from three dimensions: the video information, the video recommendation strategy of the recommendation system, and the video interaction context of the target UI, the granular information output by the granular decision model fully represents the level of detail in the video description of the video to the target object under the video recommendation strategy.
[0041] In some embodiments, granularity information can be a value between 0 and 1. The larger the value, the greater the level of detail, that is, the more detailed and specific the video description.
[0042] In some embodiments, the granularity information can also be a probability distribution at predefined levels of detail. For example, embodiments of this application can set three levels of detail: The basic summary type uses the most concise language to summarize the core theme and key information of the video. It usually only includes the video type, main plot or function, and is short in length, usually only one or two sentences.
[0043] This detailed introduction builds upon the basic overview, providing further details such as the main plot development, key character introductions, unique features, and applicable scenarios or target audiences. The language is relatively detailed, the length is moderate, and it may include several paragraphs.
[0044] The in-depth analysis type not only covers the detailed introduction type, but also provides in-depth analysis from a professional perspective, such as interpreting the artistic value and cultural connotation of the video.
[0045] In some embodiments, the granular decision model can use the ReLU activation function. The ReLU activation function has a simple form and can be calculated quickly in both forward and backward propagation, reducing model training time and effectively alleviating the gradient vanishing problem. In deep networks, compared with activation functions such as Sigmoid, it allows gradients to be propagated better in the network, enabling deep neural networks to be trained effectively and ensuring model performance.
[0046] In some embodiments, the granular decision model has a two- or three-layer fully connected network. The two- or three-layer fully connected network structure is relatively simple, with a moderate number of parameters. It can learn the complex nonlinear relationship between input features and the level of detail in the video description without overfitting due to excessive complexity. It is also less demanding in terms of data volume, performing well even on small datasets. Furthermore, the simple structure facilitates understanding and debugging, allowing for faster problem identification and adjustments during model optimization and improvement, thus increasing development efficiency.
[0047] It should be understood that, prior to step S102, a granular decision model can be pre-trained, specifically through the following method: First, obtain the video browsing records of the sample objects fed back by the video recommendation system, and identify videos with longer viewing time and positive interactive behaviors (such as liking, posting positive comments, sharing, etc.) from the video browsing records as sample videos; Then, obtain the video information and video description of the sample video, the video recommendation strategy when recommending the sample video, the video interaction context of the sample object to the sample video, and then manually annotate the level of detail of the video description. Of course, the level of detail of the video description can also be analyzed through a large language model to obtain the granular information of the sample video. Subsequently, the initial model is trained based on the video interaction context of the sample objects, the video recommendation strategy, the video information of the sample videos, and the granularity information of the sample videos. The video interaction context of the sample objects, the video recommendation strategy, and the video information of the sample videos are used as training samples, and the granularity information of the sample videos is used as sample labels, thus obtaining a granular decision model. The initial model can be a single neural network model or a combination of multiple neural network models. S103. Based on the video information and granularity information, generate at least one candidate video description that matches the granularity information.
[0048] This application embodiment can construct prompt words for a large language model based on granularity information, thereby utilizing the understanding capability of the large language model to obtain at least one candidate video summary matching the granularity information based on the prompt words and video information. For example, the prompt word could be "You are a professional video summary writer. Please generate at least one video summary based on the video information. The video summary should be limited to 50 characters and only highlight the core highlights of the video." It can be understood that the phrase "the video summary should be limited to 50 characters and only highlight the core highlights of the video" can be a prompt word specifically designed for generating video summaries with lower levels of detail. This application embodiment can construct different prompt words for different levels of detail.
[0049] The embodiments of this application can generate at least one candidate video description with the same level of detail. It is understood that different candidate video descriptions may have different focuses on video content, video interaction context of target objects, and video recommendation strategies.
[0050] For example, consider a "Yoga for Beginners" video. The video content includes explanations and demonstrations of basic yoga postures, breathing techniques, and precautions for yoga practice. A video description focusing on the content could be: "This 'Yoga for Beginners' video explains breathing techniques in detail, teaches postures such as Mountain Pose and Downward-Facing Dog, and provides precautions to help you solidly begin your yoga journey." A video description focusing on the interactive context for the target audience could be: "Office workers can practice yoga at home with this video when they are tired. The instructor provides gentle guidance, the visuals are simple, and it helps you relax your mind and body and improve your quality of life." A video description focusing on video recommendation strategies (such as prioritizing videos for new users or identifying content that the target audience wants to explore) could be: "This video presents the latest yoga teaching concepts, with innovative and cutting-edge guidance on multiple postures."
[0051] Please see Figure 3 The figure illustrates a flowchart of the video processing method provided in this application embodiment. As shown, this application embodiment first obtains a video recommendation strategy, the video interaction context of the target object, and video information 1 of video 1 and video information 2 of video 2. The video information of this application includes video type. The video category can be simply classified into long videos and short videos, or a more granular classification can be made for long videos: movies, TV series, variety shows, etc., and a more granular classification can be made for short videos: short dramas, UGC short videos, etc. In this embodiment, the video category of video information 1 is long video, and the video category of video information 2 is short video. The video interaction request includes information about the page where the target object is currently standing and a sequence of content composed of the most recent N interaction behaviors. The video information includes video duration, video type, video tags, and keyframes. Then, the video information of video 1 and 2, the video recommendation strategy, and the video interaction context are input into a pre-trained granular decision model to obtain the granular information of video 1 and video 2. Prompt words are constructed according to the granular information of the two videos respectively. The video information and prompt words of each video are input into a large language model to obtain multiple candidate video summaries of video 1 and video 2 output by the large language model. Compared to existing technologies where recommending long and short videos often results in neglecting one over the other, the solution in this application can efficiently process these two different video formats within a unified framework, providing appropriate and matching video descriptions for both short films and feature films.
[0052] The video processing method provided in this application deeply couples the generation of video descriptions with the video recommendation strategy of the recommendation system and the video interaction context of the target object by acquiring the video interaction context of the target object, the video recommendation strategy, and the video information of the video. Furthermore, instead of directly generating video descriptions, it intelligently determines the granularity information of the video based on its own video information, video interaction context, and video recommendation strategy using a granular decision model. The granularity information represents the level of detail in the video description, thus providing suggestions on the number of words and focus of the description. Finally, at least one candidate video description matching the granularity information is generated based on the video information and the granularity information. The generated video descriptions are more likely to be prioritized by the video recommendation system and are more aligned with the preferences of the target object, effectively increasing the probability that the target object will discover videos that meet its needs.
[0053] Based on the above embodiments, as an optional embodiment, when there are multiple candidate video descriptions, the method further includes: Obtain the interest model of the target object; Calculate the matching degree between the description of each candidate video and the interest model; Based on the matching degree corresponding to each candidate video description, the target video description of the video is determined from at least one candidate video description.
[0054] When there are multiple candidate video descriptions provided in this application embodiment, it is necessary to further filter from these candidate video descriptions to select the final target video description. This process requires further obtaining the target object's interest model. The target object's interest model is a digital standard of the target object's interests, behavioral habits, and needs in video browsing. The data used to construct the interest model can include explicit feedback behavioral data such as the target object's ratings, comments, favorites, and shares, as well as implicit feedback behavioral data such as the recommendation system's viewing behavior of the target object, such as viewing duration, viewing progress, and number of repeated views.
[0055] In some embodiments, the interest model may include word vectors of various keywords that the target object is interested in. Further, embodiments of this application may categorize the keywords into keywords reflecting the target object's long-term interests and keywords reflecting the target object's short-term interests. Long-term interests can be reflected by the types of content the target object has watched in the past, actors, and thematic preferences, while short-term interests can be reflected by the videos the target object has recently watched and the search terms they have recently used.
[0056] In calculating the matching degree between candidate video descriptions and interest models, this embodiment first performs preprocessing such as word segmentation on the candidate video descriptions, converting them into word vector form. Next, it calculates the similarity (e.g., cosine similarity) between each word vector in the candidate video description (referred to as description word vector) and the word vectors of each keyword in the interest model (interest word vector). The maximum similarity value can represent the degree of matching between the description word vector and the interest model. Then, it determines the matching status of all description word vectors in the candidate video description, such as by averaging the results. The higher the value, the better the candidate video description matches the interest model and the more it aligns with user interests.
[0057] Based on the above embodiments, as an optional embodiment, this application also provides a scheme for semantically optimizing and filtering multiple candidate video descriptions based on the matching degree corresponding to each candidate video description, thereby determining the target video description of the video, including: Identify at least one combination of descriptions, each combination of descriptions including at least two candidate video descriptions; For each combination of descriptions, the difference between a first value and a second value is determined. The first value is the sum of the matching degrees of all candidate video descriptions in the combination, and the second value is the sum of the semantic similarities between any two candidate video descriptions in the combination. The candidate video descriptions in the combination of descriptions with the greatest differences are merged to obtain the target video description.
[0058] In this process, this application considers the matching degree between the introduction and the interests of the target audience. On the other hand, in order to avoid homogenization of recommendations, it evaluates the semantic similarity between various candidate video introductions. The goal is to select or provide a combination of candidate video introductions that can both match the interests of the target audience and ensure diverse descriptive content.
[0059] This application embodiment can traverse multiple candidate video descriptions to determine all description combinations. For example, taking four candidate video descriptions: A, B, C, and D, description combinations can include: {A,B}, {A,C}, {A,D}, {C,B}, {D,B}, {C,D}, {A,B,C}, {A,B,D}, and {D,B,C}, etc. For each description combination, this application embodiment can calculate the difference between a first value and a second value in the description combination. Taking the description combination {A,B,C} as an example, the first value of this description combination is the sum of the matching degrees corresponding to A, B, and C, and the second value is the sum of the semantic similarity between A and B, the semantic similarity between A and C, and the semantic similarity between B and C.
[0060] In some embodiments, when calculating the difference between two first values, the first value or the second value may also be weighted to balance the matching degree and diversity.
[0061] In some embodiments of this application, the following formula may be used to determine the most significant combination of differences:
[0062] Among them, D optimized D represents a concise combination. conditate Let M(U) represent the set of all candidate video descriptions, d represent a candidate video description, and M(U) represent the set of all candidate video descriptions. interest ,d) represents the candidate video description and interest model U interest The matching degree between them, SD(d) i ,d j ) represents the semantic similarity between any two candidate video descriptions in the description combination.
[0063] It should be noted that, after determining the target summary combination with the greatest difference, this application embodiment needs to merge the candidate video summaries in the target summary combination. Specifically, in order to ensure that the word count of the target video summary conforms to the granularity information, the key to the fusion process is to simplify each candidate video summary. This application counts the keywords of each candidate video summary in the target summary combination and calculates the frequency of the keywords in each candidate video summary in the target summary combination. For a keyword, if the frequency of the keyword is higher than the threshold, the statement containing the keyword is determined to be a statement that can be simplified. One statement from all statements that can be simplified related to the same keyword is retained. Then, the total number of words remaining in each candidate video summary is determined. If the total number of words still does not conform to the level of detail indicated by the granularity information, the scarcity of each keyword in the target summary combination is further counted. The statements corresponding to the keywords are retained in descending order of scarcity until the total number of words remaining in each candidate video summary is reached. If the total number of words still does not conform to the level of detail indicated by the granularity information, the scarcity of each keyword in the target summary combination is further counted. The statements corresponding to the keywords are retained in descending order of scarcity until the total number of words remaining in each candidate video summary is reached.
[0064] In some embodiments, this application can obtain video descriptions of other videos with the same or similar themes. The lower the frequency of a keyword appearing in the video descriptions of other videos, the higher the scarcity of the keyword, and the greater the attraction of the keyword to the target audience.
[0065] Please see Figure 4 The figure illustrates, exemplarily, a flowchart of an electronic device determining a target video summary according to an embodiment of this application, as shown, including: S201. Obtain the interest model of the target object; S202. Calculate the matching degree between each candidate video description and the interest model; S203. Determine at least one combination of descriptions, each combination of descriptions including at least two candidate video descriptions; S204. For each combination of descriptions, determine the difference between the first value and the second value, wherein the first value is the sum of the matching degrees corresponding to all candidate video descriptions in the combination, and the second value is the sum of the semantic similarity between each pair of candidate video descriptions in the combination. S205. Select the combination of descriptions with the greatest differences as the target combination of descriptions, and count the keywords of each candidate video description in the target combination of descriptions; S206. Calculate the frequency of each keyword in each candidate video description in the target description combination. If the frequency of the keyword is higher than the threshold, the statement with the keyword is determined to be a statement that can be simplified. Keep one statement among all the statements that can be simplified related to the same keyword. S207. Determine whether the total number of remaining words in the description of each candidate video is within the word count range that matches the granularity information. If not, proceed to S208; if yes, proceed to S209. S208. Analyze the scarcity of each keyword in the target description combination, retain the sentences corresponding to the keywords in descending order of scarcity, and return to step S207. S209. Adjust the word order of the remaining sentences in each candidate video description to make the content transition naturally and logically coherent, thus obtaining the target video description.
[0066] Based on the above embodiments, as an optional embodiment, each keyword in the interest model is converted into word vector form to obtain each interest word vector; For each candidate video description, the candidate video description is segmented into words, and each segment is converted into word vectors to obtain multiple description word vectors; For each description word vector of each candidate video description, the similarity between the description word vector and each keyword vector is determined, and the maximum similarity is taken as the target similarity of the description word vector. According to the video recommendation strategy and the type of the keyword corresponding to the target similarity, the weight of the description word vector is determined; the type refers to whether the keyword belongs to the first keyword or the second keyword. For each candidate video description, the target similarity of each description word vector is weighted and summed according to the weight of each description word vector in the candidate video description to obtain the matching degree between the candidate video description and the interest model.
[0067] Please see Figure 5 The illustration shows a schematic diagram of the process of obtaining a target video description by an electronic device according to an embodiment of this application. As shown in the figure, this application converts each keyword in the interest model into word vectors to obtain each interest word vector. On the other hand, for each candidate video description, the candidate video description is segmented into words, and each segmented word is converted into word vectors to obtain multiple description word vectors. Then, for each description word vector of each candidate video description, the similarity between the description word vector and each keyword vector is determined, and the maximum similarity is taken as the target similarity of the description word vector.
[0068] Considering that the keywords in this application include both keywords reflecting the long-term interests of the target audience and keywords reflecting short-term interests, this application comprehensively considers the current video recommendation strategy and the type of keywords corresponding to the target similarity when determining the weight of a descriptive word vector. For example, if the current video recommendation strategy is a warm start, the weight of the keyword corresponding to the target similarity as the first keyword is greater than the weight of the keyword corresponding to the target similarity as the first keyword. Conversely, if the current video recommendation strategy is a cold start, the weight of the keyword corresponding to the target similarity as the first keyword is less than the weight of the keyword corresponding to the target similarity as the first keyword.
[0069] Furthermore, for each candidate video description, the target similarity of each description word vector is weighted and summed according to its weight to obtain the matching degree between the candidate video description and the interest model. Finally, by ranking the matching degrees between each candidate video description and the interest model, the candidate video description with the highest matching degree is selected as the target video description.
[0070] Based on the above embodiments, as an optional embodiment, this application embodiment inputs the video information, video recommendation strategy, and video interaction context into a pre-trained granular decision model to obtain granular information, including: The video information, video recommendation strategy, and video interaction context are respectively feature-encoded to obtain the feature vectors of the video information, video recommendation strategy, and video interaction context. The feature vectors of the video information, video recommendation strategy, and video interaction context are fused to obtain the decision input vector of the video. The decision input vector is then input into a granular decision model to obtain the granular information.
[0071] This application embodiment performs feature encoding on video information, video recommendation strategies, and video interaction scenarios separately. This allows for the detailed extraction of key features from each piece of information, transforming them into feature vectors. This separate encoding process avoids information confusion and omissions, comprehensively covering all factors affecting video recommendations and laying a solid foundation for subsequent accurate recommendations. Furthermore, the three feature vectors are fused to obtain a decision input vector. This decision input vector is then input into a granular decision model, enabling the granular decision model to analyze from a holistic perspective, comprehensively considering the impact of various factors on granular information, and improving the rationality of the decision.
[0072] In some embodiments, this application may concatenate the feature vectors of video information, video recommendation strategy and video interaction context in a preset order to obtain a decision input vector. This concatenation method is simple to operate and can completely preserve the original feature information.
[0073] Furthermore, in this embodiment of the application, based on the video information and granularity information, at least one candidate video description matching the granularity information is generated, including: Based on the granularity information, feature vector group is used for feature decoding to obtain at least one candidate video description of the video. The feature vector group includes at least one of the feature vectors of the video information, the video recommendation strategy, and the video interaction context, and includes at least the feature vector of the video information.
[0074] This application obtains candidate video descriptions through feature decoding. The feature vector group includes at least the feature vectors of video information, and may also include at least one of the feature vectors of the video recommendation strategy and the video interaction context.
[0075] Please see Figure 6 The figure exemplifies a flowchart illustrating the process by which an electronic device obtains candidate feature video summaries through a summary generation model, as provided in this application embodiment. As shown, the summary generation model includes a video information encoder, a video recommendation strategy encoder, a video interaction context encoder, a granular decision model, and a decoder. This application sets corresponding encoders for video information, video recommendation strategies, and the video interaction context. These three types of information are mapped to a unified dimension through their respective encoders. Then, the feature vectors of the three types of information are concatenated to obtain a decision input vector. This decision input vector is then input into the granular decision model to obtain the granular information of the video. The granular information can be a value S. granularity ∈[0,1], or a probability distribution at a predefined granularity level. Embodiments of this application prefer the former because it provides more flexible control, inputting the granular information of the video and the feature vector set into the decoder to obtain at least one candidate video summary.
[0076] In some embodiments of this application, the decision input vector X GDU It can be obtained through the following formula:
[0077] Here, Embed() represents the feature embedding function, and [;] represents the vector concatenation operation. F video Feature vectors representing video information F user_ctx Feature vectors representing the video interaction contextF rec_goal The feature vector represents the video recommendation strategy.
[0078] In some embodiments, granularity information S granularity It can be obtained through the following formula:
[0079] Where W1, b1, W2, b2 are the learnable parameters of the granular decision model, the sigmoid function reduces the output to between 0 and 1, and S... granularity The higher the value, the greater the level of detail.
[0080] In some embodiments, the decoder adjusts its strategy for generating video summaries based on granularity information, for example: Control the number of words in the generated summary, S granularity The higher the value, the more words the introduction will contain; Adjusting the focus range of the spatiotemporal attention mechanism: The decoder in this application introduces an attention mechanism, enabling the decoder to adjust the degree of attention to different key points in the video according to granular information, resulting in a lower S granularity This allows for greater focus on the overall, high-level aspects of video information, generating a macroscopic description, while higher S... granularity This will guide attention to delve deeper into the details and local events of the time sequence in the video; Activate different generation templates or vocabulary: for different S granularity The decoder will be guided to use different language styles or keywords, for example, targeting higher S... granularity Event description vocabulary (e.g., specific actions, scene details, etc.) and targeting lower S granularity General terms (such as topic categories, overall evaluation).
[0081] In this way, the embodiments of this application realize intelligent, contextualized, and targeted video description generation, ensuring that the generated descriptions are highly consistent with the needs of the recommendation system in terms of semantic content and information content, and significantly improving the efficiency and effectiveness of the entire recommendation chain.
[0082] The embodiments of this application do not limit the type of encoder, such as an embedding layer or a multilayer perceptron (MLP).
[0083] Based on the above embodiments, as an optional embodiment, the abstract generative model is trained in the following way: Multiple training samples and corresponding training labels are obtained. The training samples include sample video information, sample video recommendation strategies, and sample video interaction scenarios. The training labels include sample video descriptions corresponding to the sample video information and the actual granularity information of the sample video descriptions. For each training sample, the training sample is input into the initial encoder group to obtain the feature encoding of the sample video information, the sample video recommendation strategy and the sample video interaction context respectively. The feature encoding of the sample video information, the sample video recommendation strategy and the sample video interaction context is fused and then input into the initial granularity decision model to obtain prediction granularity information. Input the prediction granularity information and feature vector group into the initial decoder to obtain at least one predicted video summary; A first loss value is obtained based on the difference between the predicted granularity information and the actual granularity information; Determine the difference in word segmentation complexity and syntactic complexity between each predicted video description and the actual video description to obtain a second loss value; The total loss function value is obtained based on the first loss function value and the second loss value. The initial introduction generation model is trained based on the total loss function value until the total loss function converges. The initial introduction generation model at the end of training is taken as the introduction generation model.
[0084] Specifically, the descriptive generation model of this application embodiment can be trained in the following ways: S301. Based on the feedback results of each sample object to the recommended video, determine multiple sample videos that meet the conditions for the feedback results of each sample object and the actual video description of each sample video, and obtain the video recommendation strategy, the video interaction context and the video information of the sample videos. S302. For the actual video description of each sample video, manually annotate the actual granularity information of the video description. S303. Construct multiple training samples with training labels. Each training sample is a triple: video recommendation strategy, video interaction context, and video information of a sample video. The training labels include the actual granularity information of the sample video and a brief description of the actual video. S304. For each training sample, input the training sample into the video information encoder, the video recommendation strategy encoder, and the video interaction context encoder to obtain the feature vector of the triple. Concatenate the feature vectors of the triple to obtain the decision input vector. Input the decision input vector into the granular decision model to obtain the predicted granularity information of the training sample. Calculate the difference between the predicted granularity information and the actual granularity information as the first loss value. In some embodiments, the predicted granularity information and the actual granularity information can both be numbers between 0 and 1, so the absolute value of the difference between the two can be used as the first loss value. S305. Input the prediction granularity information and feature vector group into the decoder to obtain at least one predicted video description and construct at least one description pair, each description pair including a predicted video description and an actual video description. S306. For each description pair, the cosine similarity between the word vectors of the two video descriptions in the description pair is taken as the word similarity S. lex After standardizing the sentence length and the number of clauses, the Euclidean distance D is calculated. syn And convert it into syntactic similarity S syn =1-D syn ; S307. For each video description in the description pair, calculate the inverse document frequency (IDF) of each word in the video description and take the mean. This gives the mean IDF of the description. Then, take the mean of the two IDF means of the description pair to obtain the lexical complexity C of the description pair. lex ; S308. For each video description in the description pair, calculate the average sentence length of the video description based on the quotient between the total number of words and the number of sentences. Use the average sentence length of the two video descriptions as the syntactic complexity C. syn ; S309, Lexical complexity C of the introduction pair lex Syntactic complexity C syn After normalization, the word weights w corresponding to the word similarity are obtained. lex and the syntactic weight w corresponding to the syntactic similarity. syn .
[0085] S310. Based on the vocabulary similarity S of the introduction pair lex Syntactic similarity S syn The corresponding weights are used to weight and sum the lexical similarity and syntactic similarity to obtain the comprehensive similarity of the introduction pair; S311. The difference between the preset value and the comprehensive similarity of the description pair is taken as the difference between the predicted video description and the actual video description in the description pair in terms of word segmentation and syntax. In some embodiments, it can be calculated by the following formula: 1-S lex ×w lex - S syn ×w syn ; S312. Based on the differences in word segmentation and syntax between all predicted video descriptions and the actual video descriptions, obtain the second loss value; S313. The sum of the first loss value and the second loss value is taken as the total loss. The parameters of the description generation model are adjusted according to the total loss. The parameters of the description generation model in this embodiment include the parameters in the video information encoder, the video recommendation strategy encoder, the video interaction context encoder, the granular decision model, and the decoder.
[0086] Based on the above embodiments, as an optional embodiment, this application embodiment further includes: The target video description of the video is feature-encoded to obtain the information vector of the video; A multi-objective recommendation model is used to obtain multi-objective prediction and recommendation results based on the information vector of the video and the interest model.
[0087] This application embodiment encodes the target video description into an information vector using feature encoding, thus converting unstructured text into a structured semantic representation. It is understood that even the least detailed video description typically includes information across multiple dimensions, albeit with some omissions. Low-dimensional vectors (e.g., 2-3 dimensions) can only represent simple features, while high-dimensional vectors can correspond to different semantic subspaces across different dimensions. Furthermore, during video recommendation, video descriptions with different semantics in the low-dimensional space may be compressed to similar positions due to dimensional limitations, while the high-dimensional space adds independent dimensions, providing a "separation channel" for similar but different content, thereby reducing semantic conflicts and improving distinguishability. Therefore, this application can encode the target video description into a high-dimensional information vector using natural language processing technology. By increasing the number of dimensions, it can capture more complex and nuanced semantic information from the text summary.
[0088] In some embodiments, the information vector of this application has a dimension number greater than 100.
[0089] Traditional video recommendation systems are often driven by a single objective (such as click-through rate), making them prone to "local optima." This application introduces a multi-objective recommendation model that can simultaneously optimize multiple business objectives. This application does not specifically limit the type of business objective; it can include, for example, click-through rate, viewing time, completion rate, and interaction rate (likes / comments / shares). For example, for educational videos, the multi-objective recommendation model balances click-through rate and completion rate objectives, using weighted fusion or Pareto optimization algorithms to generate recommendation priorities that consider both short-term interaction and long-term value, avoiding over-recommendation of low-quality short videos or high-barrier long content.
[0090] Please see Figure 7 The figure illustrates an exemplary flowchart of video recommendation by an electronic device according to an embodiment of this application. First, the video information, video recommendation strategy, and video interaction scenario are respectively input into the corresponding encoders for feature encoding to obtain the feature vectors of the video information, video recommendation strategy, and video interaction scenario respectively. The feature vectors of the video information, video recommendation strategy, and video interaction context are fused to obtain the decision input vector of the video. The decision input vector is then input into a granular decision model to obtain the granular information. The granular information and feature vector group are input into the decoder for feature decoding to obtain m candidate video summaries of the video (m is an integer). Calculate the matching degree between the description of each candidate video and the interest model of the target audience; Based on the matching degree corresponding to each candidate video description, the target video description of the video is determined from at least one candidate video description. Feature encoding is performed on the target video description to obtain the video's information vector; Using a multi-objective recommendation model, the target recommendation priority of the video under the constraints of multiple business objectives is obtained based on the information vector of the video and the interest model.
[0091] Based on the above embodiments, as an optional embodiment, a multi-objective recommendation model is used to obtain the recommendation priority of the video under the constraints of multiple business objectives according to the information vector of the video and the interest model, including: For each video, the information vector of the video is fused with the decision input vector to obtain the content enhancement vector of the video; The content enhancement vector of the video and the interest model are used as state information and respectively input into the reward functions of the multi-objective recommendation model to obtain the reward value of the state information for each business objective under different recommendation priorities; The recommendation priority that maximizes the sum of the reward values of all business objectives based on the aforementioned status information is used as the target recommendation priority.
[0092] As can be seen from the above embodiments, the decision input vector of this application is a multimodal vector that integrates video information, video recommendation strategy and the video interaction context, while the target video description is semantic description information optimized by the recommendation target. By performing feature fusion on the vector of the target video description, that is, the information vector and the decision input vector, the video content can be represented more comprehensively, more accurately and more in line with the recommendation (that is, it conforms to the recommendation strategy and does not conform to the interests of the target object).
[0093] Single-objective recommendation models focus on only one business objective (e.g., only click-through rate), which may lead to one-sided recommendation results. The multi-objective recommendation model in this application embodiment considers multiple business objectives simultaneously. Different reward functions are designed for different business objectives, which can quantify the contribution (reward value) of state information to each business objective under different recommendation priorities, and more comprehensively evaluate the recommendation effect. Furthermore, in order to ensure global optimality, this application embodiment aims to maximize the sum of reward values, comprehensively weighing all business objectives, avoiding the situation where excessive pursuit of a single business objective harms other business objectives, making the recommendation results more in line with the overall business needs, and improving the overall performance and user experience of the recommendation system.
[0094] The sum of reward values of the multi-objective recommendation model in this application embodiment R ( s , a This can be expressed by the following formula: R ( s , a )=
[0095] Where s represents the video's content enhancement vector and interest model, a This refers to the action of determining the recommendation priority of videos. Indicates the first i A reward function The weights are assigned to each other, and the sum of all weights is 1.
[0096] Understandably, the total reward value R ( s , a In state s, after action a is performed, the feedback given by the environment is the result of the action. Each reward function is a function related to the current state s, a, and the result of the action, and measures the reward value of a business goal. This application's embodiments represent the total reward value as a weighted sum of multiple reward functions, enabling a comprehensive and flexible consideration of multiple business objectives. Each reward function corresponds to a specific business objective, such as user click-through rate, viewing time, content novelty, etc., and can accurately measure the reward value of that business objective under a given state (s, the combination of video content enhancement vector and user interest model) and action (a, determining recommendation priority). Weight α iThe settings assign different levels of importance to different business objectives. This design allows the model to comprehensively weigh multiple business objectives, avoiding the limitations of a single-objective approach and improving the comprehensiveness and rationality of the recommendation results. Simultaneously, determining the recommendation priority based on maximizing the total reward value ensures the global optimality of the recommendation strategy. This provides users with video recommendations that better match their interests and needs and contribute to achieving the platform's business goals, enhancing user experience and platform competitiveness. It has high practical value and promising prospects in the field of video recommendation.
[0097] In some embodiments, the recommendation priority can be a numerical value between 0 and 100, with a higher value indicating a higher recommendation priority for the video. This method is simple and intuitive, facilitating sorting and comparison. For example, when video A has a recommendation priority of 90 and video B has a priority of 60, it is immediately clear that video A ranks higher than video B in the recommendation ranking. In some embodiments, the recommendation priority can also be a probability distribution of the ranking results, for example, the probability of ranking first is 90%, the probability of ranking second is 78%, and so on.
[0098] Based on the above embodiments, as an optional embodiment, the business objective is related to at least one of the following: 1) The probability and degree to which the target object triggers a preset interactive behavior on the video.
[0099] The preset interactive behaviors in this application embodiment can include clicking, liking, commenting, sharing, and saving. The business objective focuses on predicting the probability that a target audience will trigger these interactive behaviors with a video, while also further evaluating the level of interaction, such as the number of words in a comment, the frequency of sharing, and the duration of viewing. From the target audience's perspective, this allows for the recommendation of videos that better match their interests and needs and are more likely to elicit interaction, improving the user experience and enhancing audience stickiness to the platform. From the platform's perspective, videos with high probability and intensity of interaction often bring higher traffic and exposure, which is beneficial for the dissemination and conversion of videos and promotional information (such as advertisements), increasing the platform's revenue and commercial value.
[0100] 2) The degree of difference between the videos in the recommendation list across multiple content dimensions.
[0101] Content dimensions include video themes (e.g., technology, food, sports), styles (e.g., humor, serious, artistic), and formats (e.g., animation, live-action, edited videos). The business objective is to ensure that the videos in the recommendation list have a certain degree of diversity across these dimensions, avoiding overly monotonous content. For example, a recommendation list might contain live-action technology videos, animated food videos, and edited sports videos. For the target audience, diverse content can satisfy their different interests and needs, preventing viewer fatigue and allowing them to obtain richer information from a single recommendation list.
[0102] 3) The novelty of each video in the recommendation list to the target object.
[0103] Novelty primarily considers whether the video's content, format, and viewpoint are unfamiliar or rarely encountered by the user. This application's embodiments analyze the user's browsing and collection history to determine their familiarity with different types of videos, thereby assessing the novelty of new videos relative to the user. For example, for users who frequently watch traditional food preparation videos, recommending food videos that incorporate modern technology or innovative cooking methods would result in a higher degree of novelty. The business objective is to bring novelty and surprise to the user, stimulating their desire to explore and increasing their time spent and activity on the platform. Simultaneously, novel videos may also lead to new trends and create new hot topics and discussions for the platform, enhancing its innovation capabilities and industry influence.
[0104] Based on the above embodiments, as an optional embodiment, this application embodiment further includes: Obtain feedback information on recommending the video to the target object based on the target recommendation priority; Based on the feedback information, at least the model parameters of the granular decision model and the multi-objective recommendation model should be adjusted.
[0105] This application embodiment, after recommending videos to target objects according to target recommendation priority, collects feedback information from the target objects regarding the videos. This application embodiment can set up multiple feedback collection methods during the object's browsing of recommended videos. On the one hand, prominent interactive buttons, such as likes, dislikes, and comments, can be set on the video playback interface to record the object's direct evaluation of the video; on the other hand, indirect feedback can be obtained by analyzing the object's behavioral data, such as the video's complete playback rate, number of repeated playbacks, and playback duration. Simultaneously, a short questionnaire can pop up when the object exits browsing, asking the object about their overall satisfaction with the recommended videos and their feelings about the level of detail in the video descriptions.
[0106] In some embodiments, if an object has a high dislike rate for a video, comments mention that the video is unclear, and the complete playback rate is low, it indicates that the level of detail in the video description may be inappropriate and misleading to the object. In this case, the weight of parameters related to increasing content complexity in the granular decision model can be increased, making the granular decision model tend to generate a more detailed description. Conversely, if the object has a high like rate and a high complete playback rate, the weight of the relevant parameters can be appropriately reduced to avoid the description being too lengthy.
[0107] In some embodiments, if an object interacts frequently with a video, the reward function weight of that type of video in the recommendation priority calculation for relevant business objectives (such as click-through rate and viewing time) can be increased. If an object frequently complains about insufficient diversity in the recommendation list, the parameters measuring video differences in the model can be adjusted to increase the recommendation priority of different types of videos.
[0108] Please see Figure 8 The figure exemplifies a flowchart illustrating the process of optimizing video recommendation effects based on video descriptions, as provided in an embodiment of this application. Compared to... Figure 7 In the embodiment shown, after obtaining the target recommendation priority of the video, the present application further recommends the video to the target object according to the target recommendation priority, then obtains the target object's feedback information on the video, and adjusts the model parameters of the summary generation model and the multi-target recommendation model according to the feedback information.
[0109] The embodiments of this application form a closed loop of continuous learning and self-evolution. The entire system can dynamically adjust its internal logic based on real user interaction data, thereby continuously improving its performance in both long and short video recommendation scenarios. In particular, it can effectively solve the cold start problem and significantly enhance the exploratory nature and diversity of content.
[0110] This application provides a method for improving video recommendation performance, including: S1. Obtain key information from the video, including video information, video interaction context, and video recommendation strategy; Video Information F video This includes video metadata information, such as video duration, video category (movie, TV series, variety show, short drama, UGC short video, etc.), and existing tag / keyword embedding vectors, etc. Video Interaction Context F user_ctx This is used to capture the user's current interaction context, such as the page the user is currently on (short video feed, long video detail page, search results page, etc.), and the preference for content granularity or type reflected in the user's most recent N interaction behaviors (clicks, views, dwell time). This can be an embedding vector based on user behavior sequence modeling.
[0111] Video recommendation strategy F rec_goal This represents the main optimization objective of the current recommendation system. For example, a learnable vector can be used to represent different objectives such as "emphasizing diversity," "emphasizing precise matching," and "cold start exploration."
[0112] S2, video information, video interaction context, and video recommendation strategy are each mapped to a unified dimension by their respective encoders (such as embedding layers or small MLPs), and then fused to obtain a comprehensive decision input vector x. GDU : x GDU =[Embed(F video Embed(F) user_ctx Embed(F) rec_goal )] Here, Embed() represents the feature embedding function, and [;] represents the vector concatenation operation.
[0113] S3. Input the decision input vector xGDU into a granular decision network f. GDU (For example, it could be a two- or three-layer fully connected network with a ReLU activation function), granular decision network f GDU Learning how to infer the most appropriate descriptive granularity from complex decision input vectors xGDU, and granular decision network f GDU The output can be a continuous granularity score Sgranularity∈[0,1], or a probability distribution at a predefined granularity level. This application's embodiments prefer the former because it provides more flexible control.
[0114] S granularity =sigmoid(W2×ReLU(W1×x GDU +b1)+b2) Where W1, b1, W2, and b2 are the learnable parameters of the network, and the sigmoid function reduces the output to between 0 and 1. The higher the value of Sgranularity, the finer and more detailed the description is required.
[0115] S4. Input Sgranularity into the adaptive description decoder. The adaptive description decoder adjusts the description strategy, including at least one of the following: Control the length of the generated description: Length ∝ S granularity Adjusting the focusing range of the spatiotemporal attention mechanism: lower S granularity This may allow attention to be focused more on high-level global features, generating a macroscopic description; a higher S granularity This will guide attention to delve deeper into the time-series details and local events within the video.
[0116] Activate different generation templates or vocabulary: For different granularities, the decoder may be guided to use different language styles or keywords, such as "event description vocabulary" for high scores and "generalization vocabulary" for low scores.
[0117] If S granularity This means generating a high-level, concise description, while the decoder focuses on extracting the core information from the video; If S granularityTo generate hierarchical, fine-grained descriptions, the decoder incorporates a spatiotemporal attention mechanism. This not only focuses on key objects within video frames but also tracks their changes and interactions over time, resulting in descriptions that include information such as chronological order, location transitions, and event progression. For example, for long videos, chapter-level descriptions can be generated first, followed by event-level descriptions for each chapter.
[0118] Ultimately, the decoder will output a set of one or more initial candidate video descriptions.
[0119] S5. The generated set of candidate video descriptions is closely integrated with the recommendation target and incorporated into the content representation of the video.
[0120] S51, Constructing an Interest Model U interest This includes both long-term and short-term interests of the object, and represents the interests through vector embedding; Semantic optimization and filtering: This application does not directly use the entire set of candidate video descriptions, but instead performs semantic optimization and filtering on the set of candidate video descriptions: A. Object Interest Matching Degree: Measures the relevance of each candidate description di to the interest model U. interest Matching degree M (Uinterest,di) For example, if the target audience likes "funny" content, this application embodiment will prioritize selecting descriptions containing words such as "hilarious" and "humorous plot"; B. Recommendation Diversity: To avoid homogenization of recommendations, this application embodiment evaluates the semantic distance SD between each pair of candidate descriptions. (di,dj) For example, by calculating the cosine distance of the descriptive text embedding, the goal is to select or recombine descriptive combinations that can both match the object's interests and ensure the diversity of recommended content. This optimization process can be viewed as a problem of finding the optimal combination of descriptions, which can be solved using reinforcement learning or heuristic search algorithms. Its goal is to maximize the balance between matching degree and diversity. For example:
[0121] Where λ is a weighting coefficient that balances matching degree and diversity.
[0122] Ultimately, a set of filtered video descriptions, Doptimized, was obtained.
[0123] S52. Encode Doptimized into a high-dimensional description vector V. desc ; S53, Describe the vector V desc With decision input vector x GDU By merging, an enhanced content representation V can be obtained.enhanced ; S6. Utilizing the interest model U interest and enhanced content representation V enhanced Through efficient recall algorithms such as the dual-tower model, candidate video sets that may be related to the object can be quickly filtered from a massive video library.
[0124] In the sorting stage, this application's embodiments introduce multi-objective reinforcement learning to dynamically optimize the recommendation list. Accuracy: Click-through rate (CTR), viewing time.
[0125] Diversity: The variety of video types, themes, and styles in the recommended list.
[0126] Novelty: Recommended videos that the target audience has not seen before or that are newly launched on the platform.
[0127] Audience satisfaction: This is determined by the audience's subsequent interactive behaviors, such as likes, comments, shares, and return visit rates.
[0128] MRL learns a policy π(a|s), where s are the features of the object-video pair (including V). enhanced and object U interest ), 'a' is the action of sorting the videos.
[0129] The reward function R(s,a) will combine the above multiple objectives:
[0130] Where α, β, γ, δ are weighting coefficients, R CTR R represents the reward value associated with click-through rate. watchtime R represents the reward value associated with viewing time. Diversity R represents the reward value related to diversity. Novelty This represents the reward value associated with novelty. In this way, the recommender system can adaptively balance different recommendation objectives based on platform policies and user feedback.
[0131] Furthermore, MRL not only optimizes the final recommendation ranking, but the object feedback signals it learns (e.g., clicks / non-clicks, viewing duration, etc. related to an object's recommendation for a specific description) can also be backpropagated, influencing and adjusting in real time. The granularity decision unit determines the granularity of the description.
[0132] The matching degree calculation and diversity balancing are described in the semantic optimization and filtering module.
[0133] In this way, description generation, description optimization, video content representation enhancement, and final recommendation ranking form a closed loop of continuous learning and self-evolution. The entire system can dynamically adjust its internal logic based on real-world object interaction data, thereby continuously improving performance in both long and short video recommendation scenarios. In particular, it effectively solves the cold start problem and significantly enhances the exploratory nature and diversity of content.
[0134] This application provides a video processing apparatus, such as... Figure 9 As shown, the video processing device may include: an information acquisition module 901, a granularity determination module 902, and a description generation module 903, wherein, The information acquisition module 901 is used to obtain the video recommendation strategy, the video interaction context of the target object, and the video information of the video; Granularity determination module 902 is used to input the video information, the video recommendation strategy and the video interaction context into a pre-trained granularity decision model to obtain the granularity information of the video, wherein the granularity information is used to represent the level of detail of the video description; The description generation module 903 is used to generate at least one candidate video description that matches the granularity information based on the video information and the granularity information.
[0135] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.
[0136] Based on the above embodiments, as an optional embodiment, when there are multiple candidate video descriptions, the video processing device further includes a description filtering module; The profile filtering module is used for: Obtain the interest model of the target object; Calculate the matching degree between the description of each candidate video and the interest model; Based on the matching degree corresponding to each candidate video description, the target video description of the video is determined from at least one candidate video description.
[0137] Based on the above embodiments, as an optional embodiment, the description filtering module determines the target video description of the video from at least one candidate video description according to the matching degree corresponding to each candidate video description, including: Identify at least one combination of descriptions, each combination of descriptions including at least two candidate video descriptions; For each combination of descriptions, the difference between a first value and a second value is determined. The first value is the sum of the matching degrees of all candidate video descriptions in the combination, and the second value is the sum of the semantic similarities between any two candidate video descriptions in the combination. The candidate video descriptions in the target description combination with the greatest differences are merged to obtain the target video description.
[0138] Based on the above embodiments, as an optional embodiment, the interest model includes multiple keywords, including at least one first keyword and at least one second keyword, wherein the first keyword is used to represent the long-term interests of the target object, and the second keyword is used to reflect the short-term interests of the target object; The description filtering module calculates the matching degree between each candidate video description and the interest model in the following way: Each keyword in the interest model is converted into a word vector to obtain the interest word vector. For each candidate video description, the candidate video description is segmented into words, and each segment is converted into word vectors to obtain multiple description word vectors; For each description word vector of each candidate video description, the similarity between the description word vector and each keyword vector is determined, and the maximum similarity is taken as the target similarity of the description word vector. According to the video recommendation strategy and the type of the keyword corresponding to the target similarity, the weight of the description word vector is determined; the type refers to whether the keyword belongs to the first keyword or the second keyword. For each candidate video description, the target similarity of each description word vector is weighted and summed according to the weight of each description word vector in the candidate video description to obtain the matching degree between the candidate video description and the interest model.
[0139] Based on the above embodiments, as an optional embodiment, the granularity determination module inputs the video information, video recommendation strategy, and video interaction context into a pre-trained granularity decision model to obtain granularity information, including: The video information, video recommendation strategy, and video interaction scenario are respectively input into the encoder group of the summary generation model to obtain the feature vectors of the video information, video recommendation strategy, and video interaction scenario respectively. The feature vectors of the video information, video recommendation strategy, and video interaction context are fused to obtain the decision input vector of the video. The decision input vector is then input into the granular decision model in the summary generation model to obtain the granular information. The description generation module generates at least one candidate video description that matches the granularity information based on the video information and granularity information, including: The granular information and feature vector group are input into the decoder of the description generation model to obtain at least one candidate video description for the video. The introductory generation model includes a cascaded encoder group, a granular decision model, and a decoder; the feature vector group includes at least one of the feature vectors of the video information, the video recommendation strategy, and the video interaction context, and includes at least the feature vector of the video information.
[0140] Based on the above embodiments, as an optional embodiment, the description generation model is trained in the following manner: Multiple training samples and corresponding training labels are obtained. The training samples include sample video information, sample video recommendation strategies, and sample video interaction scenarios. The training labels include sample video descriptions corresponding to the sample video information and the actual granularity information of the sample video descriptions. For each training sample, the training sample is input into the initial encoder group to obtain the feature encoding of the sample video information, the sample video recommendation strategy and the sample video interaction context respectively. The feature encoding of the sample video information, the sample video recommendation strategy and the sample video interaction context is fused and then input into the initial granularity decision model to obtain prediction granularity information. Input the prediction granularity information and feature vector group into the initial decoder to obtain at least one predicted video summary; A first loss value is obtained based on the difference between the predicted granularity information and the actual granularity information; The differences between each predicted video description and the actual video description in terms of word segmentation and syntax are determined to obtain a second loss value; The total loss function value is obtained based on the first loss value and the second loss value. The initial introduction generation model is trained based on the total loss function value until the total loss function converges. The initial introduction generation model at the end of training is taken as the introduction generation model.
[0141] Based on the above embodiments, as an optional embodiment, the video processing apparatus further includes a recommendation priority determination module, which is used for: Obtain the interest model of the target object; The target video description of the video is feature-encoded to obtain the information vector of the video; Using a multi-objective recommendation model, the target recommendation priority of the video under the constraints of multiple business objectives is obtained based on the information vector of the video and the interest model.
[0142] Based on the above embodiments, as an optional embodiment, the recommendation priority determination module utilizes a multi-objective recommendation model to obtain the recommendation priority of the video under the constraints of multiple business objectives based on the information vector of the video and the interest model, including: For each video, the information vector of the video is fused with the decision input vector to obtain the content enhancement vector of the video; The content enhancement vector of the video and the interest model are used as state information and input into each reward function in the multi-objective recommendation model to obtain the reward value of the state information for each business objective under different recommendation priorities; The recommendation priority that maximizes the sum of the reward values of all business objectives based on the aforementioned status information is used as the target recommendation priority.
[0143] Based on the above embodiments, as an optional embodiment, the business objective is related to at least one of the following: The probability and degree to which the target object triggers a preset interactive behavior in response to the video; The degree of difference between the videos in the recommendation list across multiple content dimensions; The novelty of each video in the recommendation list to the target object.
[0144] Based on the above embodiments, as an optional embodiment, the video processing device further includes a feedback adjustment module, which is used to: obtain feedback information on recommending the video to a target object according to the target recommendation priority; and adjust the model parameters of at least the granular decision model and the multi-target recommendation model according to the feedback information.
[0145] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of a video processing method. Compared with related technologies, this method can achieve the following: by acquiring the video interaction context of the target object, the video recommendation strategy, and the video information of the video, the generation of the video description is deeply coupled with the video recommendation strategy of the recommendation system and the video interaction context of the target object. Furthermore, instead of directly generating the video description, before generating the video description, the granularity information of the video is intelligently determined through a granular decision model based on the video information, video interaction context, and video recommendation strategy. The granularity information represents the level of detail of the video description, thereby providing suggestions on the number of words and focus of the video description. Finally, at least one candidate video description matching the granularity information is generated based on the video information and the granularity information. The generated video description is more likely to be prioritized by the video recommendation system and is also more in line with the preferences of the target object, effectively increasing the probability that the target object will find videos that meet its needs.
[0146] In one alternative embodiment, an electronic device is provided, such as Figure 10 As shown, Figure 10 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.
[0147] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0148] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, bus 4002 is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus.
[0149] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.
[0150] The memory 4003 stores computer programs that execute embodiments of this application, and its execution is controlled by the processor 4001. The processor 4001 executes the computer programs stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0151] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.
[0152] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0153] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the illustrations or text descriptions.
[0154] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.
[0155] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.
Claims
1. A video processing method, characterized in that, include: Obtain video recommendation strategies, the video interaction context of target objects, and video information of the videos; The video information, the video recommendation strategy, and the video interaction context are input into a pre-trained granular decision model to obtain the granular information of the video, which is used to represent the level of detail in the video description. Based on the video information and the granularity information, at least one candidate video description matching the granularity information is generated.
2. The method according to claim 1, characterized in that, When there are multiple candidate video descriptions, the method further includes: Obtain the interest model of the target object; Calculate the matching degree between the description of each candidate video and the interest model; Based on the matching degree corresponding to each candidate video description, the target video description of the video is determined from at least one candidate video description.
3. The method according to claim 2, characterized in that, The step of determining the target video description of the video from at least one candidate video description based on the matching degree corresponding to each candidate video description includes: Identify at least one combination of descriptions, each combination of descriptions including at least two candidate video descriptions; For each combination of descriptions, the difference between a first value and a second value is determined. The first value is the sum of the matching degrees of all candidate video descriptions in the combination, and the second value is the sum of the semantic similarities between any two candidate video descriptions in the combination. The candidate video descriptions in the target description combination with the greatest differences are merged to obtain the target video description.
4. The method according to claim 2, characterized in that, The interest model includes multiple keywords, including at least one first keyword and at least one second keyword. The first keyword is used to represent the long-term interests of the target object, and the second keyword is used to reflect the short-term interests of the target object. The matching degree between each candidate video description and the interest model is calculated in the following way: Each keyword in the interest model is converted into a word vector to obtain the interest word vector. For each candidate video description, the candidate video description is segmented into words, and each segment is converted into word vectors to obtain multiple description word vectors; For each description word vector of each candidate video description, the similarity between the description word vector and each keyword vector is determined, and the maximum similarity is taken as the target similarity of the description word vector. According to the video recommendation strategy and the type of the keyword corresponding to the target similarity, the weight of the description word vector is determined; the type refers to whether the keyword belongs to the first keyword or the second keyword. For each candidate video description, the target similarity of each description word vector is weighted and summed according to the weight of each description word vector in the candidate video description to obtain the matching degree between the candidate video description and the interest model.
5. The method according to claim 1, characterized in that, The step of inputting the video information, video recommendation strategy, and video interaction context into a pre-trained granular decision model to obtain granular information includes: The video information, video recommendation strategy, and video interaction scenario are respectively input into the encoder group of the summary generation model to obtain the feature vectors of the video information, video recommendation strategy, and video interaction scenario respectively. The feature vectors of the video information, video recommendation strategy, and video interaction context are fused to obtain the decision input vector of the video. The decision input vector is then input into the granular decision model in the summary generation model to obtain the granular information. The step of generating at least one candidate video description matching the granularity information based on the video information and granularity information includes: The granular information and feature vector group are input into the decoder of the description generation model to obtain at least one candidate video description for the video. The introductory generation model includes a cascaded encoder group, a granular decision model, and a decoder; the feature vector group includes at least one of the feature vectors of the video information, the video recommendation strategy, and the video interaction context, and includes at least the feature vector of the video information.
6. The method according to claim 5, characterized in that, The description generation model was trained in the following way: Multiple training samples and corresponding training labels are obtained. The training samples include sample video information, sample video recommendation strategies, and sample video interaction scenarios. The training labels include sample video descriptions corresponding to the sample video information and the actual granularity information of the sample video descriptions. For each training sample, the training sample is input into the initial encoder group to obtain the feature encoding of the sample video information, the sample video recommendation strategy and the sample video interaction context respectively. The feature encoding of the sample video information, the sample video recommendation strategy and the sample video interaction context is fused and then input into the initial granularity decision model to obtain prediction granularity information. Input the prediction granularity information and feature vector group into the initial decoder to obtain at least one predicted video summary; A first loss value is obtained based on the difference between the predicted granularity information and the actual granularity information; The differences between each predicted video description and the actual video description in terms of word segmentation and syntax are determined to obtain a second loss value; The total loss function value is obtained based on the first loss value and the second loss value. The initial introduction generation model is trained based on the total loss function value until the total loss function converges. The initial introduction generation model at the end of training is taken as the introduction generation model.
7. The method according to claim 5, characterized in that, The method further includes: Obtain the interest model of the target object; The target video description of the video is feature-encoded to obtain the information vector of the video; Using a multi-objective recommendation model, the target recommendation priority of the video under the constraints of multiple business objectives is obtained based on the information vector of the video and the interest model.
8. The method according to claim 7, characterized in that, The step of using a multi-objective recommendation model to obtain the recommendation priority of the video under the constraints of multiple business objectives based on the video's information vector and the interest model includes: For each video, the information vector of the video is fused with the decision input vector to obtain the content enhancement vector of the video; The content enhancement vector of the video and the interest model are used as state information and input into each reward function in the multi-objective recommendation model to obtain the reward value of the state information for each business objective under different recommendation priorities; The recommendation priority that maximizes the sum of the reward values of all business objectives based on the aforementioned status information is used as the target recommendation priority.
9. The method according to claim 7, characterized in that, The business objective relates to at least one of the following: The probability and degree to which the target object triggers a preset interactive behavior in response to the video; The degree of difference between the videos in the recommendation list across multiple content dimensions; The novelty of each video in the recommendation list to the target object.
10. The method according to claim 5, characterized in that, Also includes: Obtain feedback information on recommending the video to the target object based on the target recommendation priority; Based on the feedback information, at least the model parameters of the granular decision model and the multi-objective recommendation model should be adjusted.
11. A video processing apparatus, characterized in that, include: The information acquisition module is used to obtain video recommendation strategies, the video interaction context of target objects, and video information of the videos; The granularity determination module is used to input the video information, the video recommendation strategy, and the video interaction context into a pre-trained granularity decision model to obtain the granularity information of the video, wherein the granularity information is used to represent the level of detail of the video description; The description generation module is used to generate at least one candidate video description that matches the granularity information based on the video information and the granularity information.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the video processing method according to any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the video processing method according to any one of claims 1-10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video processing method according to any one of claims 1-10.