Video recommendation method and device, storage medium and electronic equipment
By analyzing the behavior feedback results between the target object and the reference video, filtering and labeling high-potential explosive points in long videos, and recommending them based on video tags and description information, the problems of low labeling of long videos and poor video recommendations in the existing technology are solved, and more efficient video recommendations are achieved.
Patent Information
- Application Number
- CN202510182559.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology has low efficiency in labeling of explosive points for medium and long videos and poor video recommendation effects.
By determining the interaction value based on the behavior feedback results between the target object and multiple reference videos, high-potential target videos are selected, and video tag information and description information matching the target videos are determined from the recommendation pool.
It improves the efficiency of long video explosive point labeling and video recommendation effects, and can more accurately understand video content and user interests.
Smart Images

Figure CN120104830A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video recommendation, and in particular to a video recommendation method and device, a storage medium, and an electronic device. Background Art
[0002] Hot spots are a form of short videos. Compared with the traditional short videos that are complete, independent and carefully produced, hot spots are essentially based on a larger long video content system. The difference between hot spots and independently created short videos is that hot spots are more of a strategy for content extraction and presentation, and are not directly equivalent to a complete video work. Hot spots are attached to the content of long videos. By accurately marking a certain starting point and ending point in the long video, the most attractive and most interesting or hotly discussed clips in the video are highlighted. The playback content is essentially still part of the original long video.
[0003] Identifying and labeling hot spots in long videos is a complex and challenging technical problem. It not only requires high accuracy, but also requires a huge amount of data to be processed. In addition, how to use hot spots to drive the viewing and recommendation of long videos has become an urgent problem to be solved. In other words, the existing technology has the technical problems of low efficiency in labeling hot spots in long videos and poor video recommendation effect.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present application provide a video recommendation method and device, a storage medium, and an electronic device to at least solve the technical problems in the prior art of low efficiency in marking hot spots in long videos and poor video recommendation effect.
[0006] According to one aspect of an embodiment of the present application, a video recommendation method is provided, comprising: determining an interaction value of each of multiple reference videos based on behavioral feedback results between a target object and the multiple reference videos; determining at least one target video that meets a screening condition from the multiple reference videos based on the interaction value; determining video tag information that matches the at least one target video, and description information that matches each marked segment in the at least one target video, wherein the marked segment is a segment that generates interactive behavior with the target object, and the description information includes tag description information and evaluation index information, and the evaluation index information is used to indicate the degree of influence of the marked segment on the video viewing time; determining a target recommended video that matches the target object from a recommendation pool based on the video tag information and the description information.
[0007] According to another aspect of an embodiment of the present application, a video recommendation device is also provided, including: a first determination unit, determining the interaction value of each of multiple reference videos according to the behavior feedback results between the target object and the multiple reference videos; a screening unit, determining at least one target video that meets the screening conditions from the multiple reference videos according to the interaction value; a second determination unit, determining video tag information that matches the at least one target video respectively, and description information that matches each marked segment in the at least one target video, wherein the marked segment is a segment that generates interactive behavior with the target object, and the description information includes tag description information and evaluation index information, and the evaluation index information is used to indicate the degree of influence of the marked segment on the video viewing time; a recommendation unit, determining a target recommended video that matches the target object from a recommendation pool according to the video tag information and the description information.
[0008] Optionally, the second determination unit is further used to determine the video tag that matches the target video and the fragment tag that matches the marked fragment; determine the video tag score indicated by the video tag information according to the first tag hit rate, wherein the first tag hit rate is used to indicate the degree to which the target object interacts with the video type corresponding to the video tag; determine the fragment tag score indicated by the description information according to the second tag hit rate, wherein the second tag hit rate is used to indicate the degree to which the target object interacts with the fragment type corresponding to the fragment tag; determine the viewing conversion rate indicated by the evaluation index information according to the ratio of the total playback time matching the marked fragment to the number of viewing users.
[0009] Optionally, the second determination unit includes: a matching module, configured to use a generalized video tag matching the target video as a video tag matching the target video; determine a segment refinement tag matching the marked segment, and use a combination of the segment refinement tags as a segment tag matching the marked segment.
[0010] Optionally, the above-mentioned second determination unit includes: a calculation module, used to determine a first weight value of a first feedback factor matching the target video according to a first label hit rate, and to determine a second weight value of a second feedback factor matching the marked segment according to a second label hit rate, wherein the feedback factor is a behavior parameter corresponding to the interactive behavior generated by the target object; determine a video label score according to a cumulative exposure value matching the target video and a first evaluation value, wherein the first evaluation value is the product of a first weight value and a first feedback factor; determine a segment label score according to a cumulative exposure value matching the marked segment and a second evaluation value, wherein the second evaluation value is the product of a second weight value and a second feedback factor.
[0011] Optionally, the above-mentioned recommendation unit is also used to obtain a target video tag whose video tag score is greater than a first threshold, a first target segment tag whose segment tag score is greater than a second threshold, and a second target segment tag whose viewing conversion rate is greater than a third threshold; determine the target recommendation description information according to the product of the target video tag, the first target segment tag and the second target segment tag and the weight coefficients that match them respectively; determine the target recommended video that matches the target recommendation description information from the recommendation pool.
[0012] Optionally, the first determination unit is further used to obtain the interaction behavior parameters of the target object in different time periods and the weight values matching therewith; and determine the interaction value of the reference video according to the product of the interaction behavior parameters in different time periods and the weight values matching therewith.
[0013] Optionally, the above-mentioned first determination unit also includes: a third determination module, used to determine a reference video whose interaction value is greater than a target threshold as a target video; determine interaction description information that matches the target video, wherein the interaction description information is used to indicate the interaction position of the target object with respect to the target video; and determine a marked segment that matches the target video based on the interaction description information.
[0014] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned video recommendation method when running.
[0015] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the video recommendation method through the computer program.
[0016] In an embodiment of the present application, the interaction value of each of the multiple reference videos is determined according to the behavioral feedback results between the target object and the multiple reference videos; at least one target video that meets the screening conditions is determined from the multiple reference videos according to the interaction value, thereby helping the system to quickly screen out long videos with high potential; video tag information that matches the at least one target video is determined, and description information that matches each marked segment in the at least one target video is determined, so as to more accurately understand the video content and user interests; and the target recommended video that matches the target object is determined from the recommendation pool according to the video tag information and the description information, thereby solving the technical problems in the prior art of low efficiency in marking hot spots in long videos and poor video recommendation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 is a schematic diagram of an application environment of an optional video recommendation method according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of an optional video recommendation method according to an embodiment of the present application;
[0020] Figure 3 is a flowchart of another optional video recommendation method according to an embodiment of the present application;
[0021] Figure 4 is a flowchart of another optional video recommendation method according to an embodiment of the present application;
[0022] Figure 5 is a schematic structural diagram of an optional video recommendation device according to an embodiment of the present application;
[0023] Figure 6 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] According to one aspect of the embodiments of the present application, a video recommendation method is provided. Optionally, as an optional implementation, the video recommendation method can be but is not limited to being applied to: Figure 1 in the environment shown.
[0027] The server 112 includes a database 114 and a processing engine 116. The server 112 executes S102 to determine the interaction value of each of the multiple reference videos according to the behavior feedback results between the target object and the multiple reference videos;
[0028] S104, determining at least one target video that meets the screening condition from the multiple reference videos according to the interaction value;
[0029] S106, determining video tag information respectively matching at least one target video, and description information respectively matching each marked segment in at least one target video, wherein the marked segment is a segment that generates an interactive behavior with the target object, and the description information includes tag description information and evaluation index information, and the evaluation index information is used to indicate the degree of influence of the marked segment on the video viewing time;
[0030] S108: Determine a target recommended video matching the target object from the recommendation pool according to the video tag information and the description information.
[0031] The server 112 executes S110 to the terminal device 102 through the network 110 to push the target recommended video; the terminal device 102 includes a display 108, a processor 106, and a memory 104. Finally, the terminal device 102 executes S112 to display the target recommended video.
[0032] Optionally, in this embodiment, the terminal device may be a terminal device configured with a target client, which may include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop, a tablet computer, a PDA, a MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, etc. The target client may be a video client, an instant messaging client, a browser client, an education client, etc. The network may include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The server may be a single server, or a server cluster consisting of multiple servers, or a cloud server. The above is only an example, and this embodiment does not impose any limitation on this.
[0033] Optionally, as an optional implementation, as Figure 2 As shown, the above video recommendation method includes:
[0034] S202, determining the interaction value of each of the multiple reference videos according to the behavior feedback results between the target object and the multiple reference videos;
[0035] S204, determining at least one target video that meets the screening condition from the multiple reference videos according to the interaction value;
[0036] S206, determining video tag information that matches the at least one target video, and description information that matches each marked segment in the at least one target video, wherein the marked segment is a segment that generates an interactive behavior with the target object, and the description information includes tag description information and evaluation index information, and the evaluation index information is used to indicate the degree of influence of the marked segment on the video viewing time;
[0037] S208: Determine a target recommended video matching the target object from the recommendation pool according to the video tag information and the description information.
[0038] In step S202, the interaction values of the multiple reference videos are determined according to the behavioral feedback results between the target object and the multiple reference videos; the above-mentioned reference videos are, for example, long videos that the recommended users have watched and interacted with in a historical time period, and may also have similar attributes; for example, the historical time period may be weekends and weekdays from 8 pm to 10 pm; the above-mentioned behavioral feedback results may specifically be the number of likes and replay levels of the reference videos by users, and the feedback results may be positive feedback or negative feedback. It can be understood that positive feedback results have a positive correlation with the interaction value of the reference video, and negative feedback results reduce the interaction value of the reference video; the interaction value directly expresses the user's liking and willingness to share, and these videos may contain hot spots that are more attractive to the target object.
[0039] In step S204, at least one target video that meets the screening conditions is determined from multiple reference videos based on the interaction values; the above screening conditions include but are not limited to: selecting videos ranked in the top X% of the interaction values; setting a time window to screen out videos whose interaction values have continued to grow in the recent period (such as the past month); adjusting the calculation method of the interaction value based on the release time of the video and the user's active time, and screening content with higher interaction values during the user's active time period; combining the interaction value with the tag preference to obtain the screening conditions, etc. The above target video can be a long video that has a certain appeal to users and can hit the mark.
[0040] Further executing S206, determining video tag information respectively matching at least one target video, and description information respectively matching each marked segment in at least one target video, wherein the marked segment is a segment that generates an interactive behavior with the target object, and the description information includes tag description information and evaluation index information, and the evaluation index information is used to indicate the degree of influence of the marked segment on the video viewing time;
[0041] Optionally, the above-mentioned video tag information can specifically be a video tag score value of a long video, the above-mentioned marked segment can specifically be a corresponding hot spot, and each marked segment (hot spot) matches an ID number and a starting point and an end point; the above-mentioned interactive behavior can be likes, sharing, dragging the progress bar, adjusting the playback speed, collecting and other behaviors; the above-mentioned tag description information can specifically be a score value of the tag corresponding to the "hot spot"; the above-mentioned evaluation index information can be conversion time (the time conversion brought by the hot spot to the long video), viewing completion rate (a measure of whether the hot spot prompts users to continue watching the full video), user retention time (the average retention time of users on the platform after watching the hot spot), cross-content viewing (whether the user further explored other videos of the same theme, series or type after watching the hot spot of a certain video), etc., and no specific restrictions are made here.
[0042] Optionally, in the above step S208, a target recommended video matching the target object is determined from the recommendation pool according to the video tag information and the description information. It should be noted that when playing the target recommended video to the target object, it can be played directly from the starting position corresponding to a more specific target burst point, or played from the beginning;
[0043] The above step S208 is described in a complete implementation method:
[0044] S1 builds a personalized profile of each user based on the user's historical behavior, preferences, interests, and interaction records.
[0045] S2-1, label each hot spot in multiple dimensions, including content labels (such as stars, emotions, behaviors) and user feedback labels (such as likes, comments, and replays). If a hot spot receives a high number of likes and comments under a star label, the star-related label score will increase. S2-2, label long videos and calculate label scores based on user interaction records on long videos.
[0046] S3, statistics the conversion time of each hot spot, that is, the average time that users continue to watch the full long video after watching the hot spot.
[0047] S4, incorporates the scores of the above three aspects and the conversion time into a comprehensive scoring model. The goal is to maximize user satisfaction with the recommended content while optimizing viewing time and retention rate.
[0048] S5, in the comprehensive scoring model, dynamically adjust the weights of different factors. For example, if the system finds that users are particularly sensitive to celebrity tags, then the weight of celebrity tags can be increased when recommending. Similarly, conversion time should also be given appropriate weight in the recommendation algorithm based on the overall strategy.
[0049] S6, sort the hot spots according to the comprehensive score, and select the top hot spots to recommend to users. When making recommendations, you can consider additional information such as the user's real-time location, time, device type, etc. to further improve the relevance of the recommendation and user experience. In addition, the recommendation system needs to continuously collect user feedback on the recommended content, including click-through rate, viewing time, retention rate, etc., to adjust user portraits, label scores, and model parameters to achieve long-term optimization and improve personalized recommendation effects.
[0050] The overall process can be shown in the flowchart Figure 3 To explain:
[0051] According to S302-1, long video data calculation; and S302-2, long video content understanding; in combination with the above two steps, S304 is performed, manual analysis and decision-making; and S306 is further performed, a hot spot is hit in the long video;
[0052] Based on the database, execute S308, data inversion, perform data inversion in the database, create indexes, and optimize query speed to retrieve and analyze hot spot data more quickly; S310, data recall, based on user preferences, viewing history and other information, recall long video-hot spot combinations that match the user's interests from the database; S312, data sorting, sort the recalled long video-hot spot combinations, and give priority to displaying those that are more in line with the user's interests and have a higher viewing probability; S314, re-arrangement, based on user portraits and real-time user behavior, personalize the sorting results to further optimize the matching degree of recommended content;
[0053] Then, the result of S316 is output to the user, and S318 is executed to calculate the user portrait, which is used for data inversion. It should also be noted that the offline calculated long video-hot spot combination is stored in the database and can be directly pulled in the real-time recommendation and decision-making process.
[0054] Through the above-mentioned implementation mode of the present application, the interaction value of each of the multiple reference videos is determined according to the behavioral feedback results between the target object and the multiple reference videos; at least one target video that meets the screening conditions is determined from the multiple reference videos according to the interaction value, thereby helping the system to quickly screen out long videos with high potential; video tag information that matches the at least one target video respectively, and description information that matches each marked segment in the at least one target video are determined, so as to more accurately understand the video content and user interests; it is achieved that the target recommended video that matches the target object is determined from the recommendation pool according to the video tag information and the description information, thereby solving the technical problems in the prior art of low efficiency in marking hot spots in long videos and poor video recommendation effect.
[0055] In an optional implementation, determining video tag information that matches at least one target video, and description information that matches each marked segment in at least one target video, includes:
[0056] S1, determine the video tags matching the target video and the segment tags matching the marked segments;
[0057] S2, determining a video tag score indicated by the video tag information according to the first tag hit rate, wherein the first tag hit rate is used to indicate the degree to which the target object interacts with the video type corresponding to the video tag;
[0058] S3, determining the segment tag score indicated by the description information according to the second tag hit rate, wherein the second tag hit rate is used to indicate the degree of interaction of the target object with the segment type corresponding to the segment tag;
[0059] S4, determining the viewing conversion rate indicated by the evaluation indicator information according to the ratio of the total playback time matching the marked segment to the number of viewing users.
[0060] In the above step S1, the video tags matching the target video and the segment tags matching the marked segment are determined; specifically, for example, taking movies as an example, the above video tags can be: category (movie), subject (science fiction, adventure, alien exploration), emotion (surprise, excitement, tension), which are intended to describe the general type and emotional color of the entire movie, and are suitable for recommending movies to user groups with a wide range of interests; the above segment tags can be: star tags (Wang X, villain alien leader), emotion and action combination tags: thrilling duel (tension, action), the hero's heroic performance (excitement, action), special attention is paid to the specific character performance, emotional climax and action elements in the video clip, and the user's feedback on this clip, which is only an example here;
[0061] In steps S2-S3, the video tag score indicated by the video tag information is determined according to the first tag hit rate, wherein the first tag hit rate is used to indicate the degree of interaction of the target object with the video type corresponding to the video tag; the fragment tag score indicated by the description information is determined according to the second tag hit rate, wherein the second tag hit rate is used to indicate the degree of interaction of the target object with the fragment type corresponding to the fragment tag.
[0062] As an optional implementation, the video tag score of the target video and the segment tag score of the marked segment may be calculated by the following formula:
[0063]
[0064] Among them, n represents the number of times a certain tag appears, including tags hit by behaviors such as play, click, like, favorite, search, etc. δ will give different weights according to user behavior, for example, the weights of search behavior and click behavior will be higher than the weights of like and share behavior. α will be appropriately adjusted according to different tag types, τ is the decay rate (usually a positive number that affects the speed of decay), Δt is the time difference between the exposure time of the tag and the statistical time, ε is a random factor, β is the adjustment weight for user behavior feedback, and U ij It refers to the positive and negative behavioral feedback from users on videos or hot spots, including collection, likes, comments, dragging the progress bar, speeding up, and replaying.
[0065] Specifically, the positive and negative feedback behavior index is used to measure the user's preference for the content and can be calculated according to the following formula:
[0066]
[0067] Among them, a 1 to a 7 is the weight of each behavior, L is the number of likes, C is the number of comments, H is the number of favorites, S is the number of shares, is n times the speed adjustment, T is the speed, is the number of replays. Is the number of times the progress bar is dragged.
[0068] In the above step S4, the viewing conversion rate indicated by the evaluation indicator information is determined according to the ratio of the total playback time matching the marked segment to the number of viewing users;
[0069] As an optional implementation, the conversion duration of the hot spot to the long video is calculated as follows:
[0070]
[0071] in, is the cumulative sum of n playback durations, and U is the number of users. It should be noted that because the comparison is made between the duration conversions brought about by different hot spots in the same long video, the duration of the long video does not need to be processed in the denominator. It should also be noted that the hot spots on the long video have corresponding ids as well as start and end times.
[0072] In an optional implementation, determining a video tag matching the target video and a segment tag matching the marked segment includes:
[0073] S1, taking the generalized video label that matches the target video as the video label that matches the target video;
[0074] S2, determining a segment refinement label that matches the marked segment, and using a combination of the segment refinement labels as a segment label that matches the marked segment.
[0075] In the above steps S1-S2, the video generalization label matching the target video is used as the video label matching the target video; the segment refinement label matching the marked segment is determined, and the combination of the segment refinement labels is used as the segment label matching the marked segment.
[0076] As an optional implementation, assume that the target video is a classic movie "X". Video generalization tags (as video tags that match the target video): category (movie), subject matter (crime, drama, inspirational), etc. Generalization tags describe the overall type, storyline, emotional color and other characteristics of the movie, helping the system to identify videos with similar attributes to "X" and thus make generalized recommendations.
[0077] Marked segment (hot spot): a classic scene in a movie, such as character A successfully escaping from prison and standing in the rain to celebrate freedom. Segment refinement label (as a segment label that matches the marked segment): star label (Wang X); emotional feeling and action behavior combination label: exciting prison break (tension, action), celebration of freedom (joy, action), other refinement labels: classic lines ("Get busy living, or get busy dying."), high interaction (a large number of users like and comment), etc. The refinement label focuses on the specific characters, emotional climax, action details and other features of the marked segment, helping the system to identify segments with similar attributes to this hot spot, so as to make refined recommendations.
[0078] Through the implementation of the present application, the application of generalized labels of videos combined with refined labels of hot spots in video recommendation not only improves the personalization and accuracy of recommendations, but also increases the exposure of medium and long-tail content, and enhances the interaction between users and the platform, thereby improving the recommendation effect of videos.
[0079] In an optional implementation, after determining the video tag matching the target video and the segment tag matching the marked segment, the method further includes:
[0080] S1, determining a first weight value of a first feedback factor matching a target video according to a first tag hit rate, and determining a second weight value of a second feedback factor matching a marked segment according to a second tag hit rate, wherein the feedback factor is a behavior parameter corresponding to an interactive behavior generated by the target object;
[0081] S2, determining a video tag score according to the exposure accumulation value matched with the target video and a first evaluation value, wherein the first evaluation value is a product of a first weight value and a first feedback factor;
[0082] S3, determining a segment label score according to the exposure accumulation value matching the marked segment and a second evaluation value, wherein the second evaluation value is a product of a second weight value and a second feedback factor.
[0083] In the above step S1, a first weight value of a first feedback factor matching the target video is determined according to a first tag hit rate, and a second weight value of a second feedback factor matching the marked segment is determined according to a second tag hit rate, wherein the feedback factor is a behavior parameter corresponding to the interactive behavior generated by the target object.
[0084] As an optional implementation, the calculation of the hit rate can be based on the frequency and intensity of the user's interaction with the content with these tags. For example, if the user has a high like rate and viewing rate for videos of star X, high-definition special effects, and thrilling actions, the hit rate of the corresponding tag is high; the above feedback factors can be quantitative behavior parameters corresponding to behaviors such as likes, comments, sharing, dragging the progress bar, speed playback, and replay; specifically, the hit rate has a positive correlation with the weight value of the corresponding feedback factor;
[0085] In the above steps S2-S3, the video label score is determined according to the exposure accumulation value matching the target video and the first evaluation value, wherein the first evaluation value is the product of the first weight value and the first feedback factor; the segment label score is determined according to the exposure accumulation value matching the marked segment and the second evaluation value, wherein the second evaluation value is the product of the second weight value and the second feedback factor.
[0086] As an optional implementation, the label calculation is divided into two parts. One is the cumulative function of multiple exposures of the label. The score of each exposure is calculated through the time decay function. The closer the exposure time, the higher the score of the label in the time decay function (i.e., the above-mentioned cumulative exposure value). As time goes by, the score of the label will become lower and lower. The other is the positive and negative feedback factors of user behavior, which calculate the results of these behaviors converted to the label (i.e., the above-mentioned evaluation value). These behavior feedback times are calculated offline and stored in the database during data cleaning. For example, the above-mentioned video label score and segment label score can be calculated by the following formula:
[0087]
[0088] Among them, n represents the number of times a certain tag appears, including tags hit by behaviors such as play, click, like, favorite, search, etc. δ will give different weights according to user behavior, for example, the weights of search behavior and click behavior will be higher than the weights of like and share behavior. α will be adjusted appropriately according to different tag types, τ is the decay rate (usually a positive number, affecting the speed of decay), Δt is the time difference between the exposure time of the tag and the statistical time, ε is a random factor, β is the adjustment weight for user behavior feedback (i.e. corresponding to the first weight value and the second weight value), U ij It refers to the positive and negative behavioral feedback of users on videos or hot spots (i.e. the first feedback factor and the second feedback factor mentioned above), including collection, like, comment, drag progress bar, speed increase, and replay.
[0089] By calculating the scores for long videos and the labels corresponding to the hot spots in long videos respectively, the accuracy of the relevant parameters in the video recommendation decision model is further improved, thereby improving the accuracy of video recommendations and improving the video recommendation effect.
[0090] In an optional implementation, determining a target recommended video matching the target object from a recommendation pool according to the video tag information and the description information includes:
[0091] S1, obtaining a target video tag whose video tag score is greater than a first threshold, a first target segment tag whose segment tag score is greater than a second threshold, and a second target segment tag whose viewing conversion rate is greater than a third threshold;
[0092] S2, determining target recommendation description information according to the product of the target video tag, the first target segment tag, and the second target segment tag and the weight coefficients respectively matched thereto;
[0093] S3, determining a target recommended video that matches the target recommended description information from the recommendation pool.
[0094] The above process S1-S3 is described in a specific implementation manner:
[0095] For example, the user profile of the target object (user A) is a fan of science fiction movies, likes the performance of star B, and often watches and rewatches video clips with tense or high-energy action scenes. In the recommendation pool of user A, there are many videos and clips, including science fiction movies, action movies, and other works of star B.
[0096] First, label score calculation and threshold screening are performed. For example, for video Z (science fiction movie): generalized label (video label): science fiction, movie, director C, high rating; label score: assume that the score of science fiction is 80, the score of movie is 60, the score of director C is 30, and the score of high rating is 75; screening: set the first threshold to 70, and only retain science fiction and high rating as target video labels.
[0097] For a hot spot clip in Z "Star B drives a spaceship through a wormhole": refine the labels (clip labels): Star B, wormhole, tension, high-energy action, high-definition; label scores: Star B's score is 90, tension is 85, high-energy action is 75, and high-definition is 80; screening: set the second threshold to 75, and retain Star B, tension, high-energy action, and high-definition as the first target clip labels; assuming that the viewing conversion rate of this clip is 40%, set the third threshold to 30%, so the conversion rate of this clip meets the conditions and is used as the second target clip label.
[0098] Furthermore, for example, the weight coefficients are set as follows: Science Fiction = 1.0, High Rating = 0.8; Star B = 1.2, Intense = 0.9, High Energy Action = 1.0, High Definition = 0.8. The weighted label scores are calculated as follows: Science Fiction = 80, High Rating = 60; Star B = 108, Intense = 76.5, High Energy Action = 75, High Definition = 64; the above target recommendation description information may be "Highly rated science fiction movie with intense high energy action scenes featuring Star B".
[0099] Finally, for example, the matching videos found in the recommendation pool include A1 (intense and high-energy action scenes with star B, which is in line with the positioning of a highly rated science fiction movie), A2 (also contains intense and high-energy action scenes, but may have a slightly lower score in terms of high ratings), and A3 (a highly rated science fiction movie, but without the participation of star B and no clear intense and high-energy action labels). The video that best matches the target description information is determined based on the weighted score of each video and is recommended to user A, such as A1. The above implementation method improves the accuracy of the recommendation results and the recommendation efficiency.
[0100] In an optional implementation, determining the interaction value of each of the plurality of reference videos according to the behavior feedback results between the target object and the plurality of reference videos further includes:
[0101] S1, obtaining the interaction behavior parameters of the target object in different time periods and the weight values matching them;
[0102] S2, determining the interaction value of the reference video according to the product of the interaction behavior parameters in different time periods and the weight values matching them.
[0103] In the above step S1, the interaction behavior parameters of the target object in different time periods and the corresponding weight values are obtained. This can be understood as obtaining the interaction behavior of the target object A with the reference video in different time periods such as March, the first week of March, and every weekend of March. That is, the target object has different interaction frequencies and degrees of interaction with the reference video in different time periods. Therefore, the user's short-term interests and long-term preferences are comprehensively considered to obtain more accurate analysis results.
[0104] Further in step S2, for example, the formula can be: interaction value = short-term interaction value * short-term weight + medium-term interaction value * medium-term weight + long-term interaction value * long-term weight. Assume that user X liked "Celebrity Y Highlight Moment" 30 times in the past week. Under the influence of the short-term weight (for example, set to 0.5), the interaction value is 15 (30*0.5).
[0105] Specifically, the following formula can be used to calculate the interaction value of a long video, and then filter out the long videos that can hit the hot spots. Because the ultimate goal is to take the hot spots in the long video, the analysis data of the collection is integrated into the long video.
[0106]
[0107] Among them, β and γ are weight parameters for balancing long videos and collections; n is the divided time period, corresponding to short-term, long-term, and ultra-long-term (i.e., considering the analysis of video interaction at different time lengths); m represents multiple index factors; the index factors include user participation index, replay index, click-through rate index, and viewing index; L ij is the different exponential factors under long video; a ij is the weight parameter of different exponential factors under long video; C ij are different exponential factors under the collection; b ij is the weight parameter of different exponential factors under the collection; N is the number of long videos contained in the current collection; ε is a random factor.
[0108] Through the above-mentioned implementation methods recorded in this application, by automatically calculating the interaction value of the video, long videos from which hot spots can be extracted are screened out, so that the recommendation system can more accurately identify the climax parts and other content that attracts users in the video, and improve the efficiency of marking effective hot spots in the video.
[0109] In an optional implementation, after determining the interaction value of the reference video according to the product of the interaction behavior parameter and the weight value matching the interaction behavior parameter in different time periods, the method further includes:
[0110] S1, determining the reference video whose interaction value is greater than the target threshold as the target video;
[0111] S2, determining interaction description information matching the target video, wherein the interaction description information is used to indicate an interaction position of the target object with respect to the target video;
[0112] S3, determining a marked segment matching the target video according to the interaction description information.
[0113] As an optional implementation, the interaction description information may specifically be the distribution of the start point of the replay of the long video, the number of comments, and the number of likes. For example, "User X has high interaction behavior in the segment from 1 hour 23 minutes 45 seconds to 1 hour 26 minutes 30 seconds in StarCraft XX, especially replaying, liking, and commenting. This specific segment can then be used as a marked segment that matches the target video StarCraft XX."
[0114] The following is a complete description of the execution process of video recommendation using a specific implementation method:
[0115] S1, using recommendation data to identify hot spots. First, in order to solve the problem of large video volume and capturing user interest points, relying on the data feedback of the recommendation system, we conduct in-depth statistics, mining and analysis on the viewing data of a large number of long videos and their collections. The core of this process is to identify which collections and long videos are worth promoting, and then hit the hot spots.
[0116] In addition, when conducting these data analyses, the integration of ultra-long-term, long-term, and short-term data was comprehensively considered, because the performance of medium and long-tail content in user behavior data often shows the characteristics of being dispersed over a long time span and relatively sparse in the short term. Data analysis is not limited to coarse-grained intuitive indicators of user shallow interaction such as viewing time and click-through rate, but also comprehensively considers the user's deep-level active interactive behavior: dragging the progress bar, adjusting the playback speed, and reviewing the time period statistics of specific clips in long videos. The adjustment of these actions can more finely reflect the user's real interest and preference for video content, and the number of comments in each time period of long videos can tap the audience's emotional peaks, thereby providing strong support for screening out more potential explosive content in the manual annotation stage.
[0117] Secondly, artificial intelligence algorithms can comprehensively analyze video content and provide strong support for the annotation and recommendation of hot spots. This multi-dimensional analysis method not only improves the intelligence level of the content recommendation system, but also lays a solid foundation for accurate recommendations and improved user experience.
[0118] Specifically, the following formula is used to filter out long videos that can hit the hot spots. Because the hot spots in the long videos are ultimately selected, the analysis data of the collection is integrated into the long videos.
[0119]
[0120] Among them, β and γ are weight parameters for balancing long videos and collections; n is the divided time period, corresponding to short-term, long-term, and ultra-long-term; m represents multiple index factors; the index factors include user participation index, replay index, click-through rate index, and viewing index; L ij is the different exponential factors under long video; a ijis the weight parameter of different exponential factors under long video; C ij are different exponential factors under the collection; b ij is the weight parameter of different exponential factors under the collection; N is the number of long videos contained in the current collection; ε is a random factor.
[0121] It should be noted that the user engagement index is used to measure the degree of user interaction with the video, and the calculation formula can be:
[0122]
[0123] Among them, L is the number of likes, C is the number of comments, and S is the number of shares. is n times the speed adjustment, T is the speed, and U is the number of users. When EI is the user engagement index for a long video, the corresponding data for the number of likes and comments is from the long video; when EI is the user engagement index for a collection, the corresponding data for the number of likes and comments is from the collection.
[0124] The replay index is used to measure the degree of repeated viewing of content by users. The calculation formula can be:
[0125]
[0126] in, It is the cumulative sum of n replay times, and U*D is the product of the number of users and the length of the corresponding long video or collection.
[0127] The click-through rate index is the click-through rate of a long video or collection, and the calculation formula can be:
[0128]
[0129] Among them, C is the number of clicks and S is the number of exposures.
[0130] The viewing index is used to measure the viewing of long videos or collections. The calculation formula can be:
[0131]
[0132] Among them, V is the number of views and U is the number of users.
[0133] In addition to calculating the score of long videos in the above way, we will also count the distribution of the starting points of the long video, the number of comments, and the number of likes. At the same time, we use the existing artificial intelligence algorithm to analyze the audio and video content and identify the emotional climax points that the audience may be interested in. The above data is output for reference for the specific points of the hit points, as shown in the flow chart Figure 4 As shown:
[0134] S402, recommend offline data; S404, calculate long video data;
[0135] S406, manual preliminary screening; combined with S402-1, review the distribution of the starting point, number of comments, and number of likes, and S406-1, long video content understanding; then execute S408, manual analysis and decision-making;
[0136] S410, hit the hot spots in the long video; finally obtain S412, the hot spot database;
[0137] S2, in order to deeply understand the interests and needs of the audience, conducted in-depth user behavior analysis. These data are mainly mined from the user's viewing history, interaction records and search preferences. Different types of labels are obtained for the long video behavior and hot spot behavior watched by users. Hot spot behavior focuses on obtaining detailed labels such as stars, and combined labels of emotional feelings and action behaviors, while long videos focus on obtaining generalized labels such as classification and subject matter.
[0138] The user's interactive records in this application mainly use the positive and negative feedback of collection, like, comment, drag progress bar, speed, and replay to adjust the weight of the corresponding tag, thereby affecting the recall and sorting links. For example, if a user collects a feature film, and the category and subject of the feature film hit the tag in the user's portrait, then this type of tag will be weighted to increase the weight of the tag in the user's tag system, and the content hit by the tag will have a better chance of being launched. Search preferences are mainly weighted for the collection, star, and subject tags of the search.
[0139] The label calculation is divided into two parts. One is the cumulative function of multiple exposures of the label. The score of each exposure is calculated through the time decay function. The closer the exposure time, the higher the score of the label in the time decay function. As time goes by, the score of the label will become lower and lower. The other is the positive and negative feedback factors of the user behavior mentioned above. These behaviors are converted to the label. These behavior feedback times are calculated offline and stored in the database during data cleaning. When calculating user behavior, they are directly pulled from the database. The specific method of label calculation is as follows:
[0140]
[0141] Among them, n represents the number of times a certain tag appears, including tags hit by behaviors such as play, click, like, favorite, search, etc. δ will give different weights according to user behavior, for example, the weights of search behavior and click behavior will be higher than the weights of like and share behavior. α will be appropriately adjusted according to different tag types, τ is the decay rate (usually a positive number that affects the speed of decay), Δt is the time difference between the exposure time of the tag and the statistical time, ε is a random factor, β is the adjustment weight for user behavior feedback, and U ijIt refers to the positive and negative behavioral feedback from users on videos or hot spots, including collection, likes, comments, dragging the progress bar, speeding up, and replaying.
[0142] Specifically, the positive and negative feedback behavior index is used to measure the user's preference for the content, and the calculation formula can be:
[0143]
[0144] Among them, a 1 to a 7 is the weight of each behavior, L is the number of likes, C is the number of comments, H is the number of favorites, S is the number of shares, is n times the speed adjustment, T is the speed, is the number of replays. Is the number of times the progress bar is dragged.
[0145] S3, associate the hot spots with long videos for recommendation.
[0146] Specifically, by offline calculating the click-through rate of the hot spots in the same long video, and the conversion of the duration of each hot spot to the long video, we can know the instantaneous conversion ability of different hot spots and the conversion ability to the duration of the long video. In this way, we can recommend hot spots with higher instantaneous conversion to users, and then recommend hot spots that can increase the duration of the video to users in a combination of multiple recommendation forms on the on-demand page.
[0147] By recommending embedding points, we can know which hotspot generated the long video playback. The conversion duration of the long video is calculated as follows:
[0148]
[0149] in, It is the cumulative sum of n playback durations, and U is the number of users. It should be noted that because the comparison is made between the duration conversions brought about by different burst points in the same long video, there is no need to process the duration of the long video in the denominator.
[0150] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0151] According to another aspect of the embodiment of the present application, a video recommendation device for implementing the above-mentioned video recommendation method is also provided. Figure 5 As shown, the device comprises:
[0152] A first determining unit 502 determines the interaction value of each of the multiple reference videos according to the behavior feedback results between the target object and the multiple reference videos;
[0153] A screening unit 504, determining at least one target video satisfying a screening condition from a plurality of reference videos according to the interaction value;
[0154] The second determining unit 506 determines video tag information that matches the at least one target video, and description information that matches each marked segment in the at least one target video, wherein the marked segment is a segment that generates an interactive behavior with the target object, and the description information includes tag description information and evaluation index information, and the evaluation index information is used to indicate the degree of influence of the marked segment on the video viewing time;
[0155] The recommendation unit 508 determines a target recommended video matching the target object from the recommendation pool according to the video tag information and the description information.
[0156] Optionally, the second determination unit 506 is further used to determine the video tag that matches the target video and the fragment tag that matches the marked fragment; determine the video tag score indicated by the video tag information according to the first tag hit rate, wherein the first tag hit rate is used to indicate the degree to which the target object interacts with the video type corresponding to the video tag; determine the fragment tag score indicated by the description information according to the second tag hit rate, wherein the second tag hit rate is used to indicate the degree to which the target object interacts with the fragment type corresponding to the fragment tag; determine the viewing conversion rate indicated by the evaluation index information according to the ratio of the total playback time matching the marked fragment to the number of viewing users.
[0157] Optionally, the second determination unit 506 includes: a matching module, configured to use a generalized video tag matching the target video as a video tag matching the target video; determine a segment refinement tag matching the marked segment, and use a combination of the segment refinement tags as a segment tag matching the marked segment.
[0158] Optionally, the above-mentioned second determination unit 506 includes: a calculation module, used to determine a first weight value of a first feedback factor matching the target video according to the first label hit rate, and determine a second weight value of a second feedback factor matching the marked segment according to the second label hit rate, wherein the feedback factor is a behavior parameter corresponding to the interactive behavior generated by the target object; determine the video label score according to the exposure cumulative value matching the target video and the first evaluation value, wherein the first evaluation value is the product of the first weight value and the first feedback factor; determine the segment label score according to the exposure cumulative value matching the marked segment and the second evaluation value, wherein the second evaluation value is the product of the second weight value and the second feedback factor.
[0159] Optionally, the recommendation unit 508 is further used to obtain a target video tag whose video tag score is greater than a first threshold, a first target segment tag whose segment tag score is greater than a second threshold, and a second target segment tag whose viewing conversion rate is greater than a third threshold; determine the target recommendation description information according to the product of the target video tag, the first target segment tag and the second target segment tag and the weight coefficients that match them respectively; and determine the target recommended video that matches the target recommendation description information from the recommendation pool.
[0160] Optionally, the first determination unit 502 is further used to obtain the interaction behavior parameters of the target object in different time periods and the weight values matching therewith; and determine the interaction value of the reference video according to the product of the interaction behavior parameters in different time periods and the weight values matching therewith.
[0161] Optionally, the above-mentioned first determination unit 502 also includes: a third determination module, used to determine a reference video whose interaction value is greater than a target threshold as a target video; determine interaction description information that matches the target video, wherein the interaction description information is used to indicate the interaction position of the target object to the target video; and determine a marked segment that matches the target video based on the interaction description information.
[0162] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned video recommendation method when running.
[0163] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned video recommendation method is also provided. The electronic device may be Figure 1 The terminal device or server shown in the figure. This embodiment is described by taking the electronic device as a mobile phone or a computer as an example. Figure 6 As shown, the electronic device includes a memory 602 and a processor 604. The memory 602 stores a computer program, and the processor 604 is configured to execute the steps in any of the above method embodiments through the computer program.
[0164] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0165] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0166] S1, determining the interaction value of each of the multiple reference videos according to the behavior feedback results between the target object and the multiple reference videos;
[0167] S2, determining at least one target video that meets the screening condition from the multiple reference videos according to the interaction value;
[0168] S3, determining video tag information that matches at least one target video, and description information that matches each marked segment in at least one target video, wherein the marked segment is a segment that generates an interactive behavior with the target object, and the description information includes tag description information and evaluation index information, and the evaluation index information is used to indicate the degree of influence of the marked segment on the video viewing time;
[0169] S4, determining a target recommended video matching the target object from the recommendation pool according to the video tag information and the description information.
[0170] Alternatively, a person skilled in the art may understand that: Figure 6 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Figure 6 The structure of the electronic device is not limited. Figure 6 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 6 Different configurations are shown.
[0171] Among them, the memory 602 can be used to store software programs and modules, such as the program instructions / modules corresponding to the video recommendation method and device in the embodiments of the present application. The processor 604 executes various functional applications and data processing by running the software programs and modules stored in the memory 602, that is, to implement the above-mentioned video recommendation method. The memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 602 may further include a memory remotely located relative to the processor 604, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. As an example, if Figure 6 As shown, the memory 602 may include but is not limited to the first determination unit 502, the screening unit 504, the second determination unit 506 and the recommendation unit 508 in the video recommendation device. In addition, it may also include but is not limited to other module units in the video recommendation device, which will not be repeated in this example.
[0172] Optionally, the transmission device 606 is used to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one example, the transmission device 606 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 606 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0173] In addition, the electronic device mentioned above further includes: a display 608 and a connection bus 610, which is used to connect various module components in the electronic device mentioned above.
[0174] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. The nodes may form a point-to-point network, and any form of computing device, such as a server, terminal or other electronic device, may become a node in the blockchain system by joining the point-to-point network.
[0175] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the method provided in the above various optional implementations;
[0176] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0177] S1, determining the interaction value of each of the multiple reference videos according to the behavior feedback results between the target object and the multiple reference videos;
[0178] S2, determining at least one target video that meets the screening condition from the multiple reference videos according to the interaction value;
[0179] S3, determining video tag information that matches at least one target video, and description information that matches each marked segment in at least one target video, wherein the marked segment is a segment that generates an interactive behavior with the target object, and the description information includes tag description information and evaluation index information, and the evaluation index information is used to indicate the degree of influence of the marked segment on the video viewing time;
[0180] S4, determining a target recommended video matching the target object from the recommendation pool according to the video tag information and the description information.
[0181] Optionally, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0182] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0183] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0184] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0185] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0186] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0188] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A video recommendation method, characterized in that: include: Determining the interaction value of each of the plurality of reference videos according to the behavior feedback results between the target object and the plurality of reference videos; Determine at least one target video that meets a screening condition from the plurality of reference videos according to the interaction value; Determine video tag information that matches at least one of the target videos, and description information that matches each of the marked segments in at least one of the target videos, wherein the marked segments are segments that interact with the target object, and the description information includes tag description information and evaluation index information, and the evaluation index information is used to indicate the degree of influence of the marked segments on the video viewing time; A target recommended video matching the target object is determined from a recommendation pool according to the video tag information and the description information.
2. The method according to claim 1, characterized in that The determining of the video tag information respectively matching at least one of the target videos and the description information respectively matching each marked segment in at least one of the target videos includes: Determining a video tag matching the target video and a segment tag matching the marked segment; Determining a video tag score indicated by the video tag information according to a first tag hit rate, wherein the first tag hit rate is used to indicate the degree to which the target object interacts with the video type corresponding to the video tag; Determining the segment tag score indicated by the description information according to the second tag hit rate, wherein the second tag hit rate is used to indicate the degree to which the target object interacts with the segment type corresponding to the segment tag; The viewing conversion rate indicated by the evaluation indicator information is determined according to the ratio of the total playback duration matching the marked segment to the number of viewing users.
3. The method according to claim 2, characterized in that The determining of the video tag matching the target video and the segment tag matching the marked segment includes: Using the video generalization tag matching the target video as the video tag matching the target video; A fragment refinement label matching the marked fragment is determined, and a combination of the fragment refinement labels is used as the fragment label matching the marked fragment.
4. The method according to claim 2, characterized in that: After determining the video tag matching the target video and the segment tag matching the marked segment, the method further comprises: Determining a first weight value of a first feedback factor matching the target video according to the first tag hit rate, and determining a second weight value of a second feedback factor matching the marked segment according to the second tag hit rate, wherein the feedback factor is a behavior parameter corresponding to the interactive behavior generated by the target object; Determine the video tag score according to the exposure accumulation value matched with the target video and a first evaluation value, wherein the first evaluation value is the product of the first weight value and the first feedback factor; The segment label score is determined according to the exposure accumulation value matched with the marked segment and a second evaluation value, wherein the second evaluation value is the product of the second weight value and the second feedback factor.
5. The method according to claim 2, characterized in that: The step of determining a target recommended video matching the target object from a recommendation pool according to the video tag information and the description information includes: Acquire a target video tag whose video tag score is greater than a first threshold, a first target segment tag whose segment tag score is greater than a second threshold, and a second target segment tag whose viewing conversion rate is greater than a third threshold; Determine target recommendation description information according to the product of the target video tag, the first target segment tag, and the second target segment tag and the weight coefficients respectively matched thereto; Determine a target recommended video matching the target recommended description information from the recommendation pool.
6. The method according to claim 1, characterized in that The step of determining the interaction value of each of the plurality of reference videos according to the behavior feedback results between the target object and the plurality of reference videos further includes: Obtaining interaction behavior parameters of the target object in different time periods and weight values matching the interaction behavior parameters; The interaction value of the reference video is determined according to the product of the interaction behavior parameters in different time periods and the weight values matching them.
7. The method according to claim 6, characterized in that After determining the interaction value of the reference video according to the product of the interaction behavior parameters in different time periods and the weight values matched thereto, the method further comprises: Determine the reference video whose interaction value is greater than a target threshold as the target video; Determining interaction description information matching the target video, wherein the interaction description information is used to indicate an interaction position of the target object with respect to the target video; The marked segment matching the target video is determined according to the interaction description information.
8. A video recommendation device, characterized in that: include: A first determining unit, determining the interaction value of each of the plurality of reference videos according to the behavior feedback results between the target object and the plurality of reference videos; A screening unit, for determining at least one target video satisfying a screening condition from the plurality of reference videos according to the interaction value; A second determination unit is configured to determine video tag information that matches at least one of the target videos, and description information that matches each of the marked segments in at least one of the target videos, wherein the marked segments are segments that interact with the target object, and the description information includes tag description information and evaluation index information, wherein the evaluation index information is used to indicate the degree of influence of the marked segments on the video viewing time; The recommendation unit determines a target recommended video matching the target object from a recommendation pool according to the video tag information and the description information.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program is executed by a processor to perform the method described in any one of claims 1 to 7.
10. An electronic device, comprising: A memory and a processor, wherein the memory is used to store a computer program, and wherein the processor is used to execute the steps of the method according to any one of claims 1 to 7 when calling the computer program.