Data-driven streaming web content recommendation method, system, and medium
By combining shot segmentation with saliency and aesthetic visual analysis, the problem of lack of targeting and low matching degree of user characteristics in content recommendation on streaming media platforms has been solved, achieving accurate content recommendation and improving user experience.
Patent Information
- Application Number
- CN202511299354.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-09-12
AI Technical Summary
Existing content recommendation methods on streaming platforms struggle to accurately capture user preferences, lacking features such as shot segmentation, saliency assessment, and aesthetic visual analysis. This results in a low degree of matching between recommended content and user needs, negatively impacting user experience.
The panoramic video is segmented and encoded in blocks using a shot segmentation mechanism. Combined with saliency analysis and aesthetic visual analysis, a video stream saliency descending sequence and a target cover set are generated to form an accurate content recommendation decision and push it to the target user's streaming media network visual interface.
It enables precise content recommendation based on user characteristics, improves the effectiveness of content recommendation and user experience on streaming media networks, and ensures that recommended content meets users' aesthetic preferences and content needs.
Smart Images

Figure CN120812353B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of content recommendation technology, specifically to data-driven streaming media network content recommendation methods, systems, and media. Background Technology
[0002] With the rapid development of streaming media technology, users' demand for personalized content recommendations is growing. Existing streaming media platforms often struggle to accurately capture user preferences, lacking a systematic integration of shot segmentation, saliency assessment, and aesthetic visual analysis in video content processing. This results in a low degree of matching between recommended content and user needs, negatively impacting user experience.
[0003] Existing technologies suffer from a lack of targeted content recommendations and low matching with user characteristics, resulting in poor recommendation performance and a poor user experience. Summary of the Invention
[0004] This application provides a data-driven method, system, and medium for recommending content on streaming media networks, which addresses the technical problems of poor recommendation performance and poor user experience caused by the lack of targeting and low matching degree of user characteristics in existing streaming media network content recommendations.
[0005] In view of the above problems, this application provides a data-driven method, system and medium for recommending content on streaming media networks.
[0006] A first aspect of this application provides a data-driven method for recommending content on a streaming media network, the method comprising:
[0007] The panoramic video is segmented using a shot segmentation mechanism to obtain a first-shot video. This first-shot video is then block-encoded to obtain an encoding result, which includes multiple video streams. A saliency analysis strategy is introduced to evaluate the saliency of these multiple video streams, and a descending saliency sequence is generated based on the evaluation results. An aesthetic visual analysis strategy is introduced to perform aesthetic visual analysis on the target video streams in the descending saliency sequence to form a target cover set. This target cover set is then filtered to determine the target cover corresponding to the target user. A target information content recommendation decision is generated by combining the descending saliency sequence and the target cover. Based on the target information content recommendation decision, the panoramic video is recommended to the target user's streaming media network visual interface.
[0008] A second aspect of this application provides a data-driven streaming media network content recommendation system, the system comprising:
[0009] The encoding result acquisition module is used to segment the panoramic video according to the shot segmentation mechanism to obtain the first shot video, and to encode the first shot video in blocks to obtain the encoding result, wherein the encoding result includes multiple sets of video streams; the saliency evaluation module is used to introduce a saliency analysis strategy to evaluate the saliency of the multiple sets of video streams, and to generate a video stream saliency descending sequence based on the saliency evaluation result; the target cover filtering module is used to introduce an aesthetic visual analysis strategy to perform aesthetic visual analysis on the target video streams in the video stream saliency descending sequence to form a target cover set, and to filter the target cover set to determine the target cover corresponding to the target user; the content recommendation decision generation module is used to combine the video stream saliency descending sequence and the target cover to generate a target information content recommendation decision; the panoramic video recommendation module is used to recommend the panoramic video to the target user's streaming media network visual interface according to the target information content recommendation decision.
[0010] In a third aspect of this application, a computer-readable storage medium is provided storing a computer program for executing the data-driven streaming media network content recommendation method provided in this application.
[0011] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0012] The panoramic video is segmented using a shot segmentation mechanism to obtain the first shot video, which is then block-encoded to obtain the encoding result. A saliency analysis strategy is introduced to evaluate the saliency of the multiple video streams, and a descending saliency sequence of the video streams is generated based on the evaluation results. An aesthetic visual analysis strategy is then introduced to perform aesthetic visual analysis on the target video streams in the descending saliency sequence to form a target cover set, and the target cover set is selected to determine the target cover corresponding to the target user. A target information content recommendation decision is generated. Based on the target information content recommendation decision, the panoramic video is recommended to the target user's streaming media network visual interface. This achieves accurate content recommendation based on user characteristics, improving the effectiveness of streaming media network content recommendation and enhancing the user experience. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart illustrating the data-driven streaming media network content recommendation method provided in an embodiment of this application.
[0015] Figure 2 This is a schematic diagram of the structure of a data-driven streaming media network content recommendation system provided in an embodiment of this application.
[0016] Figure labeling: Encoding result acquisition module 10, saliency evaluation module 20, target cover screening module 30, content recommendation decision generation module 40, panoramic video recommendation module 50. Detailed Implementation
[0017] This application provides a data-driven method, system, and medium for recommending content on streaming media networks, aiming to address the technical problems in existing technologies where content recommendation on streaming media networks lacks specificity and has low matching degree with user characteristics, resulting in poor recommendation performance and poor user experience.
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0019] Example 1, as Figure 1 As shown, this application provides a data-driven method for recommending content on streaming media networks, the method comprising:
[0020] Step S100: The panoramic video is segmented according to the shot segmentation mechanism to obtain the first shot video, and the first shot video is block-encoded to obtain the encoding result, wherein the encoding result includes multiple video streams.
[0021] Specifically, when segmenting panoramic video according to the shot segmentation mechanism, adjacent frame groups (including the first panoramic image and the second panoramic image) are randomly extracted from the panoramic video. The first feature value and the second feature value of the two are obtained in sequence. If the difference between the feature values does not meet the predetermined threshold, the two panoramic images are used as segmentation nodes. Based on the segmentation nodes, the segmentation result of the panoramic video is obtained, and then the first shot video is obtained. Subsequently, the first shot video is block-encoded to obtain the encoding result containing multiple video streams.
[0022] Step S200: Introduce a saliency analysis strategy to evaluate the saliency of the multiple video streams, and generate a video stream saliency descending sequence based on the saliency evaluation results.
[0023] Specifically, when introducing a saliency analysis strategy to evaluate the saliency of multiple video streams, the first video stream in the multiple video streams is first subjected to offset enhancement calibration processing to obtain a first enhanced image. Then, according to the saliency analysis strategy, the first texture coefficient and the first tone coefficient of the first enhanced image are weighted and calculated to obtain a first saliency index. Finally, the first video stream is sorted in descending order based on the first saliency index to generate a video stream saliency descending sequence.
[0024] Step S300: Introduce an aesthetic visual analysis strategy to perform aesthetic visual analysis on the target video stream in the saliency descending sequence of the video stream to form a target cover set, and filter the target cover set to determine the target cover corresponding to the target user.
[0025] Specifically, when introducing an aesthetic visual analysis strategy to perform aesthetic visual analysis on the most significant target video stream in the descending saliency sequence of video streams, the target enhanced image of the target video stream is first acquired, and a target aesthetic triplet containing visual metadata, contextual metadata, and compositional metadata is constructed. The target aesthetic weight is obtained by analyzing these three types of metadata. Then, the consistency between the target aesthetic weight and the target weight in the aesthetic visual analysis strategy, which is based on a user aesthetic database (containing aesthetic judgment data of multiple users on cover image samples with partial visual elements, partial context elements, and partial compositional elements), is calculated. If the consistency reaches a predetermined limit, the target enhanced image is used as the first cover of the first shot video. At the same time, the second cover of the second shot video is acquired, and a target cover set is constructed with the first cover. Then, the target aesthetic preferences of the target users are matched in the user aesthetic database, and this is used as a filtering constraint to filter the target cover set and determine the target cover corresponding to the target user. If the consistency of the target weight does not reach a predetermined limit, the first shot video is marked without a cover.
[0026] Step S400: Combine the video stream saliency descending sequence with the target cover to generate a target information content recommendation decision.
[0027] Specifically, when generating target information content recommendation decisions by combining the descending saliency sequence of the video stream with the target cover, the system follows the matching rules of video stream saliency and clarity (the earlier the video stream appears in the descending sequence and the higher its saliency, the higher the corresponding clarity). This ensures smooth presentation of recommended content while also saving bandwidth. At the same time, the target cover corresponding to the selected target user is used as the entry point for displaying recommended content. By combining the saliency ranking of the video stream with the target cover information, the final target information content recommendation decision is formed to achieve accurate recommendation of panoramic videos.
[0028] Step S500: Based on the target information content recommendation decision, recommend the panoramic video to the target user's streaming media network visual interface.
[0029] Specifically, based on the generated target information content recommendation decision, panoramic videos are pushed to the target user's streaming media network visual interface. This recommendation process is based on the resolution settings matched with the video stream salience descending sequence, ensuring that users can enjoy the high-definition experience of high-salience video streams while achieving a balance between data saving and smooth playback. At the same time, the target cover corresponding to the selected target user is used as the display entry point, so that the panoramic video is presented in the user's streaming media network visual interface in a form that conforms to their aesthetic preferences and content needs, thus completing the final recommendation process.
[0030] In one possible implementation, step S100 further includes:
[0031] Step S110: Randomly extract adjacent frame groups from the panoramic video, wherein the adjacent frame groups include a first panoramic image and a second panoramic image.
[0032] Step S120: Sequentially obtain the first feature value of the first panoramic image and the second feature value of the second panoramic image.
[0033] Step S130: If the difference between the first feature value and the second feature value does not meet a predetermined threshold, then the first panoramic image and the second panoramic image are used as segmentation nodes according to the lens segmentation mechanism.
[0034] Step S140: Obtain the segmentation result of the panoramic video based on the segmentation nodes, and obtain the first shot video based on the segmentation result.
[0035] Specifically, two consecutive frames are randomly selected from the panoramic video as adjacent frame groups. These adjacent frame groups include a first panoramic image and a second panoramic image. Subsequent analysis of these two adjacent frames lays the foundation for determining the segmentation nodes based on the lens segmentation mechanism.
[0036] The first and second panoramic images are processed separately using image feature extraction algorithms: First, the first panoramic image is preprocessed (e.g., noise reduction and grayscale conversion). Then, feature extraction methods such as SIFT (Scale Invariant Feature Transform) or HOG (Histogram of Oriented Gradients) are used to extract key features (e.g., texture, edges, color distribution) that can characterize the image content from the first panoramic image, and these features are quantized into first feature values. Subsequently, features are extracted from the second panoramic image using the same preprocessing and feature extraction methods and quantized to obtain second feature values. In this way, feature values of the two frames are obtained sequentially, providing a data basis for subsequent calculation of feature value differences.
[0037] After obtaining the first feature value of the first panoramic image and the second feature value of the second panoramic image, the feature value difference between the two is calculated. If the feature value difference does not reach the predetermined threshold, it means that there is a significant difference in the content features of the two frames. At this time, according to the lens segmentation mechanism, the first panoramic image and the second panoramic image are set as segmentation nodes, which serve as the boundary markers for panoramic video segmentation.
[0038] Using defined segmentation nodes as boundaries, the panoramic video is segmented to obtain the segmentation result. That is, the panoramic video is divided into multiple consecutive shot segments according to the segmentation nodes. The segment from the beginning of the panoramic video to the first segmentation node, or the consecutive segment that is first among multiple segmentation nodes, is the first shot video obtained based on the segmentation result.
[0039] In one possible implementation, step S200 further includes:
[0040] Step S210: Perform offset enhancement calibration processing on the first video stream in the multiple video streams to obtain the first enhanced image.
[0041] Step S220: According to the saliency analysis strategy, the first texture coefficient and the first tone coefficient of the first enhanced image are weighted to obtain the first saliency index.
[0042] Step S230: Sort the first video stream in descending order based on the first saliency index to obtain the video stream saliency descending sequence.
[0043] Specifically, for the first video stream among multiple video streams, offset enhancement calibration is performed. This corrects the positional offset that may occur during the transmission or processing of the video stream, while enhancing the detail and contrast of the image, making the image content clearer and the features more prominent. Finally, the first enhanced image after calibration and enhancement is obtained, providing a better image data foundation for subsequent saliency assessment.
[0044] First, the weight ratio of the first texture coefficient and the first hue coefficient is determined based on the saliency analysis strategy (e.g., the weight values of texture features and hue features on saliency are preset according to the strategy); then, the first texture coefficient (obtained by calculating feature parameters such as the thickness, direction, and repetition frequency of the texture in the image) and the first hue coefficient (obtained by analyzing parameters such as the proportion, saturation, and brightness of the dominant hue in the image) are extracted from the first enhanced image; finally, the first texture coefficient and the first hue coefficient are weighted and summed according to the determined weight ratio, and the result is the first saliency index, which is used to quantify the saliency of the first enhanced image.
[0045] The first significance index, obtained through weighted calculation, is used as the benchmark for measuring the significance of the first video stream. The first video streams in multiple groups of video streams are sorted in descending order according to the first significance index from high to low. The resulting sequence is the video stream significance descending sequence, which can intuitively reflect the difference in significance level of each video stream.
[0046] In one possible implementation, step S300 further includes:
[0047] Step S310: Obtain the target enhanced image of the target video stream and construct the target aesthetic triplet of the target enhanced image; wherein, the target aesthetic triplet includes visual metadata, contextual metadata and compositional metadata.
[0048] Step S320: Analyze the visual metadata, the contextual metadata, and the composition metadata to obtain the target aesthetic weight.
[0049] Step S330: Calculate the consistency between the target aesthetic weight and the predetermined aesthetic weight in the aesthetic visual analysis strategy.
[0050] Step S340: When the target weight consistency reaches a predetermined limit, the target enhanced image is used as the first cover of the first shot video.
[0051] Step S350: Obtain the second cover of the second shot video and form the target cover set with the first cover; wherein, it further includes: when the target weight consistency does not reach the predetermined limit, the first shot video is marked without a cover.
[0052] Specifically, the target enhanced image is first obtained by processing the target video stream in the saliency descending sequence of the video stream. Then, a target aesthetic triplet is constructed for the target enhanced image. The triplet specifically includes visual metadata (feature data reflecting the visual presentation of the image), contextual metadata (background or situational data related to the image content), and compositional metadata (data related to the compositional structure of the image), providing a basic data framework for subsequent aesthetic visual analysis.
[0053] The Analytic Hierarchy Process (AHP) is used to achieve the following: First, a hierarchical model is constructed with visual metadata, contextual metadata, and compositional metadata as the criteria layers. The relative importance of each metadata is determined by pairwise comparisons to form a judgment matrix. Next, the largest eigenvalue of the judgment matrix and its corresponding eigenvector are calculated. After consistency verification, the eigenvectors are normalized to obtain the weight proportion of each metadata. Finally, these weight proportions are integrated to form the target aesthetic weight that comprehensively reflects the aesthetic features of the target enhanced image.
[0054] First, the target aesthetic weight and the predetermined aesthetic weight are converted into structured weight vectors. Then, they are processed using vector similarity calculation methods (such as the Pearson correlation coefficient method): the two weight vectors are substituted into the correlation coefficient calculation formula, and a coefficient value between -1 and 1 is obtained by calculating the degree of linear correlation between the two. This coefficient value is the consistency of the target weight. The closer the coefficient value is to 1, the higher the consistency between the target aesthetic weight and the predetermined aesthetic weight, and vice versa.
[0055] After the target weight consistency calculation is completed, the consistency result is compared with the predetermined limit. If the target weight consistency reaches the predetermined limit, it means that the aesthetic features of the target enhancement image meet the expected aesthetic standards. At this time, the target enhancement image is determined as the first cover of the first shot video and presented to the user as a representative image of the shot video.
[0056] After determining the first cover image for the first shot video, the second cover image for the second shot video is obtained using the same aesthetic and visual analysis process. The first and second covers are then integrated to form the target cover set. It is important to note that the number of covers in the target cover set will not exceed the number of video streams in the multiple video streams, nor will it exceed the number of shots in the panoramic video, thus ensuring a reasonable correspondence between the cover set and the video content in terms of quantity.
[0057] After analyzing the target aesthetic triplet of the target enhancement image and calculating the target weight consistency, if the consistency result does not reach the predetermined limit set in the aesthetic visual analysis strategy, it indicates that the target enhancement image does not meet the aesthetic standards for a cover. At this time, no cover is set for the first shot video, and this state is marked to distinguish it from the shot video with a cover.
[0058] In one possible implementation, step S330 further includes:
[0059] Step S331: Establish a user aesthetic database for the streaming media network, wherein the user aesthetic database includes aesthetic judgment data of multiple users on cover image samples.
[0060] Step S332: Obtain the first aesthetic data of the first user on the cover image sample based on the aesthetic judgment data of multiple users on the cover image sample, and obtain the first aesthetic preference.
[0061] Step S333: Analyze the first aesthetic preference to determine the predetermined aesthetic weight; wherein, the cover image sample includes visual meta-samples, contextual meta-samples and compositional meta-samples.
[0062] Specifically, a user aesthetic database for the streaming media network will be established. The core of this database is to collect and store aesthetic judgments made by multiple users on various cover image samples. This data covers users' preferences for different cover images in terms of visual appeal, context, composition, etc., providing basic data support for subsequent aesthetic analysis.
[0063] In the user aesthetic database of the streaming media network, the relevant data of the first user was extracted from the aesthetic judgment data of multiple users on cover image samples (including visual meta-samples, context meta-samples, and composition meta-samples). The total number of visual meta-samples, context meta-samples, and composition meta-samples liked by the user was counted. By comparing the total number of these three types of samples, the sample type with the largest total number was determined as the first user's first aesthetic preference.
[0064] When analyzing the first aesthetic preference to determine the predetermined aesthetic weight, the visual meta-samples, contextual meta-samples, and compositional meta-samples contained in the cover image sample are combined. The total number of users who obtained the preference for each type of sample is counted. Then, the weight of the corresponding sample type in the aesthetic evaluation is determined according to the proportion of the total number of users who obtained the preference for different types of samples, thereby forming the predetermined aesthetic weight, so as to quantify the degree of influence of different metadata types in aesthetic judgment.
[0065] In one possible implementation, step S300 further includes:
[0066] Step S360: Match the target user's target aesthetic preferences in the user aesthetic database.
[0067] Step S370: Filter the target cover set using the target aesthetic preference as a filtering constraint to determine the target cover.
[0068] Specifically, the system uses the unique identifier of the target user (such as user account, ID, etc.) to perform a precise search in the user's aesthetic database, retrieves the user's past aesthetic judgment data on visual meta-samples, contextual meta-samples, and compositional meta-samples, including the types of samples they like, their ratings or selection records for various types of samples, etc., and then extracts and summarizes the target user's tendency in cover aesthetics from this data, thereby matching the target user's target aesthetic preferences.
[0069] The K-nearest neighbor algorithm based on feature vector matching is adopted. First, the target aesthetic preference is transformed into a feature vector containing the weight ratio of visual elements, contextual elements, and compositional elements. At the same time, visual metadata, contextual metadata, and compositional metadata of each cover in the target cover set are extracted and corresponding feature vectors are constructed. Then, the cosine similarity between the target aesthetic preference feature vector and the feature vector of each cover is calculated to measure the degree of matching between the two. Finally, the cover with the highest similarity is selected as the target cover, thereby realizing the screening process constrained by the target aesthetic preference.
[0070] In one possible implementation, step S100 further includes:
[0071] Step S150: Construct a target profile based on the target user's multidimensional target information.
[0072] Step S160: Using the target profile as a matching constraint, match and obtain the historical profile of the most similar historical user.
[0073] Step S170: Analyze the information content form of the most similar historical user to obtain the panoramic video.
[0074] Specifically, the process involves collecting multidimensional information about target users, including their basic attributes (such as age and gender), streaming media usage behavior data (such as viewing time, content type, and interaction history), and possible preference settings. After cleaning the data to remove invalid information, feature extraction techniques are used to extract key features (such as content preference tags and viewing habits) from this multidimensional information. These key features are then integrated into structured data to construct a target profile that comprehensively reflects the characteristics of the target users.
[0075] The target profile is transformed into a feature vector containing multi-dimensional user features (such as content preferences, viewing habits, etc.), and the historical profiles of historical users are also transformed into corresponding feature vectors. Then, based on the feature vector of the target profile, the similarity between it and the feature vectors of each historical user is calculated (using the cosine similarity algorithm or Euclidean distance formula). The historical profile corresponding to the historical user feature vector with the highest similarity is selected, and the historical profile of the most similar historical user is obtained by matching in this way, ensuring that the matching result is constrained by the target profile.
[0076] The system retrieves a form containing information about the most similar historical users, including records of panoramic videos they have watched, saved, liked, or tagged. The content in the form is then categorized and statistically analyzed to extract the types, themes, and related features of panoramic videos that the historical user frequently interacted with. Finally, by combining the similarity between the target user's profile and that of the historical user, panoramic videos that meet the potential needs of the target user are selected from the statistical results, thus determining the final panoramic videos to be recommended.
[0077] In one possible implementation, step S500 further includes:
[0078] Step S510: Obtain the target user's target subscription information and analyze it to obtain the target subscription index.
[0079] Step S520: Match the target advertising ratio corresponding to the target subscription index according to the collaborative revenue mechanism.
[0080] Step S530: The panoramic video is traversed and filtered in the advertising database to obtain the target advertising content.
[0081] Step S540: Combine the target advertising ratio and the target advertising content to generate a target advertising content recommendation decision for the target user.
[0082] Step S550: Insert advertisements during the process of the target user watching the panoramic video based on the target advertisement content recommendation decision.
[0083] Specifically, the user information management module of the streaming media network retrieves subscription-related data of target users, including the type of content service they subscribe to, subscription duration, subscription level, and whether they renew continuously. After structuring this information, a quantitative scoring method is used to convert subscription information of different dimensions into calculable values (e.g., subscription levels are divided into basic, advanced, and premium, corresponding to 1 to 3 points respectively, and 1 point is awarded for each full month of subscription). These quantitative values are then weighted and summarized to obtain a target subscription index that can comprehensively reflect the subscription status of target users.
[0084] The interval mapping algorithm first sets multiple interval ranges of the target subscription index and corresponding advertising ratio parameters based on the collaborative revenue mechanism (for example, an index of 0-30 corresponds to a 20% advertising ratio, 31-60 corresponds to a 12% advertising ratio, and 61-100 corresponds to a 5% advertising ratio), and constructs an index-ratio mapping table. By substituting the calculated target subscription index into this mapping table, its interval is located, and then the advertising ratio parameter corresponding to that interval is called as the target advertising ratio, thereby achieving the matching of the target subscription index and the target advertising ratio.
[0085] The algorithm uses content similarity matching to first extract key features from the panoramic video, such as topic tags, scene categories, and keywords, and convert them into feature vectors. Then, it traverses all advertising content in the advertising database, extracts key features from each advertisement, and converts them into feature vectors. By calculating the cosine similarity between the panoramic video feature vector and the feature vectors of each advertisement, advertising content with a similarity higher than a preset threshold is selected and used as target advertising content.
[0086] First, clarify the proportion of total ad time or the limit on the number of ads specified by the target ad ratio. Then, select ads from the target ad content that match the panoramic video theme and are relevant to the target user's preferences, and sort them according to indicators such as the relevance and attractiveness of the ad content. Subsequently, allocate the playback time or frequency of each ad according to the target ad ratio, determine the time node for ad insertion (such as the transition between video chapters), and form a structured scheme that includes a list of ad content, playback order, duration allocation, and insertion position, thereby generating the target ad content recommendation decision for the target user.
[0087] Based on the generated target ad content recommendation decisions, during the target user's viewing of the panoramic video, the selected target ad content is precisely inserted into the video stream according to the ad insertion time nodes specified in the decisions (such as at natural video segmentation points, chapter transitions, etc.), ad playback order, and duration. Simultaneously, it ensures that the total duration and frequency of ad insertion match the target ad ratio, achieving effective ad content presentation and completing the ad delivery targeted to the specific target user.
[0088] Example 2, based on the same inventive concept as the data-driven streaming media network content recommendation method in the foregoing examples, such as... Figure 2 As shown, this application provides a data-driven streaming media content recommendation system. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0089] The encoding result acquisition module 10 is used to segment the panoramic video according to the shot segmentation mechanism to obtain the first shot video, and to encode the first shot video in blocks to obtain the encoding result, wherein the encoding result includes multiple video streams.
[0090] The saliency assessment module 20 is used to introduce a saliency analysis strategy to assess the saliency of the multiple video streams and generate a video stream saliency descending sequence based on the saliency assessment results.
[0091] The target cover filtering module 30 is used to introduce an aesthetic visual analysis strategy to perform aesthetic visual analysis on the target video streams in the saliency descending sequence of the video streams to form a target cover set, and to filter the target cover set to determine the target cover corresponding to the target user.
[0092] The content recommendation decision generation module 40 is used to combine the video stream saliency descending sequence with the target cover to generate a target information content recommendation decision.
[0093] The panoramic video recommendation module 50 is used to recommend the panoramic video to the target user's streaming media network visual interface based on the target information content recommendation decision.
[0094] Furthermore, the system is also used to implement the following functions:
[0095] Randomly extract adjacent frame groups from the panoramic video, wherein the adjacent frame groups include a first panoramic image and a second panoramic image; sequentially obtain a first feature value of the first panoramic image and a second feature value of the second panoramic image; if the feature value difference between the first feature value and the second feature value does not meet a predetermined threshold, then according to the shot segmentation mechanism, the first panoramic image and the second panoramic image are used as segmentation nodes; obtain the segmentation result of the panoramic video according to the segmentation nodes, and obtain the first shot video based on the segmentation result.
[0096] Furthermore, the system is also used to implement the following functions:
[0097] The first video stream in the plurality of video streams is subjected to offset enhancement calibration processing to obtain a first enhanced image; according to the saliency analysis strategy, the first texture coefficient and the first tone coefficient of the first enhanced image are weighted to obtain a first saliency index; the first video stream is sorted in descending order based on the first saliency index to obtain the video stream saliency descending sequence.
[0098] Furthermore, the system is also used to implement the following functions:
[0099] The process involves: acquiring a target enhanced image from the target video stream and constructing a target aesthetic triplet for the target enhanced image; wherein the target aesthetic triplet includes visual metadata, contextual metadata, and compositional metadata; analyzing the visual metadata, contextual metadata, and compositional metadata to obtain a target aesthetic weight; calculating the consistency between the target aesthetic weight and the target weight of the predetermined aesthetic weight in the aesthetic visual analysis strategy; when the target weight consistency reaches a predetermined limit, using the target enhanced image as the first cover of the first shot video; acquiring a second cover of the second shot video and constructing the target cover set with the first cover; further comprising: when the target weight consistency does not reach the predetermined limit, marking the first shot video as having no cover.
[0100] Furthermore, the system is also used to implement the following functions:
[0101] A user aesthetic database for a streaming media network is constructed, wherein the user aesthetic database includes aesthetic judgment data of multiple users on cover image samples; based on the aesthetic judgment data of multiple users on cover image samples, a first user's first aesthetic data on the cover image sample is obtained to obtain a first aesthetic preference; the first aesthetic preference is analyzed to determine the predetermined aesthetic weight; wherein the cover image sample includes visual meta-samples, context meta-samples, and composition meta-samples.
[0102] Furthermore, the system is also used to implement the following functions:
[0103] The target aesthetic preferences of the target user are matched against the user aesthetic database; the target cover set is filtered using the target aesthetic preferences as a filtering constraint to determine the target cover.
[0104] Furthermore, the system is also used to implement the following functions:
[0105] A target profile is constructed based on the target user's multidimensional target information; using the target profile as a matching constraint, the historical profile of the most similar historical user is obtained; the information content form of the most similar historical user is analyzed to obtain the panoramic video.
[0106] Furthermore, the system is also used to implement the following functions:
[0107] Obtain the target user's target subscription information and analyze it to obtain the target subscription index; match the target advertising ratio corresponding to the target subscription index according to the collaborative revenue mechanism; traverse and filter the panoramic video in the advertising database to obtain target advertising content; combine the target advertising ratio and the target advertising content to generate a target advertising content recommendation decision for the target user; insert advertisements during the target user's viewing of the panoramic video according to the target advertising content recommendation decision.
[0108] Example 3: Based on the same inventive concept as the data-driven streaming media network content recommendation method in the preceding examples, this example provides a computer-readable storage medium for storing software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data-driven streaming media network content recommendation method in this application. The processor executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory, thereby implementing the aforementioned data-driven streaming media network content recommendation method.
[0109] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0110] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0111] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A data-driven method for recommending content of a streaming web, characterized in that, The method comprises the following steps: segmenting the panoramic video according to a shot segmentation mechanism to obtain a first shot video, and block encoding the first shot video to obtain an encoding result, wherein the encoding result comprises a plurality of video streams; introducing a saliency analysis strategy to evaluate the saliency of the plurality of video streams, and generating a saliency descending sequence of video streams according to the saliency evaluation result; introducing an aesthetic visual analysis strategy to analyze the target video stream in the saliency descending sequence of video streams to form a target cover set, and screening the target cover set to determine the target cover corresponding to the target user; combining the saliency descending sequence of video streams and the target cover to generate a target information content recommendation decision; recommending the panoramic video to the streaming media network visual interface of the target user according to the target information content recommendation decision; wherein introducing a saliency analysis strategy to evaluate the saliency of the plurality of video streams, and generating a saliency descending sequence of video streams according to the saliency evaluation result comprises: performing offset enhancement calibration processing on a first video stream in the plurality of video streams to obtain a first enhanced image; according to the saliency analysis strategy, weighting the first texture coefficient and the first tone coefficient of the first enhanced image to obtain a first saliency index; arranging the first video stream in descending order based on the first saliency index to obtain the saliency descending sequence of video streams; introducing an aesthetic visual analysis strategy to analyze the target video stream in the saliency descending sequence of video streams to form a target cover set, comprising: obtaining a target enhanced image of the target video stream and assembling a target aesthetic triple of the target enhanced image; wherein the target aesthetic triple comprises visual metadata, context metadata and composition metadata; analyzing the visual metadata, the context metadata and the composition metadata to obtain a target aesthetic weight; calculating the target weight consistency of the target aesthetic weight and the predetermined aesthetic weight in the aesthetic visual analysis strategy; when the target weight consistency reaches a predetermined limit, the target enhanced image is taken as the first cover of the first shot video; obtaining a second cover of a second shot video and assembling the target cover set with the first cover; wherein it further comprises: when the target weight consistency does not reach the predetermined limit, marking the first shot video without cover.
2. The data-driven streaming web content recommendation method of claim 1, wherein, segmenting the panoramic video according to a shot segmentation mechanism to obtain a first shot video, comprising: randomly extracting a group of adjacent frames in the panoramic video, wherein the group of adjacent frames comprises a first panoramic image and a second panoramic image; obtaining a first feature value of the first panoramic image and a second feature value of the second panoramic image in turn; if the feature value difference between the first feature value and the second feature value does not satisfy a predetermined threshold, then according to the shot segmentation mechanism, taking the first panoramic image and the second panoramic image as a segmentation node; obtaining a segmentation result of the panoramic video according to the segmentation node, and obtaining the first shot video based on the segmentation result.
3. The data-driven streaming web content recommendation method of claim 1, wherein, Before the consistency of the target aesthetic weight and the target weight of the predetermined aesthetic weight in the aesthetic visual analysis strategy is calculated, comprising: Assembling a user aesthetic database of the streaming media network, wherein the user aesthetic database comprises aesthetic judgment data of multiple users on cover image samples; Obtaining first aesthetic data of a first user on the cover image samples according to the aesthetic judgment data of multiple users on the cover image samples, to obtain first aesthetic preferences; Analyzing the first aesthetic preferences to determine the predetermined aesthetic weight; Wherein the cover image samples include bias visual meta-sample, bias context meta-sample and bias construction meta-sample.
4. The data-driven streaming web content recommendation method of claim 3, wherein, Screening the target cover set to determine a target cover corresponding to the target user, comprising: Matching the target aesthetic preferences of the target user in the user aesthetic database; Screening the target cover set with the target aesthetic preferences as a screening constraint to determine the target cover.
5. The data-driven streaming web content recommendation method of claim 1, wherein, Further comprising: Constructing a target portrait based on target multi-dimensional information of the target user; Matching a historical portrait of the most similar historical user with the target portrait as a matching constraint; Analyzing information content forms of the most similar historical user to obtain the panoramic video.
6. The data-driven streaming web content recommendation method of claim 1, wherein, Further comprising: Obtaining target subscription information of the target user and analyzing to obtain a target subscription index; Matching a target advertisement proportion corresponding to the target subscription index according to a cooperative benefit mechanism; Traversing and screening the panoramic video in an advertisement database to obtain target advertisement content; Combining the target advertisement proportion and the target advertisement content to generate a target advertisement content recommendation decision for the target user; According to the target advertisement content recommendation decision, inserting advertisements into the process of the target user watching the panoramic video.
7. A data-driven streaming web content recommendation system, characterized by, The system is used to implement the data-driven streaming media network content recommendation method of any one of claims 1-6, and the system comprises: An encoding result acquisition module is configured to segment a panoramic video into first shot videos according to a shot segmentation mechanism, and block-encode the first shot videos to obtain an encoding result, wherein the encoding result comprises multiple groups of video streams; A saliency evaluation module is configured to introduce a saliency analysis strategy to evaluate the saliency of the multiple groups of video streams, and generate a video stream saliency descending sequence according to the saliency evaluation result; A target cover screening module is configured to introduce an aesthetic visual analysis strategy to analyze the aesthetic visual of a target video stream in the video stream saliency descending sequence to form a target cover set, and screen the target cover set to determine a target cover corresponding to a target user; A content recommendation decision generation module is configured to combine the video stream saliency descending sequence and the target cover to generate a target information content recommendation decision; A panoramic video recommendation module is configured to recommend the panoramic video to a streaming media network visual interface of the target user according to the target information content recommendation decision.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by a processor to implement the data-driven streaming media network content recommendation method of any one of claims 1-6.
Citation Information
Patent Citations
Audio and video SDK (Software Development Kit) interface for swan-gap system
CN118803348A
Short video recommendation method and system based on user data
CN120238678A