Big data-based short play recommendation method and system

By constructing a user interest distribution matrix and a content semantic feature matrix, combining multi-dimensional data analysis and historical interaction weights, the ranking of short drama recommendations is optimized, which solves the problem of inaccurate recommendations in existing technologies and achieves a higher personalized recommendation effect.

CN120602697AInactive Publication Date: 2025-09-05YISHANG (SHENZHEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510745061.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing short drama recommendation methods rely too much on a single data dimension, resulting in inaccurate recommendation results, poor user experience, and ignoring the diversity of user behavior and the complexity of short drama content.

Method used

By collecting multi-dimensional user behavior data, constructing a user interest distribution matrix and a content semantic feature matrix, and combining historical interaction weights, we can judge the degree of fit between the short drama content and user preferences, and use historical distribution data to optimize the sorting and form a recommendation closed loop.

Benefits of technology

It improves the accuracy and personalization of short drama recommendations, and enhances user experience and platform operation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602697A_ABST
    Figure CN120602697A_ABST
Patent Text Reader

Abstract

The invention relates to the field of movies, and discloses a movie recommendation method and system based on big data, and the method comprises the steps: collecting the multi-dimensional behavior data of a user to form an initial behavior data set; based on the initial behavior data set, calculating preference scores of the user for different categories of short play contents to obtain a user interest distribution matrix; dividing user groups and labeling user category labels; forming an initial short episode content feature set; constructing a content semantic feature matrix; if the content semantic feature matrix and the user category label have the matching requirement, judging the integrating degree of the short episode content and the user preference, and outputting a preliminarily recommended short episode content list; and on the basis of the preliminarily recommended episode content list, sorting and optimizing the recommended episode content by using click rate and complete playing rate indexes in the historical distribution data to obtain a final recommended episode content list. The method has the following effect that the accuracy and the individuation degree of short play recommendation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of short plays, and in particular to a short play recommendation method and system based on big data. Background Art

[0002] In today's digital entertainment era, short plays, as an emerging form of short drama content, are widely popular because of their brevity and tight rhythm.

[0003] However, many current recommendation methods still have significant shortcomings. Many systems overly rely on a single data dimension, such as making recommendations based solely on a user's click or viewing history, ignoring the diversity of user behavior and the complexity of the skit content itself. This simplistic approach often leads to inaccurate recommendations, poor user experience, and even a mismatch between the skit content and user interests.

[0004] How to solve the above technical problems is a technical difficulty that needs to be overcome by those skilled in the art.

[0005] Inventing skit content The present invention provides a short drama recommendation method based on big data to at least partially solve the above technical problems.

[0006] In order to solve the above technical problems, the present invention provides a short drama recommendation method based on big data, comprising: Collect multi-dimensional behavioral data of users to form an initial behavioral data set; Calculating the user's preference scores for different categories of skit content based on the initial behavior dataset to obtain a user interest distribution matrix; Analyzing the content category distribution and viewing habits of short dramas based on the user interest distribution matrix to divide user groups and label user categories; Obtaining video metadata from a skit content database to form an initial skit content feature set; Extracting the skit content semantic vector based on the initial skit content feature set and building a content semantic feature matrix in combination with the emotional tone matching requirements; If there is a matching requirement between the content semantic feature matrix and the user category label, the degree of compatibility between the short play content and the user preference is determined and a preliminary list of recommended short play content is output; Based on the preliminary recommended short play content list, the recommended short play content is sorted and optimized using the click rate and completion rate indicators in the historical distribution data to obtain the final recommended short play content list.

[0007] In an optional embodiment, collecting multi-dimensional behavior data of users to form an initial behavior data set includes: Obtain multi-dimensional behavioral data from the interaction log data, including viewing time percentage, like frequency, comment interaction depth, viewing habits during different time periods, and repeat viewing times, and perform preliminary cleaning and formatting to obtain a behavioral data set; Based on the behavioral data set, normalizing the viewing time ratio and the like frequency to obtain a standardized behavioral feature set; Determining whether the viewing habit data for a time period is within a first preset threshold range based on the behavioral feature set; if so, correlating and matching the viewing habit data for the time period with the repeated viewing count data to determine the intensity of the user's viewing preference for the corresponding time period; Extract comment interaction depth data based on viewing preference strength judgment results, and use cluster analysis method to group users' interactive behaviors to obtain user interaction pattern classification; Based on the results of the interaction pattern classification, the correlation data of the like frequency and the number of repeated viewings is obtained. If the like frequency is higher than a second preset threshold, the like frequency and the number of repeated viewings are weighted to determine the user's preference weight for the skit content; Based on the calculation results of the preference weights, all dimensions of the multi-dimensional behavior data are integrated and a decision tree model is used to conduct a comprehensive analysis of user behavior characteristics to obtain an initial behavior data set.

[0008] In an optional embodiment, calculating the user's preference scores for different categories of skit content based on the initial behavior dataset to obtain a user interest distribution matrix includes: Obtain the initial user behavior dataset; Performing weighted aggregation calculation on the multi-dimensional data in the initial user behavior dataset using a preset weight distribution algorithm to obtain a preliminary preference score of the user for the content of the skit; Extracting data related to user behavior and skit content based on the preliminary preference scores; if a user's preference score exceeds a preset threshold, marking the user behavior data and determining it as a high-preference user group; Obtaining viewing time and viewing completion rate data for the high-preference user group, and determining the user's preference for a specific short drama content category by comparing and analyzing the distribution of their behavior on different short drama content; Based on the tendency judgment results, the like frequency and comment depth data are integrated and the user behaviors are grouped using a cluster analysis method to obtain a behavioral pattern classification of different user groups; Extracting relevant data of the user interest distribution matrix through the behavioral pattern classification; wherein, if the user interest distribution matrix data of a certain classification group meets the preset conditions, the preference score is adjusted twice to determine the final user interest distribution; According to the final user interest distribution and the characteristic data of the short play content, a decision tree model is used to associate the user behavior with the short play content characteristics to obtain the user's preference for the short play content category and obtain the user interest distribution matrix.

[0009] In an optional embodiment, analyzing the distribution of short drama content categories and viewing habits by time period based on the user interest distribution matrix to divide user groups and label user categories includes: Obtaining the user interest distribution matrix and using a clustering method to stratify the user groups to obtain preliminary user category division results; Based on the preliminary user category classification results, obtain the short drama content category data and distribution feature data in the user interest distribution matrix; if the distribution feature of a certain user category meets the preset threshold, normalize the user preference scores within the category to determine an adjusted category feature set; Based on the adjusted category feature set, analyzing the correlation data between viewing habits and viewing preferences in different time periods, using statistical tools to calculate the user's activity in different time periods in segments, and obtaining the time period preference distribution of the user group; Based on the time period preference distribution, the category preference weights related to the time period are determined in combination with the skit content categories and the preliminarily labeled user category tags; if the activity level of a certain time period is higher than a preset threshold, the skit content category preferences of the user group in that time period are weighted; Compare the time period-related category preference weight distributions under different user category labels to determine the user group's preference characteristics for the short drama content category in the corresponding time period; The tendency features and user interest data are integrated, and a clustering method is used to adjust the user category labels of the user group to obtain the final user category labels.

[0010] In an optional embodiment, extracting the skit content semantic vector based on the initial skit content feature set and constructing a content semantic feature matrix in combination with the emotional tone matching requirement includes: Obtain the plot type and emotional tone description information from the initial skit content feature set, use semantic analysis technology to perform word segmentation and text normalization on the plot type and emotional tone description information, and obtain standardized skit content semantic data; Performing vectorization based on the standardized skit content semantic data using a pre-established word embedding model to determine a semantic vector for the skit content semantic data; Based on the semantic vector, determining the degree of correlation between the semantic vector and the emotional tone by calculating the vector space distance; if the correlation degree is lower than a preset threshold, returning to the previous step and re-obtaining the semantic vector; Based on the adjusted semantic vector, a mapping of the correspondence between the content semantics and the emotional tone of the skit is constructed to generate a preliminary content semantic feature matrix; For the preliminary content semantic feature matrix, a clustering tool is used to group the distribution characteristics of plot types and emotional tones to determine the matrix structure of the classified content semantic feature matrix; The integrity of the matrix structure after classification is verified. If the integrity of a certain classification is lower than a preset threshold, the content semantic feature matrix after classification is supplemented to obtain an optimized content semantic feature matrix.

[0011] In an optional embodiment, if there is a matching requirement between the content semantic feature matrix and the user category label, the degree of compatibility between the short play content and the user preference is determined and a preliminary list of recommended short play content is output, including: Vectorize the semantics of the skit content to obtain the feature vector of the skit content; Use the pre-established user tag database to obtain tag matching information related to user categories and determine the scope of the preliminary matching user group; If the user group range of the preliminary match and the historical interaction data have an intersection, then the association strength between the user and the skit content is calculated to obtain an interaction weight value; For the interaction weight value, the similarity score between the short play content and the user preference is calculated by integrating the short play content type and preference information to obtain the degree of fit; If the similarity score is higher than a preset threshold, the corresponding short play content will be included in the preliminary recommended short play content list; The order of recommended short play contents is determined by combining the user category and the tag matching information to obtain a preliminary recommended short play content list.

[0012] In an optional embodiment, based on the preliminary recommended short drama content list, the recommended short drama contents are sorted and optimized using the click rate and completion rate indicators in the historical distribution data to obtain a final recommended short drama content list, including: Obtain click-through rate and completion rate indicators from historical data, analyze the preliminary list of recommended short drama content, and obtain initial evaluation results; Based on the initial evaluation results, the recommended short play contents are screened using a preset dynamic threshold. If the click-through rate and completion rate of a certain short play content are both lower than the dynamic threshold, it is removed from the preliminary list to determine a screened set of short play contents; Analyzing the distribution ratio of the skit content categories for the screened skit content set, and if the distribution ratio of a certain type of skit content exceeds a preset range, adjusting the skit content to obtain an adjusted skit content set; Based on the click-through rate and completion rate of the adjusted short drama content set, a logistic regression model is used to perform sorting optimization to obtain a preliminary sorting sequence; According to the preliminary sorting sequence, the category distribution characteristics of the short play content after sorting optimization are obtained. If the concentration of a certain category of short play content in the sequence is higher than the preset standard, its position is adjusted twice. After determining that no secondary adjustment is required, the final recommended short play content list is obtained.

[0013] In an optional embodiment, the method further includes: Obtaining real-time user scene data; the real-time user scene data includes geographic location, environmental information, and device status; Construct scene feature vectors based on user real-time scene data; Combine it with the short play content feature vector and user tag matching information for analysis; Based on the differences in user preferences in different scenarios, the calculation weight of the compatibility between the short drama content and user preferences is dynamically adjusted to optimize the preliminary recommended short drama content list.

[0014] In an optional embodiment, a weighted aggregation calculation is performed on the multi-dimensional data in the initial user behavior dataset using a preset weight distribution algorithm to obtain the user's preliminary preference score for the skit content, including: Construct the state space of the reinforcement learning model; combine the normalized features of the user's multi-dimensional behavior data, historical weight distribution results, and current recommendation effect indicators into a state vector; Define the action space as the weight adjustment of each behavior data dimension, and express the weight update strategy through bounded continuous values; The user's actual interaction behavior with the recommended short drama is used as the immediate reward, and the long-term reward is calculated by combining the matching degree between the predicted preference and the actual preference; Use deep reinforcement learning algorithms to train policy networks and optimize weight adjustment strategies through experience replay; Based on the trained policy network, dynamic weights are generated, and the viewing time ratio, like frequency, and comment interaction depth in the initial user behavior dataset are weighted and aggregated to calculate the user's preliminary preference score for the short drama content.

[0015] In an optional embodiment, the method further includes: Extract the average weekly viewing time, monthly interaction times, and comment sentiment intensity data from user historical behavior data and calculate the activity index A; Determine the fusion weight of personalized recommendation and diversity recommendation based on the calculated activity index A; Based on the user interest distribution matrix, a collaborative filtering algorithm is used to calculate the similarity scores between candidate short dramas and user preferences, and the top N short dramas are selected to generate a personalized recommendation list; Construct category coverage evaluation function and novelty evaluation indicators; The personalized recommendation list and the diversity recommendation list are integrated according to the determined integration weight; The generated final recommendation list is checked to ensure that any 10 consecutive short plays in the list contain at least 4 different categories; if this condition is not met, the list is adjusted until it meets the requirements; When the user activity index A is less than the preset value, the proportion of categories that the user has not watched in the final recommendation list is checked; if the proportion is lower than the first proportion, short dramas of the corresponding category are added from the candidate library until it is greater than the first proportion.

[0016] In a second aspect, the present invention provides a short drama recommendation system based on big data, comprising: The first processing module is used to collect multi-dimensional behavior data of users to form an initial behavior data set; A second processing module is configured to calculate the user's preference scores for different categories of skit content based on the initial behavior data set to obtain a user interest distribution matrix; The third processing module is used to: analyze the content category distribution and viewing habits of the short dramas based on the user interest distribution matrix to divide the user groups and mark the user category labels; A fourth processing module is configured to: obtain video metadata from a skit content database to form an initial skit content feature set; A fifth processing module is configured to extract a skit content semantic vector based on the initial skit content feature set and construct a content semantic feature matrix in combination with an emotional tone matching requirement; a sixth processing module, configured to: if there is a matching requirement between the content semantic feature matrix and the user category label, determine the degree of compatibility between the short play content and the user preference and output a preliminary list of recommended short play content; The seventh processing module is used to: sort and optimize the recommended short drama contents based on the preliminary recommended short drama content list using the click rate and completion rate indicators in the historical distribution data to obtain a final recommended short drama content list.

[0017] Compared to existing technologies, this invention has at least the following beneficial effects: It constructs a user interest distribution matrix through multi-dimensional behavioral data collection and weighted aggregation, and uses clustering methods to classify and label users. It also extracts semantic features from the skit content, constructs a content semantic feature matrix, and combines historical interaction weights to determine the fit between the skit content and user preferences, generating a preliminary list of recommended skit content. Finally, it optimizes sorting using historical distribution data and dynamically adjusts the user interest model, forming a closed-loop recommendation loop. This improves the accuracy and personalization of skit recommendations, effectively enhancing the user experience and platform operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of a short drama recommendation method based on big data of the present invention.

[0019] Figure 2 This is a block diagram of a short drama recommendation system based on big data provided by the second embodiment of the present invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] Reference Figure 1 The first embodiment of the present invention provides a short drama recommendation method based on big data, comprising the following steps: S101, collecting multi-dimensional behavior data of users to form an initial behavior data set; S102, calculating the user's preference scores for different categories of skit content based on the initial behavior data set to obtain a user interest distribution matrix; S103, analyzing the distribution of short drama content categories and viewing habits in different time periods based on the user interest distribution matrix to divide user groups and label user categories; S104, acquiring video metadata from a skit content database to form an initial skit content feature set; S105, extracting the skit content semantic vector based on the initial skit content feature set and constructing a content semantic feature matrix in combination with the emotional tone matching requirement; S106: If there is a matching requirement between the content semantic feature matrix and the user category label, then the degree of compatibility between the short play content and the user preference is determined and a preliminary list of recommended short play content is output; S107 , based on the preliminary recommended short play content list, the recommended short play contents are sorted and optimized using the click rate and completion rate indicators in the historical distribution data to obtain a final recommended short play content list.

[0022] In one embodiment, collecting multi-dimensional behavioral data of users to form an initial behavioral data set includes: Obtain multi-dimensional behavioral data from the interaction log data, including viewing time percentage, like frequency, comment interaction depth, viewing habits during different time periods, and repeat viewing times, and perform preliminary cleaning and formatting to obtain a behavioral data set; Based on the behavioral data set, normalizing the viewing time ratio and the like frequency to obtain a standardized behavioral feature set; Determining whether the viewing habit data for a time period is within a first preset threshold range based on the behavioral feature set; if so, correlating and matching the viewing habit data for the time period with the repeated viewing count data to determine the intensity of the user's viewing preference for the corresponding time period; Extract comment interaction depth data based on viewing preference strength judgment results, and use cluster analysis method to group users' interactive behaviors to obtain user interaction pattern classification; Based on the results of the interaction pattern classification, the correlation data of the like frequency and the number of repeated viewings is obtained. If the like frequency is higher than a second preset threshold, the like frequency and the number of repeated viewings are weighted to determine the user's preference weight for the skit content; Based on the calculation results of the preference weights, all dimensions of the multi-dimensional behavior data are integrated and a decision tree model is used to conduct a comprehensive analysis of user behavior characteristics to obtain an initial behavior data set.

[0023] Specifically, through a pre-established user behavior collection framework, multi-dimensional behavioral data including viewing time percentage, like behavior frequency, comment interaction depth, time viewing habits, and repeat viewing times are obtained from the interaction log data, and preliminary cleaning and formatting are performed to obtain a structured behavioral data set. Among them, viewing time percentage is the ratio of the user's actual viewing time to the total length of the short drama, reflecting the user's immersion in the content; like behavior frequency is the number of times a user likes a short drama within a specific time period, reflecting the user's recognition of the content; comment interaction depth is calculated by comprehensively calculating the number of comments and the number of replies, which is used to evaluate the user's enthusiasm for participating in content interaction; time viewing habits refer to the distribution characteristics of the number of times a user views different time periods in a day, which can be used to analyze the user's active time periods; repeat viewing times are the frequency of users viewing the same short drama, which is an important indicator for measuring the intensity of user interest.

[0024] Based on the aforementioned behavioral data set, we normalized the percentage of viewing time and the frequency of likes using a normalization method to establish a unified dimensional range and obtain a standardized behavioral feature set. Normalization eliminates the impact of dimensional differences between indicators on subsequent analysis, making the data comparable.

[0025] Based on the behavioral feature set, the system determines whether viewing habits for a specific time period fall within a preset threshold. If so, the viewing habits data for that time period is correlated with the repeat viewing data. By analyzing the correlation between the number of views and the frequency of repeat viewings within a specific time period, the user's viewing preference for that time period is determined. For example, if a user's viewing frequency during the nighttime hours accounts for a high proportion and they frequently re-watch a certain type of content, it can be determined that they have a strong preference for that type of content during the nighttime hours.

[0026] Based on the results of the viewing preference intensity assessment, comment interaction depth data is extracted, and cluster analysis is used to group user interactions and categorize interaction patterns across different user groups. Cluster analysis can group users with similar interaction behaviors into categories based on features such as the number of comments and replies, for example, high-interaction, medium-interaction, and low-interaction user groups. Based on the interaction pattern classification results, correlation data on like frequency and repeat viewing is obtained. If the like frequency exceeds a preset threshold, a weighted calculation is performed on the like frequency and repeat viewing count to determine the user's preference for the skit content. This weighted calculation assigns weights based on the importance of different indicators, for example, assigning a weight of 0.3 for like frequency and 0.7 for repeat viewing count to highlight a user's continued interest in the content. Finally, based on the calculated preference weights, all dimensions of the multi-dimensional behavioral data (including viewing time percentage, like frequency, comment interaction depth, viewing habits during different time periods, and repeat viewing count) are integrated, and a decision tree model is used to comprehensively analyze the user behavior characteristics to generate an initial behavioral dataset. The decision tree model is a machine learning model that constructs a tree structure to perform classification and regression analysis on data. It can effectively extract the association rules between user behavior characteristics and form preliminary portrait data of user interests.

[0027] In one embodiment, calculating the user's preference scores for different categories of skit content based on the initial behavior dataset to obtain a user interest distribution matrix includes: Obtain the initial user behavior dataset; Performing weighted aggregation calculation on the multi-dimensional data in the initial user behavior dataset using a preset weight distribution algorithm to obtain a preliminary preference score of the user for the content of the skit; Extracting data related to user behavior and skit content based on the preliminary preference scores; if a user's preference score exceeds a preset threshold, marking the user behavior data and determining it as a high-preference user group; Obtaining viewing time and viewing completion rate data for the high-preference user group, and determining the user's preference for a specific short drama content category by comparing and analyzing the distribution of their behavior on different short drama content; Based on the tendency judgment results, the like frequency and comment depth data are integrated and the user behaviors are grouped using a cluster analysis method to obtain a behavioral pattern classification of different user groups; Extracting relevant data of the user interest distribution matrix through the behavioral pattern classification; wherein, if the user interest distribution matrix data of a certain classification group meets the preset conditions, the preference score is adjusted twice to determine the final user interest distribution; According to the final user interest distribution and the characteristic data of the short play content, a decision tree model is used to associate the user behavior with the short play content characteristics to obtain the user's preference for the short play content category and obtain the user interest distribution matrix.

[0028] Specifically, an initial user behavior dataset is obtained, including multi-dimensional data such as viewing time, like frequency, comment depth, and completion rate. Using a pre-set weighting mechanism (e.g., assigning a higher weight to viewing time percentage based on metric importance, in this example, a weight of 0.4), a weighted aggregation calculation is performed on this multi-dimensional data to obtain a preliminary user preference score for the content. This weighted aggregation calculation is implemented using a formula such as: preference score = viewing time percentage × 0.4 + normalized like frequency × 0.2 + comment engagement depth score × 0.2 + completion rate × 0.2, thereby quantifying the user's interest in different content types. Based on the preliminary preference scores, data related to user behavior and the skit content is extracted. If a user's preference score exceeds a pre-set threshold (e.g., the platform's high interest threshold of 40 points), the user's behavior data is labeled as belonging to the high-preference user group. This group represents users with a significant interest in specific content, and their behavioral characteristics require further analysis. Detailed data on viewing time and completion rates for high-preference user groups is obtained. By comparing and analyzing their behavior distribution across different short drama content categories (such as suspense and comedy), users' preferences for specific content categories are determined. For example, if a user's viewing time for suspense short dramas accounts for 70% and their completion rate is 85%, they are determined to have a strong preference for suspense. Based on this preference determination, cluster analysis methods (such as K-means clustering) are integrated with data on like frequency and comment depth to group user behavior and categorize behavioral patterns for different user groups. For example, users with high like frequency and high comment depth can be categorized as "actively interactive," while users with low like frequency and high completion rates can be categorized as "deeply engaged." Based on behavioral pattern classification, relevant data for the user interest distribution matrix is ​​extracted. If the interest matrix data for a categorized group meets preset conditions (such as the mean preference score for a category exceeding a threshold of 50 points), the preference score is adjusted secondary, such as through dynamic weighting or outlier filtering, to determine the final user interest distribution. Finally, based on the end-user interest distribution and combined with the characteristic data of the short drama content (such as plot type and emotional tone), a decision tree model was used to correlate user behavior with content features, resulting in a refined structure of user preferences for content categories and forming a user interest distribution matrix. By constructing a tree-like rule set, the decision tree model can identify association rules such as "watch time > 80% and like frequency > 1.5 times / day → suspense preference." Each element in the matrix represents the user's preference score for a specific category (ranging from 0 to 100). For example, a score of 75 for suspense in the matrix indicates a high user interest in this category.

[0029] For example, on a video platform, the system first automatically collects the user's viewing time percentage through the logging module. For example, if a user watches a 60-minute video but actually watches 45 minutes, the system calculates the viewing time percentage as 75% and stores this data in the behavior database. The system then analyzes the frequency of likes, recording that the user has liked 10 videos 8 times in the past 7 days. This calculates an average daily frequency of 1.14 likes. Compared to the platform average of 0.5 likes, the system preliminarily determines that the user has a high level of engagement with the short dramas. Next, the system evaluates the depth of comment engagement. For example, if a user posts three comments on a video, two of which are replied to by other users, the engagement depth score is (number of comments × 0.5 + number of replies × 1.0) = 3 × 0.5 + 2 × 1.0 = 3.5 points, exceeding the platform average of 2.0, indicating high user engagement. The system then analyzed viewing habits during certain time periods, noting that users' views between 8 PM and 10 PM accounted for 60% of their total views. Using a clustering algorithm, the user was categorized as a "nighttime active user" and associated with advertising strategies to improve exposure. Finally, the system counted repeat views and found that the user had watched a particular video five times, exceeding the platform average of two times. Combined with the 75% viewing time, the system used a weighted algorithm (number of repeat views × 0.3 + viewing percentage × 0.7) to calculate an interest index of 5 × 0.3 + 75 × 0.7 = 54.0, exceeding the threshold of 40.0, identifying the short as high-interest content. This data was ultimately integrated into the initial behavioral dataset, which was then used by machine learning models to further predict user interests.

[0030] In one embodiment, analyzing the distribution of short drama content categories and viewing habits by time period based on the user interest distribution matrix to divide user groups and label user categories includes: Obtaining the user interest distribution matrix and using a clustering method to stratify the user groups to obtain preliminary user category division results; Based on the preliminary user category classification results, obtain the short drama content category data and distribution feature data in the user interest distribution matrix; if the distribution feature of a certain user category meets the preset threshold, normalize the user preference scores within the category to determine an adjusted category feature set; Based on the adjusted category feature set, analyzing the correlation data between viewing habits and viewing preferences in different time periods, using statistical tools to calculate the user's activity in different time periods in segments, and obtaining the time period preference distribution of the user group; Based on the time period preference distribution, the category preference weights related to the time period are determined in combination with the skit content categories and the preliminarily labeled user category tags; if the activity level of a certain time period is higher than a preset threshold, the skit content category preferences of the user group in that time period are weighted; Compare the time period-related category preference weight distributions under different user category labels to determine the user group's preference characteristics for the short drama content category in the corresponding time period; The tendency features and user interest data are integrated, and a clustering method is used to adjust the user category labels of the user group to obtain the final user category labels.

[0031] Specifically, a user interest distribution matrix (with dimensions of number of users × number of content categories, e.g., 1000 × 10, where each element represents a user's preference score for a specific category, ranging from 0 to 1) is obtained. A clustering method (such as the K-means clustering algorithm) is then used to stratify the user population. By calculating the Euclidean distance between user preference vectors, users with similar preferences are grouped together, resulting in a preliminary user category segmentation (e.g., initially into five categories). The K-means clustering algorithm is an unsupervised learning method that iteratively updates cluster centers to minimize preference differences among users within the same category. Based on this preliminary user category segmentation, the skit content category data (e.g., news, entertainment, sports, etc.) and distribution feature data (e.g., the average preference score for a particular category within the group) are extracted from the user interest distribution matrix. If the distribution feature of a user category (e.g., the average preference score for news ≥ 0.7) meets a preset threshold, the user preference scores within that category are normalized (scaled to the range of 0–1) to eliminate dimensionality differences and determine the adjusted category feature set. Based on the adjusted set of category features, the correlation between viewing habits and preferences by time period is analyzed. Statistical tools (such as Excel pivot tables or the Python Pandas library) are used to segment user activity (percentage of viewing time) for different time periods (e.g., hourly on a 24-hour clock) to determine the time period preference distribution for each user group (e.g., 50% of a group's viewing time is between 20:00 and 22:00). Based on this time period preference distribution, combined with the skit content categories and pre-labeled user category labels (e.g., "news enthusiast"), the time period-specific category preference weights are determined. If the activity level for a particular time period exceeds a preset threshold (e.g., average viewing time percentage > 40%), the user group's skit content category preference for that time period is weighted (e.g., assigning a weight of 1.2 to news during the nighttime period to emphasize the influence of time). The time period-specific category preference weight distributions for different user category labels are compared and analyzed to determine the differences in weights. This allows the user group to identify their preferred skit content categories for the corresponding time period (e.g., "news enthusiasts" have a significantly higher preference for news during the nighttime period than during other time periods). By integrating propensity characteristics with user interest data, clustering methods (such as hierarchical clustering) are used to adjust the user category labels for user groups. For example, "news enthusiast" can be refined into "evening news active" to obtain the final user category label. Hierarchical clustering constructs a cluster tree and merges or splits categories layer by layer, achieving refined labeling based on the correlation between time period and content preferences.

[0032] For example, on a short drama platform, the system analyzes user preferences using a weighted aggregation method based on an initial behavioral dataset, constructing a user interest distribution matrix. First, the system extracts data on the percentage of time a user spends watching a short drama. Assuming a short drama is 20 minutes long and a user watches 18 minutes, the calculated percentage is 90%. Because this metric reflects user engagement, the system assigns it a weight of 0.4. Next, the system analyzes the frequency of likes. Counting that a user has liked 15 short dramas 10 times in the past five days, the average daily frequency is 2.0. Considering the platform average of 1.0, the system assigns a weight of 0.2 to this metric and normalizes it to a preference factor of 2.0. Next, the system evaluates the depth of comment engagement. For example, if a user posts five comments on a short drama, three of which receive replies, the system calculates a score of 5 × 0.3 + 3 × 0.7 = 3.6, which is 2.5 points higher than the platform average. The weight is assigned to 0.2, reflecting user engagement. The system also calculates viewing completion, finding that of the 20 short dramas the user watched in the last 30 days, 16 reached 100% completion, representing an 80% completion rate. This percentage is weighted at 0.2, serving as a reference indicator of sustained interest. Finally, the system weights and aggregates these indicators to calculate a preference score of (90 × 0.4 + 2.0 × 0.2 + 3.6 × 0.2 + 80 × 0.2) = 36.0 + 0.4 + 0.72 + 16.0 = 53.12. Combined with the short drama genre labels (e.g., suspense, comedy), the scores are mapped to the user's interest distribution matrix. The system identifies that the user has the highest preference for suspense short dramas, accounting for 45% of the matrix, thus forming a complete interest profile.

[0033] In one embodiment, extracting the skit content semantic vector based on the initial skit content feature set and constructing a content semantic feature matrix in combination with the emotional tone matching requirement includes: Obtain the plot type and emotional tone description information from the initial skit content feature set, use semantic analysis technology to perform word segmentation and text normalization on the plot type and emotional tone description information, and obtain standardized skit content semantic data; Performing vectorization based on the standardized skit content semantic data using a pre-established word embedding model to determine a semantic vector for the skit content semantic data; Based on the semantic vector, determining the degree of correlation between the semantic vector and the emotional tone by calculating the vector space distance; if the correlation degree is lower than a preset threshold, returning to the previous step and re-obtaining the semantic vector; Based on the adjusted semantic vector, a mapping of the correspondence between the content semantics and the emotional tone of the skit is constructed to generate a preliminary content semantic feature matrix; For the preliminary content semantic feature matrix, a clustering tool is used to group the distribution characteristics of plot types and emotional tones to determine the matrix structure of the classified content semantic feature matrix; The integrity of the matrix structure after classification is verified. If the integrity of a certain classification is lower than a preset threshold, the content semantic feature matrix after classification is supplemented to obtain an optimized content semantic feature matrix.

[0034] Specifically, an initial content feature set containing plot types (such as "suspense", "comedy", "family") and emotional tone description information (such as "thrilling", "warm and healing", "ups and downs") is obtained from the short play content database, and semantic analysis technology (such as word segmentation and part-of-speech tagging technology in natural language processing) is used to perform word segmentation processing on the above information (breaking the text into independent word units, such as breaking "romantic love plot" into "romance", "love", and "plot") and text normalization processing (unifying uppercase and lowercase, removing punctuation and meaningless stop words) to obtain standardized short play content semantic data.

[0035] Based on standardized semantic data, pre-trained word embedding models (such as Word2Vec and BERT, which learn distributed vector representations of words from massive amounts of text) are used for vectorization. Each word or phrase is converted into a dense vector in a high-dimensional space (for example, the vector corresponding to "suspense" is [0.65, 0.22, 0.13]). This determines the semantic vectors of the short play's content. Word embedding models capture the semantic connections between words. For example, the vectors for "romantic" and "warm" are close in space, reflecting semantic similarity.

[0036] Based on the semantic vector, the correlation between the semantic vector and the preset emotional tone (such as "warmth," "sadness," or "suspense") is determined using vector space distance calculation methods (such as cosine similarity or Euclidean distance). If the correlation falls below a preset threshold (e.g., cosine similarity < 0.5), the algorithm returns to the previous step, adjusts the word segmentation rules or model parameters, and regenerates the semantic vector. For example, if the cosine similarity between the semantic vector of the "family ethics" plot and the emotional tone of "warmth" is 0.4, which is below the threshold of 0.5, the keywords "family affection" and "care" are re-extracted and a new vector is generated.

[0037] Based on the adjusted semantic vectors, a mapping is constructed between the content semantics of the skits (e.g., plot type keywords) and their emotional tones (e.g., "suspense + thrill → suspense-type emotional tone"), generating a preliminary content semantic feature matrix. The matrix dimensions are the number of skits × the semantic dimension (e.g., 3000 × 5), with each element representing the characteristic value of a skit on the corresponding semantic dimension (e.g., the emotional intensity value of a skit is 0.8). Clustering tools (e.g., K-means clustering or hierarchical clustering) are used to group the distribution characteristics of plot types and emotional tones within the preliminary content semantic feature matrix. For example, "suspense" and "thriller" plots and "thrill" and "depressing" emotional tones are grouped into the "high tension content group," determining the matrix structure after classification. Clustering tools group content based on similarities between data points, improving the semantic cohesion of the matrix.

[0038] The matrix structure after classification is checked for completeness (e.g., checking whether each group contains a sufficient number of shorts and sufficient feature coverage). If the completeness of a particular category falls below a preset threshold (e.g., the number of shorts in a group is less than 100 or key semantic dimensions are missing), the matrix is ​​optimized by supplementing similar content data or adjusting the classification rules to obtain the final content semantic feature matrix. For example, if the "Science Fiction" plot category contains only 50 shorts, which is below the threshold of 100, content with the "Science Fiction" tag is retrieved from the database and supplemented to this category to ensure the completeness of the matrix structure.

[0039] In one embodiment, if there is a matching requirement between the content semantic feature matrix and the user category label, the degree of compatibility between the short play content and the user preference is determined and a preliminary list of recommended short play content is output, including: Vectorize the semantics of the skit content to obtain the feature vector of the skit content; Use the pre-established user tag database to obtain tag matching information related to user categories and determine the scope of the preliminary matching user group; If the user group range of the preliminary match and the historical interaction data have an intersection, then the association strength between the user and the skit content is calculated to obtain an interaction weight value; For the interaction weight value, the similarity score between the short play content and the user preference is calculated by integrating the short play content type and preference information to obtain the degree of fit; If the similarity score is higher than a preset threshold, the corresponding short play content will be included in the preliminary recommended short play content list; The order of recommended short play contents is determined by combining the user category and the tag matching information to obtain a preliminary recommended short play content list.

[0040] Specifically, key semantic information (such as plot type vector and emotional tone vector) is extracted from the content semantic feature matrix, and the content semantics of the short play are vectorized to obtain the short play content feature vector (for example, the feature vector of a suspense short play is [0.7, 0.3, 0.8], which corresponds to the suspense type weight, plot complexity, and emotional intensity, respectively).

[0041] The content feature vector maps the semantic features of a skit into a numerical representation in a multidimensional space, quantifying the degree of match between the content and user preferences. Using a pre-established user tag database (which stores user category tags, such as "nighttime suspense enthusiast" and their feature vectors), we obtain tag matching information related to user categories (e.g., the preference vector corresponding to the tag [0.8, 0.2, 0.7]). Using methods such as cosine similarity, we screen out user groups with high similarity to the content feature vectors and determine the scope of a preliminary matching user group (e.g., matching 1,000 users with the tag "suspense"). If the preliminary matching user group overlaps with historical interaction data (e.g., user click and completion records) (i.e., some users have previously interacted with the content or similar content), we calculate the strength of the association between the user and the skit content (e.g., the average click-through rate and completion rate for suspense-related content) based on this historical interaction data. Using a weighted algorithm (e.g., click-through rate × 0.6 + completion rate × 0.4), we determine the interaction weight (ranging from 0 to 1). The interaction weight value reflects the degree of influence of the user's historical behavior on the current recommendation. The higher the weight, the more stable the user's interest in this type of content.

[0042] Based on the interaction weight, the cosine similarity formula (similarity = vector dot product (|A| × |B|)) is used to calculate the similarity score between the skit content and the user's preferences, combining the skit's content type (e.g., suspense, comedy) and user preference information (e.g., the genre weight in the preference vector). For example, if the user preference vector is [0.8, 0.2] (suspense weight 0.8, comedy weight 0.2) and the skit's content vector is [0.7, 0.3], the similarity score is (0.8×0.7+0.2×0.3) / (√(0.8²+0.2²)×√(0.7²+0.3²))≈0.99, indicating a strong match. Cosine similarity measures similarity by calculating the cosine of the angle between the vectors; values ​​closer to 1 indicate a stronger match. If the similarity score exceeds a preset threshold (e.g., 0.6), the corresponding skit is included in the preliminary recommendation list. For example, if the threshold is set at 0.5, a suspense skit with a similarity score of 0.99 will be selected, while a comedy skit with a score of 0.45 will be excluded. Finally, combining user categories (such as "highly interactive suspense user") and tag matching information (such as a preference for nighttime content), the order of recommended skits is determined in descending order of similarity score, taking into account factors such as content freshness (for example, prioritizing recently released content) and diversity (avoiding excessive concentration of similar content), to generate a preliminary recommendation list. For example, the first three skits on the list are suspense skits with the highest similarity scores, interspersed with a science fiction skit that may be of interest to the user to enhance diversity.

[0043] In one embodiment, based on the preliminary recommended short drama content list, the recommended short drama content is sorted and optimized using the click-through rate and completion rate indicators in the historical distribution data to obtain a final recommended short drama content list, including: Obtain click-through rate and completion rate indicators from historical data, analyze the preliminary list of recommended short drama content, and obtain initial evaluation results; Based on the initial evaluation results, the recommended short play contents are screened using a preset dynamic threshold. If the click-through rate and completion rate of a certain short play content are both lower than the dynamic threshold, it is removed from the preliminary list to determine a screened set of short play contents; Analyzing the distribution ratio of the skit content categories for the screened skit content set, and if the distribution ratio of a certain type of skit content exceeds a preset range, adjusting the skit content to obtain an adjusted skit content set; Based on the click-through rate and completion rate of the adjusted short drama content set, a logistic regression model is used to perform sorting optimization to obtain a preliminary sorting sequence; According to the preliminary sorting sequence, the category distribution characteristics of the short play content after sorting optimization are obtained. If the concentration of a certain category of short play content in the sequence is higher than the preset standard, its position is adjusted twice. After determining that no secondary adjustment is required, the final recommended short play content list is obtained.

[0044] Specifically, the initial list of recommended shorts is analyzed based on historical distribution data, including click-through rates (the ratio of user clicks to total impressions) and completion rates (the ratio of user complete views to total play times). This initial list is then analyzed by comparing it to industry averages or recent platform data to provide an initial assessment of content quality. For example, if a short drama has an 8% click-through rate and a 65% completion rate, which are 10% and 70% lower than the platform average, it is considered low-performing content.

[0045] Based on the initial evaluation results, recommended content is filtered using preset dynamic thresholds (e.g., an average platform click-through rate of 0.10 and a completion rate of 0.60 over the past week). If both the click-through rate and completion rate of a particular short drama fall below the dynamic thresholds, it is removed from the preliminary list to avoid recommending low-engagement content to users, thereby determining the final content set. For example, content with a click-through rate of less than 0.10 and a completion rate of less than 0.60 is removed, while short dramas with at least 60% meeting the thresholds are retained. Within this filtered content set, the distribution of the short dramas by category is analyzed (e.g., 70% entertainment and 30% education). If the distribution of a particular category exceeds the preset range (e.g., >60% for a single category), adjustments are made by adding content from other categories or reducing the number of categories to avoid overly diverse recommendations, resulting in the final adjusted content set. For example, if the proportion of entertainment is too high, two or three science fiction or lifestyle short dramas can be manually inserted to adjust the ratio to 50% entertainment and 50% other categories combined. Based on the click-through rate and completion rate of the adjusted content collection, a logistic regression model (a statistical model used to predict binary classification results, here used to predict user preference for content) was used to optimize the ranking. The model uses click-through rate (weighted 0.4) and completion rate (weighted 0.6) as independent variables and outputs a ranking score (ranging from 0 to 1), with higher scores giving higher priority. This yields a preliminary ranking sequence. For example, a skit with a click-through rate of 0.15 and a completion rate of 0.75 would have a score of 0.15 × 0.4 + 0.75 × 0.6 = 0.51, ranking third out of 10 content. Based on this preliminary ranking sequence, the category distribution characteristics of the optimized skits were determined (e.g., the top five were all entertainment). If the concentration of a particular category in the ranking exceeded a preset threshold (e.g., if three or more consecutive episodes were in the same category), its position was adjusted again by interleaving content from other categories to enhance diversity. For example, two entertainment skits from the top three could be swapped with the fourth and fifth ranked technology skits, creating an alternating ranking of "entertainment-technology-entertainment-lifestyle-education." After confirming that the categories are evenly distributed and not overly concentrated, the final list of recommended short drama content is determined.

[0046] For example, specifically, in the process of optimizing the recommended short drama content, the preliminary recommended short drama content list is first analyzed based on historical distribution data. Assuming that the list contains 10 video short drama contents, historical data shows that their click-through rates are 0.12, 0.08, 0.15, 0.09, 0.11, 0.07, 0.14, 0.10, 0.13, and 0.06, and the completion rates are 0.65, 0.55, 0.70, 0.58, 0.62, 0.50, 0.68, 0.60, 0.66, and 0.52, respectively. Using dynamic threshold screening, we set the click-through rate threshold to the average click-through rate of the past week at 0.10, and the completion rate threshold to 0.60. We screened out videos that met both conditions. As a result, we retained videos with a click-through rate greater than or equal to 0.10 and a completion rate greater than or equal to 0.60. We obtained 6 videos that met the conditions, with click-through rates of 0.12, 0.15, 0.11, 0.14, 0.10, and 0.13, and corresponding completion rates of 0.65, 0.70, 0.62, 0.68, 0.60, and 0.66, respectively. Next, combining recommendation diversity adjustment, we analyzed the category distribution of these six videos and found that four of them belonged to the entertainment category and two to the education category. To avoid over-concentration of categories, we introduced diversity weights, setting the entertainment category weight to 0.8 and the education category weight to 1.2. We also calculated a comprehensive score based on click-through rate and completion rate using the formula: score = click-through rate × 0.4 + completion rate × 0.6 + category weight × 0.1. The calculated results were 0.541, 0.582, 0.534, 0.564, 0.520, and 0.549, respectively. Sorting by score from high to low yields the following sequence: Video 2 (0.582), Video 4 (0.564), Video 6 (0.549), Video 1 (0.541), Video 3 (0.534), and Video 5 (0.520). Finally, to further optimize the user experience, combined with business relevance, the thematic relevance of adjacent videos after sorting is checked. If the relevance is lower than 0.3 (calculated through semantic analysis), the order is adjusted to ensure the logic of the recommendation sequence. For example, Video 2 and Video 4 can be swapped to improve thematic coherence, ultimately forming a list of recommended short drama content.

[0047] In one embodiment, a real-time feedback mechanism is used for the final recommended short drama content list to obtain new data from user interaction behaviors, including the number of repeated viewings and interaction behavior combinations, to update the user interest distribution matrix, and form a dynamically adjusted data closed loop.

[0048] Specifically, real-time feedback data is obtained from the user interaction platform, and user behaviors and interaction combinations are recorded to obtain a preliminary behavior data set. Based on the preliminary behavior data set, the pattern of repeated viewing and interaction combinations is analyzed, and a preset threshold is used for screening to determine high-frequency interaction behavior characteristics. If the high-frequency interaction behavior characteristics meet the preset significance conditions, they are mapped to the user interest distribution matrix, and the corresponding interest weights are updated to obtain an adjusted user interest distribution matrix. With respect to the adjusted user interest distribution matrix, the recommended short drama content features related to the short drama content sequence are obtained, and the matching degree between the recommended short drama content features and the user interests is determined. If the matching degree is lower than the preset threshold, the short drama content sequence is re-sorted by a collaborative filtering algorithm to obtain an optimized recommended short drama content sequence. Based on the optimized recommended short drama content sequence, real-time feedback is provided to the user interaction platform, and new user behavior data is recorded to form a dynamically adjusted data closed loop.

[0049] Specifically, to dynamically adjust the final list of recommended short drama content, we first collect user interaction data through a real-time feedback mechanism. Suppose a user watches a movie three times in the recommended sequence, and their interaction behavior includes two likes, one comment, and one share. The system records this data in the behavior log and calculates the user's interest preference through weighted calculation. The formula is interest value = number of repeated views × 0.5 + number of likes × 0.3 + number of comments × 0.2 + number of shares × 0.1. The interest value of the movie is 3 × 0.5 + 2 × 0.3 + 1 × 0.2 + 1 × 0.1 = 2.4. The system then compares this interest value with the interest values ​​of other short drama content and updates the user interest distribution matrix. Assuming that the interest value of this type of short drama content in the original matrix is ​​1.8, the updated value is (1.8 + 2.4) / 2 = 2.1, reflecting the user's increased preference for this type of short drama content. Next, based on the updated matrix, the system uses a collaborative filtering algorithm to calculate the user's interest similarity with other users. Assuming a similarity threshold of 0.7, once similar users are found, their highly-interested skits (with an interest score greater than 2.0) are added to the recommendation pool. Finally, the system dynamically adjusts the recommendation sequence using a ranking algorithm (for example, a comprehensive score based on interest and freshness of the skit content: the formula is: comprehensive score = interest value × 0.6 + freshness × 0.4. For example, if a skit has an interest value of 2.1 and freshness of 0.8, the comprehensive score is 2.1 × 0.6 + 0.8 × 0.4 = 1.58), thus forming a closed-loop data loop. To ensure a closed-loop logic, the system can use the exposure duration of the skit content as an auxiliary indicator if the user has no significant interaction. For example, if a skit has been exposed for more than 60 seconds and has no negative feedback, the default interest score is 1.0, and this is included in the matrix update, ensuring continuous optimization of the recommendation system.

[0050] In one embodiment, the method further comprises: Obtaining real-time user scene data; the real-time user scene data includes geographic location, environmental information, and device status; Construct scene feature vectors based on user real-time scene data; Combine it with the short play content feature vector and user tag matching information for analysis; Based on the differences in user preferences in different scenarios, the calculation weight of the compatibility between the short drama content and user preferences is dynamically adjusted to optimize the preliminary recommended short drama content list.

[0051] Specifically, real-time user scenario data is collected through mobile terminal sensors, GPS positioning modules, and device logs. This data includes geographic location (e.g., city, business district, indoor / outdoor), environmental information (e.g., light intensity, noise decibels, temperature), and device status (e.g., screen brightness, network signal strength, and device type). For example, a user commuting outdoors may be located on Subway Line 1, with environmental information indicating noise levels exceeding 75 decibels and strong light levels. The device status is mobile network, and the screen brightness is at maximum. Next, based on this real-time scenario data, a scenario feature vector is constructed using feature engineering. For example, the geographic location "Subway" is encoded as [1,0,0] (subway / bus / walking), ambient noise levels exceeding 75 decibels are encoded as 1, and the device network status "Mobile Network" is encoded as 1, resulting in a scenario feature vector [1,1,1,0,...] with the same dimensionality as the scenario (e.g., 10 dimensions). This scenario feature vector is then analyzed in conjunction with the short drama content feature vector (e.g., plot type vector, emotional tone vector), and user tag matching information (e.g., "young female / preferred suspense / active at night"). Through tensor fusion, the three-dimensional correlation between scene, content, and user preferences is calculated. For example, in outdoor commuting scenarios, users have a higher preference for "fast-paced, dialogue-free" suspense shorts, necessitating an increased weighting of the content features "compactness" and "visual impact." Based on the differences in user preferences in different scenarios, the weightings for calculating the compatibility of short drama content with user preferences are dynamically adjusted. For example, in noisy outdoor scenes, the weighting of "sound clarity" is increased from 0.2 to 0.4, and the weighting of "plot complexity" is reduced from 0.5 to 0.3; the opposite is true in quiet indoor scenes. Based on the adjusted weightings, the similarity score is recalculated, content that does not match the current scenario is filtered out (e.g., eliminating slow-paced plots that require deep immersion in outdoor scenes). The initial recommendation list is then reordered to generate optimized recommendations tailored to the current scenario. This process, by real-time sensing of the user's environment and device status, incorporates the scene dimension into the recommendation model, achieving three-dimensional dynamic matching of "scene, content, and user," enhancing the user viewing experience in different scenarios.

[0052] In one embodiment, a weighted aggregation calculation is performed on the multi-dimensional data in the initial user behavior dataset using a preset weight distribution algorithm to obtain the user's preliminary preference score for the skit content, including: Construct the state space of the reinforcement learning model; combine the normalized features of the user's multi-dimensional behavior data, historical weight distribution results, and current recommendation effect indicators into a state vector; Define the action space as the weight adjustment of each behavior data dimension, and express the weight update strategy through bounded continuous values; The user's actual interaction behavior with the recommended short drama is used as the immediate reward, and the long-term reward is calculated by combining the matching degree between the predicted preference and the actual preference; Use deep reinforcement learning algorithms to train policy networks and optimize weight adjustment strategies through experience replay; Based on the trained policy network, dynamic weights are generated, and the viewing time ratio, like frequency, and comment interaction depth in the initial user behavior dataset are weighted and aggregated to calculate the user's preliminary preference score for the short drama content.

[0053] In one embodiment, the method further comprises: Extract the average weekly viewing time, monthly interaction times, and comment sentiment intensity data from user historical behavior data and calculate the activity index A; Determine the fusion weight of personalized recommendation and diversity recommendation based on the calculated activity index A; Based on the user interest distribution matrix, a collaborative filtering algorithm is used to calculate the similarity scores between candidate short dramas and user preferences, and the top N short dramas are selected to generate a personalized recommendation list; Construct category coverage evaluation function and novelty evaluation indicators; The personalized recommendation list and the diversity recommendation list are integrated according to the determined integration weight; The generated final recommendation list is checked to ensure that any 10 consecutive short plays in the list contain at least 4 different categories; if this condition is not met, the list is adjusted until it meets the requirements; When the user activity index A is less than the preset value, the proportion of categories that the user has not watched in the final recommendation list is checked; if the proportion is lower than the first proportion, short dramas of the corresponding category are added from the candidate library until it is greater than the first proportion.

[0054] Specifically, data on average weekly viewing time (the total amount of time users watch short dramas per week), monthly interactions (the total number of likes, comments, and shares per month), and comment sentiment (using natural language processing to analyze the sentiment of user comments, outputting a numerical value between -1 and 1, with higher values ​​indicating more positive sentiment) are extracted from user historical behavior data. After normalization, this data is weighted to calculate an activity index A. The calculated activity index A is used to determine the weights for combining personalized and diverse recommendations. For example, when A > 0.7, indicating high user activity, the personalized recommendation weight is set to 0.8, while the diverse recommendation weight is set to 0.2 to enhance user interest matching. When A < 0.3, indicating low user activity, the personalized recommendation weight is reduced to 0.4, while the diverse recommendation weight is increased to 0.6 to explore potential user interests. Based on the user interest distribution matrix, a collaborative filtering algorithm is used to calculate the similarity scores between candidate short dramas and the user's preferences. These short dramas are sorted from high to low by score, and the top N (e.g., 50) scoring short dramas are selected to generate a personalized recommendation list. The collaborative filtering algorithm analyzes user historical behavior data to find other users with similar interests to the target user, and then recommends content that similar users like.

[0055] A category coverage evaluation function (e.g., calculating the ratio of the number of categories covered in the recommendation list to the total number of categories) and novelty evaluation metrics (e.g., the proportion of categories the user has not viewed and the proportion of newly released content) are constructed to measure the richness and freshness of the diverse recommendation list. For example, a diverse recommendation list must meet category coverage ≥ 60% and novel content ≥ 40%. A linear combination of the personalized and diverse recommendation lists is performed according to the determined fusion weights to generate a preliminary final recommendation list. For example, the combined list may contain 60% personalized content and 40% diverse content. The list is then reordered based on the combined score (personalization score × 0.6 + diversity score × 0.4). A continuity check is performed on the generated final recommendation list to ensure that any 10 consecutive short dramas in the list contain at least four different categories (e.g., the intersection of suspense, comedy, science fiction, and family). If this condition is not met, adjustments are made by swapping positions, inserting new categories, etc., until the list meets the requirements to avoid excessive concentration of similar content that affects the user experience. When the user activity index A is less than a preset value (e.g., 0.4), the proportion of categories in the final recommendation list that the user has not viewed is checked. If this percentage is lower than the first percentage (e.g., 30%), short dramas in the corresponding categories are retrieved and added from the candidate library until the percentage of unwatched categories exceeds the first percentage, thereby stimulating the exploration interest of low-activity users. For example, if a user's historical preferences are concentrated in comedy (accounting for 80%), and the percentage of unwatched categories (such as technology and education) in the recommended list is only 20%, then 3-5 technology / education short dramas are added to increase the percentage of unwatched categories to 35%. The above process quantifies user activity and dynamically balances the accuracy of personalized recommendations with the exploratory nature of diverse recommendations. Through continuous checks and low-activity intervention strategies, it ensures that the recommended list not only meets user interests but also expands content coverage. This is especially suitable for the differentiated needs of users with different activity levels, improving overall recommendation effectiveness and user retention.

[0056] Reference Figure 2 The second embodiment of the present invention provides a short drama recommendation system based on big data, including: The first processing module is used to collect multi-dimensional behavior data of users to form an initial behavior data set; A second processing module is configured to calculate the user's preference scores for different categories of skit content based on the initial behavior data set to obtain a user interest distribution matrix; The third processing module is used to: analyze the content category distribution and viewing habits of the short dramas based on the user interest distribution matrix to divide the user groups and mark the user category labels; A fourth processing module is configured to: obtain video metadata from a skit content database to form an initial skit content feature set; A fifth processing module is configured to extract a skit content semantic vector based on the initial skit content feature set and construct a content semantic feature matrix in combination with an emotional tone matching requirement; a sixth processing module, configured to: if there is a matching requirement between the content semantic feature matrix and the user category label, determine the degree of compatibility between the short play content and the user preference and output a preliminary list of recommended short play content; The seventh processing module is used to: sort and optimize the recommended short drama contents based on the preliminary recommended short drama content list using the click rate and completion rate indicators in the historical distribution data to obtain a final recommended short drama content list.

[0057] It should be noted that the short drama recommendation system based on big data provided in an embodiment of the present invention is used to execute all the process steps of the short drama recommendation method based on big data in the above embodiment. The working principles and beneficial effects of the two correspond one to one, so they will not be repeated here.

[0058] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A short drama recommendation method based on big data, characterized in that: include: Collect multi-dimensional behavioral data of users to form an initial behavioral data set; Calculating the user's preference scores for different categories of skit content based on the initial behavior dataset to obtain a user interest distribution matrix; Analyzing the content category distribution and viewing habits of short dramas based on the user interest distribution matrix to divide user groups and label user categories; Obtaining video metadata from a skit content database to form an initial skit content feature set; Extracting the skit content semantic vector based on the initial skit content feature set and building a content semantic feature matrix in combination with the emotional tone matching requirements; If there is a matching requirement between the content semantic feature matrix and the user category label, the degree of compatibility between the short play content and the user preference is determined and a preliminary list of recommended short play content is output; Based on the preliminary recommended short play content list, the recommended short play content is sorted and optimized using the click rate and completion rate indicators in the historical distribution data to obtain the final recommended short play content list.

2. The method for recommending short plays based on big data according to claim 1, characterized in that: Collect multi-dimensional behavioral data of users to form an initial behavioral data set, including: Obtain multi-dimensional behavioral data from the interaction log data, including viewing time percentage, like frequency, comment interaction depth, viewing habits during different time periods, and repeat viewing times, and perform preliminary cleaning and formatting to obtain a behavioral data set; Based on the behavioral data set, normalizing the viewing time ratio and the like frequency to obtain a standardized behavioral feature set; Determining whether the viewing habit data for a time period is within a first preset threshold range based on the behavioral feature set; if so, correlating and matching the viewing habit data for the time period with the repeated viewing count data to determine the intensity of the user's viewing preference for the corresponding time period; Extract comment interaction depth data based on viewing preference strength judgment results, and use cluster analysis method to group users' interactive behaviors to obtain user interaction pattern classification; Based on the results of the interaction pattern classification, the correlation data of the like frequency and the number of repeated viewings is obtained. If the like frequency is higher than a second preset threshold, the like frequency and the number of repeated viewings are weighted to determine the user's preference weight for the skit content; Based on the calculation results of the preference weights, all dimensions of the multi-dimensional behavior data are integrated and a decision tree model is used to conduct a comprehensive analysis of user behavior characteristics to obtain an initial behavior data set.

3. The method for recommending short plays based on big data according to claim 1, characterized in that: Based on the initial behavior dataset, the user's preference scores for different categories of short drama content are calculated to obtain a user interest distribution matrix, including: Obtain the initial user behavior dataset; Performing weighted aggregation calculation on the multi-dimensional data in the initial user behavior dataset using a preset weight distribution algorithm to obtain a preliminary preference score of the user for the content of the skit; Extracting data related to user behavior and skit content based on the preliminary preference scores; if a user's preference score exceeds a preset threshold, marking the user behavior data and determining it as a high-preference user group; Obtaining viewing time and viewing completion rate data for the high-preference user group, and determining the user's preference for a specific short drama content category by comparing and analyzing the distribution of their behavior on different short drama content; Based on the tendency judgment results, the like frequency and comment depth data are integrated and the user behaviors are grouped using a cluster analysis method to obtain a behavioral pattern classification of different user groups; Extracting relevant data of the user interest distribution matrix through the behavioral pattern classification; wherein, if the user interest distribution matrix data of a certain classification group meets the preset conditions, the preference score is adjusted twice to determine the final user interest distribution; According to the final user interest distribution and the characteristic data of the short play content, a decision tree model is used to associate the user behavior with the short play content characteristics to obtain the user's preference for the short play content category and obtain the user interest distribution matrix.

4. The method for recommending short plays based on big data according to claim 1, characterized in that: Analyze the content category distribution and viewing habits of short dramas based on the user interest distribution matrix to divide user groups and label user categories, including: Obtaining the user interest distribution matrix and using a clustering method to stratify the user groups to obtain preliminary user category division results; Based on the preliminary user category classification results, obtain the short drama content category data and distribution feature data in the user interest distribution matrix; if the distribution feature of a certain user category meets the preset threshold, normalize the user preference scores within the category to determine an adjusted category feature set; Based on the adjusted category feature set, analyzing the correlation data between viewing habits and viewing preferences in different time periods, using statistical tools to calculate the user's activity in different time periods in segments, and obtaining the time period preference distribution of the user group; Based on the time period preference distribution, the category preference weights related to the time period are determined in combination with the skit content categories and the preliminarily labeled user category tags; if the activity level of a certain time period is higher than a preset threshold, the skit content category preferences of the user group in that time period are weighted; Compare the time period-related category preference weight distributions under different user category labels to determine the user group's preference characteristics for the short drama content category in the corresponding time period; The tendency features and user interest data are integrated, and a clustering method is used to adjust the user category labels of the user group to obtain the final user category labels.

5. The method for recommending short plays based on big data according to claim 1, characterized in that: Extracting the skit content semantic vector based on the initial skit content feature set and building a content semantic feature matrix based on the emotional tone matching requirements, including: Obtain the plot type and emotional tone description information from the initial skit content feature set, use semantic analysis technology to perform word segmentation and text normalization on the plot type and emotional tone description information, and obtain standardized skit content semantic data; Performing vectorization based on the standardized skit content semantic data using a pre-established word embedding model to determine a semantic vector for the skit content semantic data; Based on the semantic vector, determining the degree of correlation between the semantic vector and the emotional tone by calculating the vector space distance; if the correlation degree is lower than a preset threshold, returning to the previous step and re-obtaining the semantic vector; Based on the adjusted semantic vector, a mapping of the correspondence between the content semantics and the emotional tone of the skit is constructed to generate a preliminary content semantic feature matrix; For the preliminary content semantic feature matrix, a clustering tool is used to group the distribution characteristics of plot types and emotional tones to determine the matrix structure of the classified content semantic feature matrix; The integrity of the matrix structure after classification is verified. If the integrity of a certain classification is lower than a preset threshold, the content semantic feature matrix after classification is supplemented to obtain an optimized content semantic feature matrix.

6. The method for recommending short plays based on big data according to claim 1, characterized in that: If there is a matching requirement between the content semantic feature matrix and the user category label, the degree of compatibility between the short play content and the user preference is determined and a preliminary list of recommended short play content is output, including: Vectorize the semantics of the skit content to obtain the feature vector of the skit content; Use the pre-established user tag database to obtain tag matching information related to user categories and determine the scope of the preliminary matching user group; If the user group range of the preliminary match and the historical interaction data have an intersection, then the association strength between the user and the skit content is calculated to obtain an interaction weight value; For the interaction weight value, the similarity score between the short play content and the user preference is calculated by integrating the short play content type and preference information to obtain the degree of fit; If the similarity score is higher than a preset threshold, the corresponding short play content will be included in the preliminary recommended short play content list; The order of recommended short play contents is determined by combining the user category and the tag matching information to obtain a preliminary recommended short play content list.

7. The method for recommending short plays based on big data according to claim 1, characterized in that: Based on the preliminary recommended short drama content list, the recommended short drama content is sorted and optimized using the click-through rate and completion rate indicators in the historical distribution data to obtain a final recommended short drama content list, including: Obtain click-through rate and completion rate indicators from historical data, analyze the preliminary list of recommended short drama content, and obtain initial evaluation results; Based on the initial evaluation results, the recommended short play contents are screened using a preset dynamic threshold. If the click-through rate and completion rate of a short play content are both lower than the dynamic threshold, it is removed from the preliminary list to determine a screened set of short play contents. Analyzing the distribution ratio of the skit content categories for the screened skit content set, and if the distribution ratio of a certain type of skit content exceeds a preset range, adjusting the skit content to obtain an adjusted skit content set; Based on the click-through rate and completion rate of the adjusted short drama content set, a logistic regression model is used to perform sorting optimization to obtain a preliminary sorting sequence; According to the preliminary sorting sequence, the category distribution characteristics of the short play content after sorting optimization are obtained. If the concentration of a certain category of short play content in the sequence is higher than the preset standard, its position is adjusted twice. After determining that no secondary adjustment is required, the final recommended short play content list is obtained.

8. The method for recommending short plays based on big data according to claim 1, characterized in that: The method further comprises: Obtaining real-time user scene data; the real-time user scene data includes geographic location, environmental information, and device status; Build scene feature vectors based on user real-time scene data; Combine it with the short play content feature vector and user tag matching information for analysis; Based on the differences in user preferences in different scenarios, the calculation weight of the compatibility between the short drama content and user preferences is dynamically adjusted to optimize the preliminary recommended short drama content list.

9. The method for recommending short plays based on big data according to claim 3, characterized in that: A weighted aggregation calculation is performed on the multi-dimensional data in the initial user behavior dataset using a preset weight distribution algorithm to obtain the user's preliminary preference score for the skit content, including: Construct the state space of the reinforcement learning model; combine the normalized features of the user's multi-dimensional behavior data, historical weight distribution results, and current recommendation effect indicators into a state vector; Define the action space as the weight adjustment of each behavior data dimension, and express the weight update strategy through bounded continuous values; The user's actual interaction behavior with the recommended short drama is used as the immediate reward, and the long-term reward is calculated by combining the matching degree between the predicted preference and the actual preference; Use deep reinforcement learning algorithms to train policy networks and optimize weight adjustment strategies through experience replay; Generating dynamic weights based on the trained policy network, performing weighted aggregation on the viewing time percentage, like frequency, and comment interaction depth in the initial user behavior dataset, and calculating the user's preliminary preference score for the short drama content; the method further includes: Extract the average weekly viewing time, monthly interaction times, and comment sentiment intensity data from user historical behavior data and calculate the activity index A; Determine the fusion weight of personalized recommendation and diversity recommendation based on the calculated activity index A; Based on the user interest distribution matrix, a collaborative filtering algorithm is used to calculate the similarity scores between candidate short dramas and user preferences, and the top N short dramas are selected to generate a personalized recommendation list; Construct category coverage evaluation function and novelty evaluation indicators; The personalized recommendation list and the diversity recommendation list are integrated according to the determined integration weight; The generated final recommendation list is checked to ensure that any 10 consecutive short plays in the list contain at least 4 different categories; if this condition is not met, the list is adjusted until it meets the requirements; When the user activity index A is less than the preset value, the proportion of categories that the user has not watched in the final recommendation list is checked; if the proportion is lower than the first proportion, short dramas of the corresponding category are added from the candidate library until it is greater than the first proportion.

10. A short drama recommendation system based on big data, characterized in that: include: The first processing module is used to collect multi-dimensional behavior data of users to form an initial behavior data set; The second processing module is configured to calculate the user's preference scores for different categories of skit contents based on the initial behavior data set to obtain a user interest distribution matrix; The third processing module is used to: analyze the content category distribution and viewing habits of the short dramas based on the user interest distribution matrix to divide the user groups and mark the user category labels; A fourth processing module is configured to: obtain video metadata from a skit content database to form an initial skit content feature set; A fifth processing module is configured to extract a skit content semantic vector based on the initial skit content feature set and construct a content semantic feature matrix in combination with an emotional tone matching requirement; a sixth processing module, configured to: if there is a matching requirement between the content semantic feature matrix and the user category label, determine the degree of compatibility between the short play content and the user preference and output a preliminary list of recommended short play content; The seventh processing module is used to: sort and optimize the recommended short drama contents based on the preliminary recommended short drama content list using the click rate and completion rate indicators in the historical distribution data to obtain a final recommended short drama content list.

Citation Information

Cited By

  • Personalized video recommendation method and system based on big data

    CN120769084A

  • Online video content intelligent pushing method combined with learning interest model

    CN121000906A

  • Agricultural material information management system based on cloud platform

    CN121235798A

  • A cloud-based agricultural input information management system

    CN121235798B