Commodity recommendation method and system based on movie watching crowd characteristics
By extracting and analyzing the real-time viewing data of users and building dynamic viewing characteristics and interest models, the problem that the existing recommendation system cannot adapt to the changes in viewing population preferences is solved, personalized and customized product recommendations are achieved, and the accuracy and user satisfaction of the recommendation system are improved.
Patent Information
- Application Number
- CN202510314039.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing product recommendation system fails to fully consider the special characteristics of the viewer, resulting in the inability to adapt to changes in user preferences in a timely manner, resulting in a decline in the quality of recommendation results.
By collecting the user's real-time movie viewing data, extracting user characteristics, video features and movie viewing interaction features, using technologies such as autoencoder, multimodal video feature extraction framework, recurrent neural network and graph neural network, a user's movie viewing behavior map and interest evolution map are built, combined with online learning algorithms and influence communication models, the product recommendation resource pool is dynamically updated, and product recommendations that match the user's current interests are generated.
It has achieved a keen perception of changes in the preferences of the movie audience, provided the latest and most accurate basis for user preferences, improved the quality and user satisfaction of the recommendation results, avoided the lag and homogeneity of the recommendation results, and improved user shopping experience and platform operation efficiency.
Smart Images

Figure CN120338909A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method and system for recommending products based on the characteristics of movie-watching audiences. Background Art
[0002] With the popularization of the Internet and the rapid development of e-commerce, users are faced with a vast selection of products. How to find products that meet their needs and interests from among numerous products has become a challenge. Recommendation systems have emerged, which analyze users' preferences and recommend products that users may be interested in, thereby improving the users' shopping experience and the operational efficiency of the platform.
[0003] However, most existing product recommendation systems are based on general user behavior data and do not fully consider the special characteristics of movie-watching audiences. Users' preferences may change due to the release of new movies, and existing product recommendation systems cannot adapt to these changes in a timely manner, resulting in a decline in the quality of the recommendation results.
[0004] No effective solution has been proposed to address the above problems. Summary of the Invention
[0005] Embodiments of this application provide a method and system for recommending products based on the characteristics of movie-watching audiences to solve the above technical problems.
[0006] This application provides a method for recommending products based on the characteristics of movie-watching audiences, including: collecting real-time movie-watching data of users from different movie-watching platforms; extracting user characteristics, movie characteristics, and movie-watching interaction characteristics from the collected real-time movie-watching data of users; where the movie-watching interaction characteristics include movie-watching behavior characteristics, rating and feedback characteristics, preference and interest characteristics, and social interaction characteristics; obtaining the resources to be recommended in the product recommendation resource pool; extracting the product characteristics of the resources to be recommended; and updating the resources to be recommended in the product recommendation resource pool according to the product characteristics of the resources to be recommended, the user characteristics, the movie characteristics, and the movie-watching interaction characteristics.
[0007] The present application provides a commodity recommendation system based on the characteristics of the movie-watching population, including: a movie-watching data collection module for collecting real-time movie-watching data of users from different movie-watching platforms; a movie-watching feature extraction module for extracting user features, movie features, and movie-watching interaction features from the collected real-time movie-watching data of users; wherein, the movie-watching interaction features include movie-watching behavior features, rating and feedback features, preference and interest features, and social interaction features; a to-be-recommended resource acquisition module for acquiring to-be-recommended resources in a commodity recommendation resource pool; a commodity feature extraction module for extracting commodity features of the to-be-recommended resources; and a to-be-recommended resource update module for updating the to-be-recommended resources in the commodity recommendation resource pool according to the commodity features of the to-be-recommended resources, the user features, the movie features, and the movie-watching interaction features.
[0008] Further, the commodity recommendation system based on the characteristics of the movie-watching population further includes:
[0009] a resource classification and label generation module for classifying the to-be-recommended resources in the commodity recommendation resource pool according to the commodity features of the to-be-recommended resources in dimensions of commodity type, brand, price range, and applicable scenario, so as to generate commodity label information matching the to-be-recommended resources.
[0010] Further, the movie-watching feature extraction module extracts user features and movie features from the collected real-time movie-watching data of users, and is configured to:
[0011] Use an autoencoder to encode the real-time movie-watching data of users, so as to map the real-time movie-watching data of users into an embedding space to obtain an embedding representation of users; wherein, the number of features included in the embedding representation of users is less than the number of features in the real-time movie-watching data of users;
[0012] Determine the user features based on the embedding representation of users;
[0013] Use a multi-modal video feature extraction framework to perform sentiment analysis and theme mining on the real-time movie-watching data of users, so as to construct a movie knowledge graph;
[0014] Analyze the movie knowledge graph to obtain the movie features; wherein, the movie features include creation and performance features, visual appearance features, optical flow features, and audio features of the movie; the creation and performance features include creation subject features, performance subject features, and content classification features.
[0015] Further, the movie-watching feature extraction module extracts movie-watching interaction features from the collected real-time movie-watching data of users, and is configured to:
[0016] Using a recurrent neural network to model the viewing behavior sequence data in the user's real-time viewing data, combined with a scene perception algorithm, the user's viewing device, time and location are integrated into the viewing behavior sequence data;
[0017] Determining a set of movie-watching behaviors and a correlation relationship between the movie-watching behaviors according to the movie-watching behavior sequence data;
[0018] Each movie-watching behavior in the movie-watching behavior set is regarded as a node, and the edges between the nodes represent the association weights between the movie-watching behaviors, so as to construct a user movie-watching behavior graph;
[0019] Analyzing the user's movie-watching behavior graph through a graph neural network to mine potential patterns and rules of the user's movie-watching behavior and obtain the movie-watching behavior characteristics;
[0020] Perform text sentiment analysis on the rating feedback data in the user's real-time viewing data to extract the user's text sentiment features and the user's voice sentiment features when rating;
[0021] The text emotion feature and the speech emotion feature are combined to obtain the score and feedback feature;
[0022] By combining online learning algorithms with time series analysis algorithms, the user's preference features are extracted from the user's real-time viewing data, and a user's interest evolution map is constructed; wherein the interest evolution map is used to record the user's current preferences and interests, and to trace the source and development path of the interest;
[0023] According to the interest evolution graph, mining the interest associations between users in different fields to obtain the preference and interest characteristics;
[0024] Constructing a dynamic user movie-watching social network; wherein the user movie-watching social network is used to represent the user's social relationships and interactive behaviors;
[0025] Based on the user's movie-watching social network and combined with the influence propagation model, analyze the user's influence, emotional tendency and information propagation path in the social network;
[0026] The social interaction characteristics are determined based on the user's influence, emotional inclination and information dissemination path in the social network.
[0027] Further, the to-be-recommended resource updating module updates the to-be-recommended resources in the commodity recommendation resource pool according to the commodity features of the to-be-recommended resources, the user features, the film features, and the viewing interaction features, and is configured as follows:
[0028] Use the K-means based clustering algorithm to fill in the missing values of the user features, the missing values of the movie features, and the missing values of the viewing interaction features;
[0029] Among them, for the categorical features in the user features, the movie features, and the viewing interaction features, adopt adaptive coding and combine it with leave-one-out coding for processing; for the continuous features in the user features, the movie features, and the viewing interaction features, adopt a custom quantile Min-Max normalization for processing. First, segment by quantiles, and then perform Min-Max normalization within each segment;
[0030] Based on the filled viewing interaction features, the filled user features, and the filled movie features, derive user preference type score features and weighted viewing popularity rating features through the term frequency-inverse document frequency algorithm; among them, the weighted viewing popularity rating feature is obtained by weighted fusion of the filled movie features and the user preference type score features;
[0031] Perform singular value decomposition on the filled viewing interaction features, the filled user features, and the filled movie features to obtain singular values and singular vectors;
[0032] Select the singular vectors corresponding to the first y1 largest singular values as the feature vectors after dimensionality reduction; where y1 is a natural number, and the value of y1 is determined by the cumulative explained variance ratio;
[0033] Project the filled viewing interaction features, the filled user features, and the filled movie features onto the feature vectors after dimensionality reduction to obtain the data representation after dimensionality reduction;
[0034] Based on the data representation after dimensionality reduction, construct an intra-class scatter matrix and an inter-class scatter matrix; among them, the intra-class scatter matrix is used to represent the scatter of samples within the same category; the inter-class scatter matrix is used to represent the scatter of samples between different categories;
[0035] Calculate the product of the inverse matrix of the intra-class scatter matrix and the inter-class scatter matrix to obtain a projection matrix;
[0036] Calculate the eigenvalues and eigenvectors of the projection matrix; where the eigenvalues of the projection matrix represent the ability of the eigenvectors of the projection matrix in class discrimination;
[0037] According to the user preference type score features and the weighted viewing popularity rating features, select the first y2 generalized eigenvectors from the eigenvectors of the projection matrix; where y2 is a natural number;
[0038] Project the filled viewing interaction features, filled user features, and filled movie features onto the generalized feature vector to obtain a projected feature representation;
[0039] Evaluate and optimize the projected feature representation to obtain an optimized projected viewing feature;
[0040] Update the resources to be recommended in the product recommendation resource pool according to the product label information matching the resource to be recommended and the optimized projected viewing feature.
[0041] Further, the evaluating and optimizing the projected feature representation to obtain an optimized projected viewing feature is configured as:
[0042] Conduct an F-test on the projected feature representation, process the features that do not conform to the normal distribution, and reconstruct the features with significant F-statistics;
[0043] Use the gradient boosting tree as the base model to perform recursive feature elimination on the projected features after the F-test, eliminating 15% of the features with the smallest weight coefficients in each round;
[0044] Use the XGBoost model and the LightGBM model to train the projected features after recursive feature elimination, and calculate the first importance of the projected features after recursive feature elimination in the XGBoost model and the second importance in the LightGBM model respectively;
[0045] Perform weighted fusion on the first importance and the second importance of the projected features after recursive feature elimination to obtain the reference importance of the projected features after recursive feature elimination;
[0046] Retain the top 80% of the features with reference importance to form a feature candidate set;
[0047] According to the data distribution of the feature candidate set, adaptively select 7-fold cross-validation or 5-fold cross-validation;
[0048] Use 7-fold cross-validation or 5-fold cross-validation, combined with stratified sampling, to evaluate the feature candidate set;
[0049] According to the evaluation results of the feature candidate set, adopt a feature contribution degree analysis algorithm based on the Shapley value and the LIME algorithm to perform adaptive sensitivity analysis on each feature in the feature candidate set, and dynamically set the sensitivity threshold according to the importance and relevance of the features;
[0050] For the features below the sensitivity threshold, perform feature pruning;
[0051] For the pruned feature candidate set, an adaptive feature fusion algorithm based on a deep learning attention mechanism is used to fuse the features in the pruned feature candidate set to obtain the optimized projection viewing features.
[0052] Further, the within-class scatter matrix is S ω ; the between-class scatter matrix is S c ; the projection matrix is inv(S ω )×S c , and the calculation process of the projection matrix is as follows:
[0053] Calculate the inverse matrix inv(S ω ) of S ω ;
[0054] Multiply inv(S ω ) by S c to obtain inv(S ω )×S c ;
[0055] Among them, the inverse matrix inv(S ω ) is a square matrix; the eigenvectors of inv(S ω )×S c are used to maximize between-class separation and minimize within-class separation.
[0056] Further, the viewing data collection module collects real-time viewing data of users from different viewing platforms, and is configured to: collect cinema ticket sales system data, online video platform viewing record data, and viewing-related discussion data on social media; integrate the collected cinema ticket sales system data, online video platform viewing record data, and viewing-related discussion data on social media into the real-time viewing data of users.
[0057] Based on the embodiments provided in this application, the movie-watching data collection module collects the movie-watching data of users on different movie-watching platforms in real time, enabling the timely acquisition of multi-dimensional information such as users' movie-watching behaviors, rating feedback, preference interests, and social interactions for new movies. This real-time data collection method allows the system to keenly perceive the subtle changes in the preferences of the movie-watching population, such as the interest transfer caused by the release of new movies, thereby providing the latest and most accurate basis for user preferences for the recommendation system and effectively solving the problem that the existing recommendation system cannot adapt to the changes in the preferences of the movie-watching population in a timely manner. The movie-watching feature extraction module comprehensively extracts user features, movie features, and movie-watching interaction features, where the movie-watching interaction features cover multiple aspects such as movie-watching behavior features, rating and feedback features, preference and interest features, and social interaction features. This all-round and multi-level feature extraction method can comprehensively depict users' movie-watching habits and preferences, as well as the complex interaction relationship between users and movies. Compared with the recommendation system based only on general user behavior data, the system of this application can more accurately understand users' personalized needs, and then provide commodity recommendations that better match users' current interests and preferences, significantly improving the quality of recommendation results and user satisfaction. The to-be-recommended resource update module dynamically updates the to-be-recommended resources in the commodity recommendation resource pool according to the various features extracted. This mechanism ensures that the recommended resources always match the latest movie-watching features and preferences of users, can timely introduce new commodities related to users' current movie-watching interests, and eliminate old commodities that no longer meet users' needs. Through this dynamic update, the recommendation system can continuously provide users with fresh and accurate recommendation content, avoiding the lag and homogenization of recommendation results, and further enhancing users' shopping experience and the operation efficiency of the platform. The recommendation system of this application is specifically designed for the movie-watching population, fully considering the special features and needs of this group. Compared with the general recommendation system, it can better adapt to the diverse and personalized characteristics of the movie-watching population and provide customized recommendation services for users with different movie-watching preferences and different movie-watching frequencies. Description of the Drawings
[0058] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and the illustrative embodiments and descriptions thereof are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0059] Figure 1 It is a flowchart of an optional commodity recommendation method based on the characteristics of the movie-watching population according to the embodiments of this application;
[0060] Figure 2 It is a structural diagram of an optional commodity recommendation system based on the characteristics of the movie-watching population according to the embodiments of this application.
[0061] The realization of the purpose of the present invention, functional features, and advantages will be further described with reference to the embodiments and the drawings. Detailed implementation manners
[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0063] Optionally, as Figure 1 shown, the present application provides a commodity recommendation method based on the characteristics of the movie-watching crowd, including:
[0064] S101, collecting real-time movie-watching data of users from different movie-watching platforms;
[0065] S102, extracting user characteristics, movie characteristics, and movie-watching interaction characteristics from the collected real-time movie-watching data of users; among them, the movie-watching interaction characteristics include movie-watching behavior characteristics, rating and feedback characteristics, preference and interest characteristics, and social interaction characteristics;
[0066] In this embodiment, the movie-watching behavior characteristics include:
[0067] Movie-watching frequency: the number of times a user watches movies within a specific time period, reflecting the user's movie-watching activity. For example, the number of times of watching movies per month;
[0068] Movie-watching duration: the average duration of each movie watched by the user, reflecting the user's movie-watching concentration. For example, the duration of each movie-watching on average;
[0069] Movie-watching time: the preferred time period for the user to watch movies, such as weekends, evenings, holidays, etc., reflecting the user's movie-watching habit. For example, the user mainly watches movies between 8 pm and 10 pm;
[0070] Movie-watching device: the type of device used by the user to watch movies, such as mobile phones, tablets, computers, TVs, etc., reflecting the user's movie-watching device preference. For example, the user mainly uses mobile phones to watch movies.
[0071] The rating and feedback characteristics include:
[0072] Rating distribution: the rating situation of the user for different movies, reflecting the user's preference intensity. For example, the user generally gives high ratings to science fiction movies;
[0073] Comment content: the comment content of the user on the movie, reflecting the user's detailed preferences and emotional tendencies. For example, the user mentions in the comment that "like the special effects in the movie" or "don't like the plot of the movie";
[0074] Likes and shares: The like and share behaviors of users towards movies reflect their willingness to recommend. For example, users often like and share their favorite movies to social media;
[0075] Preference and interest characteristics include:
[0076] Movie genre preference: The movie genres preferred by users, such as science fiction, action, romance, comedy, etc., reflect the user's interest areas. For example, users mainly watch science fiction and action movies;
[0077] Director and actor preference: The directors and actors preferred by users reflect the user's specific preferences. For example, users like movies directed by Christopher Nolan or movies starring Tom Cruise;
[0078] Theme and subject matter preference: The movie themes and subject matters preferred by users, such as science fiction adventures, historical wars, love stories, etc., reflect the user's deep - seated interests. For example, users like war movies set in historical backgrounds;
[0079] Social interaction characteristics include:
[0080] Social network interaction: The interaction behaviors related to movies on social networks, such as following the official movie accounts, participating in movie - related topic discussions, etc., reflect the user's social preferences. For example, users often participate in movie - related Weibo topic discussions;
[0081] Friend recommendations: The behavior of users watching movies based on friend recommendations reflects the user's social influence. For example, users often watch movies based on their friends' recommendations;
[0082] Movie - watching group: With whom users often watch movies reflects the user's social circle and movie - watching habits. For example, users often watch movies with their family or friends.
[0083] S103. Obtain the resources to be recommended in the product recommendation resource pool; extract the product characteristics of the resources to be recommended;
[0084] S104. Update the resources to be recommended in the product recommendation resource pool according to the product characteristics, user characteristics, movie characteristics, and movie - watching interaction characteristics of the resources to be recommended.
[0085] Based on the embodiments provided in this application, the movie viewing data collection module collects the movie viewing data of users on different movie viewing platforms in real time, and can timely obtain multi-dimensional information such as the movie viewing behavior, rating feedback, preference interests, and social interactions of users for new movies. This real-time data collection method enables the system to keenly perceive the subtle changes in the preferences of the movie viewing population, such as the interest transfer caused by the release of new movies, thereby providing the latest and most accurate basis for user preferences for the recommendation system, effectively solving the problem that the existing recommendation system cannot adapt to the changes in the preferences of the movie viewing population in a timely manner. The movie viewing feature extraction module comprehensively extracts user features, movie features, and movie viewing interaction features. Among them, the movie viewing interaction features cover multiple aspects such as movie viewing behavior features, rating and feedback features, preference and interest features, and social interaction features. This all-round and multi-level feature extraction method can comprehensively depict the movie viewing habits and preferences of users, as well as the complex interaction relationship between users and movies. Compared with the recommendation system based only on general user behavior data, the system of this application can more accurately understand the personalized needs of users, and then provide product recommendations that better match the current interests and preferences of users, significantly improving the quality of recommendation results and user satisfaction. The to-be-recommended resource update module dynamically updates the to-be-recommended resources in the product recommendation resource pool according to the various features extracted. This mechanism ensures that the recommended resources always match the latest movie viewing features and preferences of users, can timely introduce new products related to the current movie viewing interests of users, and eliminate old products that no longer meet the needs of users. Through this dynamic update, the recommendation system can continuously provide users with fresh and accurate recommendation content, avoid the lag and homogenization of recommendation results, and further improve the shopping experience of users and the operation efficiency of the platform. The recommendation system of this application is specifically designed for the movie viewing population, fully considering the special features and needs of this group. Compared with the general recommendation system, it can better adapt to the diversification and personalization characteristics of the movie viewing population, and provide customized recommendation services for users with different movie viewing preferences and different movie viewing frequencies.
[0086] Optionally, as Figure 2 shown, this application provides a product recommendation system based on the characteristics of the movie viewing population, including:
[0087] A movie viewing data collection module 201 for collecting real-time movie viewing data of users from different movie viewing platforms;
[0088] A movie viewing feature extraction module 202 for extracting user features, movie features, and movie viewing interaction features from the collected real-time movie viewing data of users; among them, the movie viewing interaction features include movie viewing behavior features, rating and feedback features, preference and interest features, and social interaction features;
[0089] A to-be-recommended resource acquisition module 203 for acquiring to-be-recommended resources in the product recommendation resource pool;
[0090] The product feature extraction module 204 is used to extract the product features of the resources to be recommended.
[0091] The resource to be recommended update module 205 is used to update the resources to be recommended in the product recommendation resource pool according to the product features, user features, film features, and viewing interaction features of the resources to be recommended.
[0092] Furthermore, the product recommendation system based on the viewing population characteristics further includes:
[0093] The resource classification and label generation module is used to classify the resources to be recommended in the product recommendation resource pool according to the product features of the resources to be recommended in terms of the product type dimension, brand dimension, price range dimension, and applicable scenario dimension, so as to generate product label information matching the resources to be recommended.
[0094] Furthermore, the viewing feature extraction module extracts user features and film features from the collected real-time viewing data of users, and is configured to:
[0095] Use an autoencoder to encode the real-time viewing data of users, so as to map the real-time viewing data of users into the embedding space to obtain the embedding representation of users, thereby realizing the dimensionality reduction of the data; wherein, the number of features included in the embedding representation of users is less than the number of features in the real-time viewing data of users;
[0096] Determine user features based on the embedding representation of users;
[0097] Use a multi-modal video feature extraction framework to perform sentiment analysis and topic mining on the real-time viewing data of users to construct a film knowledge graph; wherein, the multi-modal video feature extraction framework may include the video_features framework;
[0098] In this embodiment, perform sentiment analysis, topic modeling, etc. on text contents such as movie introductions and reviews, and extract useful topic words and sentiment information; the Transformer model in deep learning can be used for topic mining, which can more deeply understand the topic words and keywords in the user feedback text and their emotional associations;
[0099] Analyze the film knowledge graph to obtain film features; wherein, the film features include the creation and performance features, visual appearance features, optical flow features, and audio features of the film; the creation and performance features include the creation subject features, performance subject features, and content classification features.
[0100] Furthermore, the viewing feature extraction module extracts viewing interaction features from the collected real-time viewing data of users, and is configured to:
[0101] Use recurrent neural networks to model the viewing behavior sequence data in users' real-time viewing data, and combine scene perception algorithms to integrate users' viewing devices, time, and location into the viewing behavior sequence data;
[0102] Determine the movie-watching behavior set and the movie-watching behavior association relationship based on the movie-watching behavior sequence data; wherein the movie-watching behavior association relationship includes continuous viewing, interval viewing, etc.;
[0103] Each viewing behavior in the viewing behavior set is regarded as a node, and the edges between nodes represent the association weights between the viewing behaviors, so as to construct a user viewing behavior graph;
[0104] Among them, the correlation weights between various movie-watching behaviors are determined based on the following formula:
[0105]
[0106] Among them, W ij represents the association weight between movie-watching behavior i and movie-watching behavior j; v i and v j are the vector representations of the movie-watching behavior i and the movie-watching behavior j in the embedding space; cos(v i ,v j ) is the vector v i and vector v j The cosine similarity of N is used to measure the similarity of movie-watching behaviors in terms of content features; ij is the number of times that movie-watching behavior i and movie-watching behavior j appear at the same time; N i and N j are the total number of occurrences of movie-watching behavior i and movie-watching behavior j, respectively; α and β are hyperparameters used to balance the impact of content similarity and co-occurrence frequency on the association weight;
[0107] Analyze the user's movie-watching behavior graph through graph neural network to explore the potential patterns and rules of users' movie-watching behavior and obtain the characteristics of movie-watching behavior;
[0108] Perform text sentiment analysis on the rating feedback data in the user's real-time viewing data to extract the user's text sentiment features and the user's voice sentiment features when rating;
[0109] Based on the following formula, the text sentiment features and speech sentiment features are integrated to obtain the score and feedback features;
[0110] f = tanh(W t t+W v v+b)⊙m
[0111] Among them, f is the rating and feedback feature; W t t is the text sentiment feature t through the weight matrix Wt Perform a linear transformation; W v v is the voice emotion feature v that undergoes a linear transformation through the weight matrix W v Perform a linear transformation; b is the bias vector used to adjust the translation of the fused feature; tanh is the hyperbolic tangent activation function used to introduce non-linearity so that the fused feature can better capture complex emotion information; ⊙ is the element-wise multiplication; m is the modal importance vector used to adjust the contribution degrees of different modal features;
[0112] Adopt an online learning algorithm combined with a time series analysis algorithm to extract the user's preference features from the user's real-time movie-watching data and construct an interest evolution map of the user; among them, the interest evolution map is used to record the user's current preferences and interest points, as well as to trace the source and development path of the interest;
[0113] According to the interest evolution map, mine the interest associations between different fields of the user to obtain preference and interest features; among them, different fields are such as movies, music, books, etc. For example, if it is found that the user likes the soundtrack music of a certain movie, the system will automatically recommend other movies with a similar music style to the user to achieve cross-field interest expansion and recommendation;
[0114] Construct a dynamic user movie-watching social network; among them, the user movie-watching social network is used to represent the user's social relationships and interaction behaviors;
[0115] Based on the user movie-watching social network, combined with the influence propagation model, analyze the user's influence, emotional tendency and information dissemination path in the social network; among them, the influence propagation model can include but is not limited to the independent cascade model, the linear threshold model;
[0116] Determine the social interaction features according to the user's influence, emotional tendency and information dissemination path in the social network.
[0117] Furthermore, the to-be-recommended resource update module updates the to-be-recommended resources in the commodity recommendation resource pool according to the commodity features, user features, movie features and movie-watching interaction features of the to-be-recommended resources, and is configured as:
[0118] Use the K-means-based clustering algorithm to fill in the missing values of user features, movie features and movie-watching interaction features;
[0119] Among them, for the categorical features in user features, movie features and movie-watching interaction features, adopt adaptive coding and combine it with leave-one-out coding for processing; for the continuous features in user features, movie features and movie-watching interaction features, adopt a custom quantile Min-Max normalization for processing. First, segment by quantiles, and then perform Min-Max normalization within each segment;
[0120] Among them, the customized quantile Min - Max normalization is a data normalization method that combines the concept of quantiles and the traditional Min - Max normalization method;
[0121] Based on the filled viewing interaction features, filled user features, and filled film features, user preference type score features and weighted viewing popularity score features are derived through the term frequency - inverse document frequency algorithm; among them, the weighted viewing popularity score feature is obtained by weighted fusion of the filled film features and user preference type score features;
[0122] Perform singular value decomposition on the filled viewing interaction features, filled user features, and filled film features to obtain singular values and singular vectors;
[0123] Select the singular vectors corresponding to the first y1 largest singular values as the feature vectors after dimensionality reduction; where y1 is a natural number, and the value of y1 is determined by the cumulative explained variance ratio. Usually, the number of singular values with a cumulative explained variance ratio reaching 85% - 90% is selected;
[0124] Project the filled viewing interaction features, filled user features, and filled film features onto the feature vectors after dimensionality reduction to obtain the data representation after dimensionality reduction;
[0125] Based on the data representation after dimensionality reduction, construct the within - class scatter matrix and between - class scatter matrix; among them, the within - class scatter matrix is used to represent the scatter of samples within the same class; the between - class scatter matrix is used to represent the scatter of samples between different classes;
[0126] Calculate the product of the inverse matrix of the within - class scatter matrix and the between - class scatter matrix to obtain the projection matrix;
[0127] Calculate the eigenvalues of the projection matrix and the eigenvectors of the projection matrix; where the eigenvalues of the projection matrix represent the ability of the eigenvectors of the projection matrix in class discrimination;
[0128] According to the user preference type score features and weighted viewing popularity score features, select the first y2 generalized eigenvectors from the eigenvectors of the projection matrix; where y2 is a natural number, usually less than the number of classes minus one, and these eigenvectors form the projection matrix, which can maximize the difference between classes and minimize the difference within classes;
[0129] In this embodiment, the specific steps for selecting the generalized eigenvectors may include:
[0130] Construct a feature matrix: Integrate the filled viewing interaction features, filled user features, and filled film features into a feature matrix. Each row of this matrix represents a user or a film, and each column represents a feature.
[0131] Apply singular value decomposition: Perform singular value decomposition on the feature matrix. After decomposition, three matrices will be obtained, one of which is a diagonal matrix containing singular values, and the other two are orthogonal matrices. The values on the diagonal of the singular value matrix represent the importance of each eigenvector.
[0132] Select eigenvectors: Select the top y2 largest singular values from the singular value matrix. The eigenvectors corresponding to these singular values are the most important generalized eigenvectors. Specifically, select the y2 eigenvectors with the largest singular values. These eigenvectors can capture the main changes and trends in the data and are therefore called "generalized eigenvectors".
[0133] Determine the value of y2: y2 is a natural number representing the number of eigenvectors selected. The value of y2 can be determined according to specific application scenarios and requirements. Generally, the smaller the value of y2, the more the eigenvectors can capture the main information in the data, but some details may be lost; the larger the value of y2, the more details the eigenvectors can capture, but the computational complexity will also increase.
[0134] Through the above steps, the top y2 most important generalized eigenvectors can be selected from the feature matrix. These eigenvectors can effectively represent the user preference type score features and the weighted features of movie viewing popularity scores, thereby improving the accuracy and relevance of the recommendation system.
[0135] Project the filled movie viewing interaction features, filled user features, and filled movie features onto the generalized eigenvectors to obtain projected feature representations;
[0136] Evaluate and optimize the projected feature representations to obtain optimized projected movie features;
[0137] Update the resources to be recommended in the product recommendation resource pool according to the product label information matching the resource to be recommended and the optimized projected movie features.
[0138] Furthermore, evaluating and optimizing the projected feature representations to obtain optimized projected movie features is configured to:
[0139] Perform an F-test on the projected feature representations, process the features that do not conform to the normal distribution, and reconstruct the features with significant F-statistics;
[0140] Among them, the F-test is a statistical test method used to compare two or more variances to determine whether there are significant differences between them. In practical applications, the F-test is often used in analysis of variance and regression analysis. Specifically, the F-test determines whether the hypothesis holds by calculating the F-statistic. The F-statistic is obtained by comparing the ratio of the mean square between groups and the mean square within groups. If the F-statistic is large and the corresponding p-value is small, it indicates that at least one coefficient is significantly non-zero and the overall model is meaningful.
[0141] The F-statistic is used to measure the significance of a model or feature. The larger the F-statistic, the stronger the overall significance of the model and the better the fitting effect of the model. When performing the F-test, if the calculated F-statistic is greater than the critical value at a given significance level, the null hypothesis can be rejected and it is considered that the model or feature is significant;
[0142] Taking the gradient boosting tree as the base model, recursive feature elimination is performed on the projection features after the F-test. In each round, 15% of the features with the smallest weight coefficients are eliminated; for example, if the current feature set has 100 features, then in each round, 15 features with the smallest weight coefficients will be eliminated, and 85 features will be retained for the next round of training.
[0143] Use the XGBoost model and the LightGBM model to train the projection features after recursive feature elimination, and calculate the first importance of the projection features after recursive feature elimination in the XGBoost model and the second importance in the LightGBM model respectively;
[0144] Perform weighted fusion on the first importance and the second importance of the projection features after recursive feature elimination to obtain the reference importance of the projection features after recursive feature elimination;
[0145] Retain the features in the top 80% of the reference importance to form a feature candidate set;
[0146] According to the data distribution of the feature candidate set, adaptively select 7-fold cross-validation or 5-fold cross-validation;
[0147] Use 7-fold cross-validation or 5-fold cross-validation, combined with stratified sampling, to evaluate the feature candidate set; for example, a custom comprehensive evaluation metric can be used for evaluation, fusing accuracy, recall, F1-score, and ROC-AUC, and dynamically adjusting the weights of each metric according to the requirements of different business scenarios to comprehensively evaluate the final feature candidate set;
[0148] According to the evaluation results of the feature candidate set, adopt a feature contribution degree analysis algorithm based on the Shapley value and the LIME algorithm to perform adaptive sensitivity analysis on each feature in the feature candidate set, and dynamically set the sensitivity threshold according to the importance and relevance of the features;
[0149] For features below the sensitivity threshold, feature pruning is performed;
[0150] For the pruned feature candidate set, an adaptive feature fusion algorithm based on the attention mechanism of deep learning is used to fuse the features in the pruned feature candidate set to obtain optimized projection viewing features.
[0151] Furthermore, the within-class scatter matrix is S ω ; the between-class scatter matrix is S c ; the projection matrix is inv(S ω ) × S c , and the calculation process of the projection matrix is:
[0152] Calculate the inverse matrix inv(S ω ) of S ω ;
[0153] Multiply inv(S ω ) by S c to obtain inv(S ω ) × S c ;
[0154] Among them, the inverse matrix inv(S ω ) is a square matrix; the eigenvectors of inv(S ω ) × S c are used to maximize between-class separation and minimize within-class separation.
[0155] S ω = ∑ x∈categ N categ (x - μ categ )(x - μ categ ) T
[0156] S c = ∑ x∈categ N categ (μ categ - μ)(μ categ - μ) T
[0157] Among them, x is a sample in the category categ; μ categ is the mean vector of the category categ; T is the transpose operation of the matrix; N categ is the number of samples in the category categ; μ is the overall mean vector of all samples.
[0158] Furthermore, the viewing data collection module collects real-time viewing data of users from different viewing platforms and is configured as:
[0159] Collect data from the cinema ticketing system, viewing record data from online video platforms, and movie-related discussion data on social media;
[0160] Integrate the collected data from the cinema ticketing system, viewing record data from online video platforms, and movie-related discussion data on social media into the user's real-time movie viewing data.
[0161] It should be noted that in this application, the embodiments implemented on the side of the commodity recommendation system based on the characteristics of the movie-watching population can be referred to each other with the embodiments implemented on the side of the commodity recommendation method based on the characteristics of the movie-watching population, and this application will not elaborate on them one by one.
[0162] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A method for recommending products based on the characteristics of the movie-watching population, characterized in that, including: collecting real-time viewing data of users from different viewing platforms; extracting user features, film features, and viewing interaction features from the collected real-time viewing data of users; wherein, the viewing interaction features include viewing behavior features, rating and feedback features, preference and interest features, and social interaction features; obtaining the resources to be recommended in the product recommendation resource pool; extracting the product features of the resources to be recommended; updating the resources to be recommended in the product recommendation resource pool according to the product features of the resources to be recommended, the user features, the film features, and the viewing interaction features.
2. A product recommendation system based on the characteristics of the movie-watching population, the system implementing the method as described in claim 1, wherein, including: a viewing data collection module for collecting real-time viewing data of users from different viewing platforms; a viewing feature extraction module for extracting user features, film features, and viewing interaction features from the collected real-time viewing data of users; wherein, the viewing interaction features include viewing behavior features, rating and feedback features, preference and interest features, and social interaction features; a resource to be recommended acquisition module for obtaining the resources to be recommended in the product recommendation resource pool; a product feature extraction module for extracting the product features of the resources to be recommended; a resource to be recommended update module for updating the resources to be recommended in the product recommendation resource pool according to the product features of the resources to be recommended, the user features, the film features, and the viewing interaction features.
3. The merchandise recommendation system based on the characteristics of the movie-watching crowd according to claim 2, wherein The product recommendation system based on the characteristics of the viewing population further includes: a resource classification and label generation module for classifying the resources to be recommended in the product recommendation resource pool according to the product features of the resources to be recommended in the product recommendation resource pool in terms of product type dimension, brand dimension, price range dimension, and applicable scenario dimension, so as to generate product label information matching the resources to be recommended.
4. The merchandise recommendation system based on the characteristics of the movie-watching population according to claim 2, wherein The viewing feature extraction module extracts user features and film features from the collected real-time viewing data of users, and is configured to: use an autoencoder to encode the real-time viewing data of users to map the real-time viewing data of users into an embedding space, so as to obtain an embedding representation of the users; wherein, the number of features included in the embedding representation of the users is less than the number of features in the real-time viewing data of the users; determine the user features based on the embedding representation of the users; use a multi-modal video feature extraction framework to perform sentiment analysis and topic mining on the real-time viewing data of users to construct a film knowledge graph; analyze the film knowledge graph to obtain the film features; wherein, the film features include the creation and performance features, visual appearance features, optical flow features, and audio features of the film; the creation and performance features include creation subject features, performance subject features, and content classification features.
5. The merchandise recommendation system based on the characteristics of the movie-watching population according to claim 2, wherein The viewing feature extraction module extracts viewing interaction features from the collected real-time viewing data of users, and is configured to: use a recurrent neural network to model the viewing behavior sequence data in the real-time viewing data of users, and combine a scene perception algorithm to integrate the user's viewing device, time, and location into the viewing behavior sequence data; Determining a set of movie-watching behaviors and a correlation relationship between the movie-watching behaviors according to the movie-watching behavior sequence data; Each movie-watching behavior in the movie-watching behavior set is regarded as a node, and the edges between the nodes represent the association weights between the movie-watching behaviors, so as to construct a user movie-watching behavior graph; Analyzing the user's movie-watching behavior graph through a graph neural network to mine potential patterns and rules of the user's movie-watching behavior and obtain the movie-watching behavior characteristics; Perform text sentiment analysis on the rating feedback data in the user's real-time viewing data to extract the user's text sentiment features and the user's voice sentiment features when rating; The text emotion feature and the speech emotion feature are combined to obtain the score and feedback feature; By combining online learning algorithms with time series analysis algorithms, the user's preference features are extracted from the user's real-time viewing data, and a user's interest evolution map is constructed; wherein the interest evolution map is used to record the user's current preferences and interests, and to trace the source and development path of the interest; According to the interest evolution graph, mining the interest associations between users in different fields to obtain the preference and interest characteristics; Constructing a dynamic user movie-watching social network; wherein the user movie-watching social network is used to represent the user's social relationships and interactive behaviors; Based on the user's movie-watching social network and combined with the influence propagation model, analyze the user's influence, emotional tendency and information propagation path in the social network; The social interaction characteristics are determined based on the user's influence, emotional inclination and information dissemination path in the social network.
6. The merchandise recommendation system based on the characteristics of the movie-watching population according to claim 3, wherein The to-be-recommended resource updating module updates the to-be-recommended resources in the commodity recommendation resource pool according to the commodity features of the to-be-recommended resources, the user features, the film features, and the viewing interaction features, and is configured as follows: Using a K-means-based clustering algorithm, the missing values of the user features, the missing values of the movie features, and the missing values of the movie viewing interaction features are filled; Among them, for the categorical features in the user features, the film features, and the viewing interaction features, adaptive coding is used in combination with leave-one-out coding for processing; for the continuous features in the user features, the film features, and the viewing interaction features, a self-defined quantile Min-Max normalization is used for processing, first segmenting by quantile, and then performing Min-Max normalization in each segment; Based on the filled movie interaction features, filled user features and filled film features, the user preference type score features and the movie viewing popularity score weighted features are derived through the word frequency-inverse document frequency algorithm; wherein the movie viewing popularity score weighted features are obtained by weighted fusion of the filled film features and the user preference type score features; Perform singular value decomposition on the filled-in movie viewing interaction features, filled-in user features, and filled-in movie features to obtain singular values and singular vectors; Select the singular vector corresponding to the largest singular value of the first y1 as the eigenvector after dimensionality reduction; where y1 is a natural number, and the value of y1 is determined by the cumulative explained variance ratio; Project the filled viewing interaction features, filled user features, and filled film features onto the feature vectors after dimensionality reduction to obtain the data representation after dimensionality reduction; Based on the data representation after dimensionality reduction, construct the within-class scatter matrix and the between-class scatter matrix; wherein, the within-class scatter matrix is used to represent the scatter of samples within the same category; the between-class scatter matrix is used to represent the scatter of samples between different categories; Calculate the product of the inverse matrix of the within-class scatter matrix and the between-class scatter matrix to obtain the projection matrix; Calculate the eigenvalues and eigenvectors of the projection matrix; wherein, the eigenvalues of the projection matrix represent the discrimination ability of the eigenvectors of the projection matrix; According to the user preference type score feature and the viewing popularity score weighted feature, select the top y2 generalized eigenvectors from the eigenvectors of the projection matrix; where y2 is a natural number; Project the filled viewing interaction features, filled user features, and filled film features onto the generalized eigenvectors to obtain the projection feature representation; Evaluate and optimize the projection feature representation to obtain the optimized projection viewing features; Update the resources to be recommended in the commodity recommendation resource pool according to the commodity label information matching the resource to be recommended and the optimized projection viewing features.
7. The merchandise recommendation system based on the characteristics of the movie-watching population according to claim 6, wherein The evaluation and optimization of the projection feature representation to obtain the optimized projection viewing features is configured as: Conduct an F-test on the projection feature representation, process the features that do not conform to the normal distribution, and reconstruct the features with significant F-statistics; Use the gradient boosting tree as the base model to perform recursive feature elimination on the projection features after the F-test, eliminating the 15% features with the smallest weight coefficients in each round; Use the XGBoost model and the LightGBM model to train the projection features after recursive feature elimination, and calculate the first importance of the projection features after recursive feature elimination in the XGBoost model and the second importance in the LightGBM model respectively; Perform weighted fusion on the first importance and the second importance of the projection features after recursive feature elimination to obtain the reference importance of the projection features after recursive feature elimination; Retain the features with the top 80% of the reference importance to form a feature candidate set; According to the data distribution of the feature candidate set, adaptively select 7-fold cross-validation or 5-fold cross-validation; Use 7-fold cross-validation or 5-fold cross-validation, combined with stratified sampling, to evaluate the feature candidate set; According to the evaluation results of the feature candidate set, adopt a feature contribution degree analysis algorithm based on the Shapley value and the LIME algorithm to perform adaptive sensitivity analysis on each feature in the feature candidate set, and dynamically set the sensitivity threshold according to the importance and correlation of the features; For the features below the sensitivity threshold, perform feature pruning; For the pruned feature candidate set, adopt an adaptive feature fusion algorithm based on the attention mechanism of deep learning to fuse the features in the pruned feature candidate set to obtain the optimized projection viewing features.
8. The merchandise recommendation system based on the characteristics of the movie-watching population according to claim 6, wherein The within-class scatter matrix is S ω ; The between-class scatter matrix is S c ; The projection matrix is inv(S ω )×S c , and the calculation process of the projection matrix is as follows: Calculate S ω and its inverse matrix inv(S ω ); Multiply inv(S ω ) by S c to obtain inv(S ω )×S c ; where the inverse matrix inv(S ω ) is a square matrix; the eigenvectors of inv(S ω )×S c are used to maximize between-class separation and minimize within-class separation.
9. The product recommendation system based on the characteristics of the movie-watching population according to claim 2, wherein The viewing data collection module collects real-time viewing data of users from different viewing platforms and is configured to: Collect cinema ticket sales system data, online video platform viewing record data, and viewing-related discussion data on social media; Integrate the collected cinema ticket sales system data, online video platform viewing record data, and viewing-related discussion data on social media into the real-time viewing data of users.
Citation Information
Patent Citations
Traffic information filling method based on high-quality data acquisition
CN104217002A
Collaborative filtering film recommendation processing method and device based on knowledge graph
CN118093933A
Movie item recommendation processing method and device based on knowledge graph
CN118708799A
Security propaganda and education recommendation method and system based on demand portrait and content label
CN118797173A
Construction engineering quality data management method based on artificial intelligence
CN119377209A