Big data-based movie scoring system and method

Through big data analysis and multi-task learning, combined with focus loss function, the problem of neglecting correlations in the movie scoring system is solved, and more accurate and personalized movie scoring is achieved.

CN120354239AActive Publication Date: 2025-07-22NINGBO DAHONGYING UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510438936.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the prior art, the film rating system ignores the correlation between different dimensions of the film, resulting in the inability to accurately reflect the specific evaluations and preferences of the audience or critics on different dimensions.

Method used

Using a movie scoring system based on big data, through data collection, storage and processing layer, model construction layer, service layer and user interface layer, multiple XGBoost models are trained using naive Bayes analysis and multi-task learning to predict scores in different dimensions, and combined with focus loss functions to deal with emotional tendency imbalance problem, providing personalized scores for different groups.

Benefits of technology

It improves the accuracy and personalization of film ratings, can capture the intrinsic connections between different dimensions, and improves the prediction accuracy and generalization ability of the model in different groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354239A_ABST
    Figure CN120354239A_ABST
Patent Text Reader

Abstract

The invention discloses a film scoring system and method based on big data. The film scoring system comprises a data acquisition layer, a data storage and processing layer, a model construction layer, a service layer and a user interface layer, wherein the data acquisition layer is used for performing data acquisition; the data storage and processing layer is used for carrying out data cleaning and preprocessing on the collected data and extracting features from the preprocessed data; the model construction layer is responsible for processing and analyzing big data, generating movie scores, analyzing user behavior data by using naive Bayes, quantifying positive / negative emotions, reducing weights of easy-to-classify samples by using focus loss when emotional tendency distribution of a specific group is imbalanced, and constructing a classification model based on movie features, multi-dimensional features and time features; according to the method, multiple XGBoost models are trained through multi-task learning, scores of different dimensions are predicted respectively, input of emotional tendency is provided for different dimensions through naive Bayesian analysis results, and meanwhile personalized scores are provided for different groups in combination with group characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of movie rating, and in particular, to a movie rating system and method based on big data. Background Art

[0002] The rapid development of Internet technology has made it possible to collect, store, and analyze massive amounts of data. Big data technology can mine valuable information from these data, providing strong technical support for movie rating systems. Movie rating systems can provide a relatively objective and comprehensive movie evaluation for audiences by collecting and analyzing audience evaluation data. This helps audiences have a preliminary understanding of movies before watching, avoiding wasting time and money on movies they don't like.

[0003] In the prior art, movie rating systems only focus on overall ratings and ignore the correlations between different dimensions of movies (such as plot, actor performance, director level, etc.). That is, under the single-task learning method, the rating information of different dimensions may not be considered or utilized separately. The model may simply combine these factors to analyze an overall rating. However, this method may not accurately reflect the specific evaluations and preferences of audiences or critics in different dimensions. Therefore, a movie rating system and method based on big data are proposed. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem in the prior art that under the single-task learning method, the rating information of different dimensions may not be considered or utilized separately, and thus may not accurately reflect the specific evaluations and preferences of audiences or critics in different dimensions, and to propose a movie rating system and method based on big data.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] A movie rating system based on big data includes a data collection layer, a data storage and processing layer, a model construction layer, a service layer, and a user interface layer:

[0007] Data collection layer: Perform data collection, including user behavior data, movie metadata, social media data, and professional movie review data. The user behavior data includes viewing records, search histories, comments, likes, shares, etc. The movie metadata includes basic movie information (such as director, actor, genre, release year, etc.), plot summaries, trailer view counts, poster click-through rates, etc. The social media data includes discussions and sentiment tendencies about movies crawled from social platforms such as Twitter and Reddit. The professional movie review data includes ratings and reviews from well-known movie review websites and magazines;

[0008] Data Storage and Processing Layer: Use Hadoop HDFS to store the collected data, and perform data cleaning and preprocessing on the collected data. Extract features from the preprocessed data, including user preference features (such as comedy and action preferences), group features (such as circle influence and team activity), movie features (director influence and cast), multi-dimensional movie features (such as director style score and actor acting score), and time features (initial release popularity). The user preference features are extracted from user behavior data, the group features are extracted from user behavior data and social media data, the movie features are extracted from movie metadata, the multi-dimensional movie features are extracted from social media data and professional movie review data, and the time features are extracted from user behavior data and movie metadata;

[0009] Model Building Layer: Responsible for processing and analyzing big data to generate movie ratings. Use Naive Bayes to analyze user preference features and group features, quantify positive / negative sentiment. When the sentiment tendency distribution of a specific group is unbalanced, use focal loss to reduce the weight of easy-to-classify samples. Based on movie features, multi-dimensional features, and time features, adopt multi-task learning to train multiple XGBoost models to predict ratings in different dimensions respectively. The Naive Bayes analysis results provide the input of sentiment tendency for different dimensions. At the same time, combined with group features, provide personalized ratings for different groups;

[0010] Service Layer: Responsible for providing API interfaces and real-time processing functions to support data interaction and response between the model building layer and the user interface layer;

[0011] User Interface Layer: Responsible for displaying the data generated by the model building layer and the functions provided by the service layer.

[0012] The above technical solution further includes:

[0013] Preferably, the user behavior data reflects the user's movie-watching habits, preferences, and social behaviors, which is the key to understanding user needs and preferences. The movie metadata provides the basic information and market performance of the movie. The social media data is scraped from social platforms to reveal data on public discussions and sentiment tendencies towards movies. The professional movie review data includes the ratings and comments of professional movie critics.

[0014] Preferably, the user preference features are extracted by calculating the user's viewing frequency or rating of different types of movies. Use nodes and edges in the social network to represent the relationship between users and groups, and then extract group features by calculating the degree of nodes. The movie features are extracted from movie metadata. The multi-dimensional movie features are extracted based on professional movie reviews and user comments through sentiment analysis and topic modeling. Use time series analysis to extract time features.

[0015] Preferably, the specific steps of using Naive Bayes to analyze user comments and quantify positive / negative sentiment are as follows:

[0016] Data preparation: Collect user behavior data, including viewing records, search history, comments, likes, shares, etc. Preprocess the text data in the user behavior data, such as removing stop words, punctuation marks, performing stemming or lemmatization, etc. Convert the preprocessed text data into numerical features, usually using methods such as Bag of Words, TF-IDF (Term Frequency - Inverse Document Frequency), or word embeddings (such as Word2Vec, BERT);

[0017] Bernoulli Naive Bayes model construction: Assume that the features are independent of each other (the assumption of Naive Bayes), and each feature is binary (Bernoulli distribution), that is, it only takes values of 0 or 1. According to the preprocessed data, calculate the prior probability of each feature, that is, the probability of the feature appearing in all samples. Use Bayes' formula to calculate the posterior probability of each category, that is, the probability of belonging to a certain category given the observed data. The posterior probability P(C∣X) is expressed as where P(C∣X) is the posterior probability of category C given feature X, P(X∣C) is the likelihood probability of feature X given category C, P(C) is the prior probability of category C, P(X) is the prior probability of feature X. In Bernoulli Naive Bayes, assume that the feature vector X has n binary features, that is, X = (x1, x2,..., x n ), where x i ∈{0,1}. Given category C, each feature x i appears independently, that is, the probability that x i =1 is θ i,C . The likelihood probability P(X∣C) is expressed as where is the probability that feature x i is 1 under category C, and (1 - θ i,C ) is the probability that feature x i is 0 under category C. The decision rule of the Bernoulli Naive Bayes classifier is expressed as where is the predicted category;

[0018] Sentiment tendency quantification: Input the text data into the Bernoulli Naive Bayes model to obtain the probability that each comment belongs to positive or negative sentiment. Set a threshold according to the probability value to classify the comments as positive or negative sentiment. Aggregate the sentiment tendencies of all comments to obtain the overall sentiment tendency of users towards the movie.

[0019] Preferably, when the emotional tendency distribution of a specific group is unbalanced, the specific process of using focal loss to reduce the weight of easily classifiable samples is as follows:

[0020] Analysis of emotional tendency distribution: Analyze the emotional tendency distribution of a specific group to identify whether there is an imbalance (such as the number of positive emotional samples is much larger than that of negative emotional samples). When the emotional tendency distribution of a specific group is unbalanced, that is, the number of positive or negative comments varies greatly, use focal loss to reduce the weight of easily classifiable samples;

[0021] Definition of focal loss function: For the unbalanced distribution, define the focal loss function, and introduce a modulation factor to reduce the weight of easily classifiable samples, so that the model focuses on difficult-to-classify samples. The formula of the focal loss function is FL(p i ) = -α i (1 - p i ) γ log(p i ), where p i is the probability that the sample predicted by the model belongs to the true category (for positive or negative emotions), α i is the class weight (used to handle class imbalance), and γ is the modulation factor (used to reduce the weight of easily classifiable samples);

[0022] Model optimization: During the training process of the Bernoulli Naive Bayes model, use the focal loss function to replace the cross-entropy loss function.

[0023] Preferably, based on historical rating data, multiple XGBoost models are trained using multi-task learning to predict ratings of different dimensions respectively. The Naive Bayes analysis results provide inputs of emotional tendencies for different dimensions, including the following steps:

[0024] Data preparation: Collect the features provided by the data storage and processing layer;

[0025] Construction of multi-task learning model: Adopt a multi-task learning framework to construct an XGBoost model for each rating dimension. The objective function of the XGBoost model is expressed as where N is the number of users, M is the number of rating dimensions, y ij is the true rating of user i in dimension j, is the predicted rating of the model, w k is the model parameter, and λ is the regularization coefficient;

[0026] Naive Bayes analysis: The results of sentiment analysis are input into the multi-task learning model as additional features to help the model more accurately predict ratings;

[0027] Model Training and Evaluation: The model is trained using cross-validation, and the model performance is evaluated. Based on the evaluation results, the model parameters and feature sets are adjusted to optimize the model performance;

[0028] Prediction and Output: The trained model is used to predict the ratings of new movies, and the prediction results are presented to users in an intuitive and easy-to-understand manner.

[0029] Preferably, according to the characteristics of the group to which the user belongs, the model prediction results are adjusted to provide personalized ratings for different groups. The specific steps of using the K-means clustering algorithm to group users and then adjusting the ratings according to the group characteristics are as follows:

[0030] Data Preprocessing: The collected data is preprocessed;

[0031] Data Integration and Feature Engineering: The preprocessed data is integrated to generate a unified feature matrix;

[0032] Select the Value of K: Preset the value of K in combination with the actual situation;

[0033] Apply the K-means Algorithm for Clustering: Randomly select K initial clustering centers, assign each data point to the nearest clustering center, update the centroid of the cluster according to the mean of the data points in each cluster, and repeat the assignment and update steps until the clustering centers no longer change or reach the preset maximum number of iterations;

[0034] Analyze the Clustering Results: According to the data characteristics in each cluster, analyze the representativeness of the cluster, assign a label to each cluster, classify the clusters according to user preferences, and visualize the clustering results using a three-dimensional scatter plot;

[0035] Personalized Rating: Adjust the rating results according to the classification results. For example, for the user group who likes action movies, the weight of the plot rating can be slightly increased; for the user group who likes comedy movies, the weight of the actor performance rating can be slightly increased.

[0036] Preferably, a movie rating method based on big data corresponding to a movie rating system based on big data includes:

[0037] S1: Data Collection and Integration: Data collection is carried out, including user behavior data, movie metadata, social media data, and professional movie review data;

[0038] S2: Data Storage and Processing: The collected data is stored using Hadoop HDFS, and the collected data is cleaned and preprocessed, and features are extracted from the preprocessed data;

[0039] S3: Sentiment Analysis: Use the Naive Bayes algorithm to perform sentiment analysis on user behavior data and social media data, quantify positive / negative sentiment tendencies. To address the issue of unbalanced sentiment tendency distributions among specific groups, introduce the focal loss function to reduce the weights of easily classified samples;

[0040] S4: Multi-task Learning: Based on historical rating data, use a multi-task learning framework to train multiple XGBoost models. Each model is responsible for predicting the rating of one dimension (such as overall rating, plot rating, actor performance rating, etc.). Use the sentiment tendency obtained from the Naive Bayes analysis in S3 as a feature and input it into the XGBoost model to provide reference information on sentiment tendency for different dimensions and improve the accuracy of rating prediction;

[0041] S5: Personalized Rating: Combine the characteristics of the group to which the user belongs to make personalized adjustments to the model prediction results and provide ratings that meet the preferences of different groups.

[0042] The present invention has the following beneficial effects:

[0043] 1. In the present invention, use a multi-task learning framework to train multiple XGBoost models, and each model is responsible for predicting the rating of one dimension. This method can capture the internal connections between different dimensions, make full use of the correlations between different dimensions, and improve the accuracy of rating prediction. At the same time, use the sentiment tendency obtained from the Naive Bayes analysis as a feature and input it into the XGBoost model to provide reference information on sentiment tendency for different dimensions, further improving the accuracy and personalization of the rating.

[0044] 2. In the present invention, use the Naive Bayes algorithm to perform sentiment analysis on user behavior data and social media data, which can quantify positive / negative sentiment tendencies, provide reference for movie ratings from the sentiment dimension. To address the issue of unbalanced sentiment tendency distributions among specific groups, introduce the focal loss function to reduce the weights of easily classified samples and increase the attention of the model to difficult-to-classify samples. This helps improve the prediction accuracy and generalization ability of the model among different groups. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is the system architecture diagram of a movie rating system based on big data proposed by the present invention;

[0046] Figure 2 is the flowchart of a movie rating method based on big data proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] As Figure 1 shown, a movie rating system based on big data includes a data collection layer, a data storage and processing layer, a model construction layer, a service layer, and a user interface layer:

[0049] Data collection layer: Data collection is performed, including user behavior data, movie metadata, social media data, and professional movie review data. The user behavior data includes viewing records, search histories, comments, likes, shares, etc. The movie metadata includes basic movie information (such as director, actors, genre, release year, etc.), plot summaries, trailer view counts, poster click-through rates, etc. The social media data includes discussions and sentiment tendencies about movies scraped from social platforms such as Twitter and Reddit. The professional movie review data includes ratings and comments from well-known movie review websites and magazines;

[0050] Data storage and processing layer: Hadoop HDFS is used to store the collected data, and the collected data is cleaned and preprocessed. Features are extracted from the preprocessed data. The features include user preference features (comedy, action preferences), group features (such as circle influence, team activity), movie features (director influence, actor lineup), multi-dimensional movie features (such as director style score, actor acting score), and time features (initial release popularity). The user preference features are extracted from user behavior data, the group features are extracted from user behavior data and social media data, the movie features are extracted from movie metadata, the multi-dimensional movie features are extracted from social media data and professional movie review data, and the time features are extracted from user behavior data and movie metadata;

[0051] Model construction layer: Responsible for processing and analyzing big data to generate movie ratings. Naive Bayes is used to analyze user preference features and group features, quantify positive / negative sentiment. When the sentiment tendency distribution of a specific group is unbalanced, focal loss is used to reduce the weight of easily classified samples. Based on movie features, multi-dimensional features, and time features, multiple XGBoost models are trained using multi-task learning to predict ratings in different dimensions. The results of the naive Bayes analysis provide input on sentiment tendency for different dimensions, thus helping the system more accurately predict user ratings in each dimension. At the same time, combined with the characteristics of the group to which the user belongs, personalized ratings are provided for different groups;

[0052] Service layer: Responsible for providing API interfaces and real-time processing functions, supporting data interaction and response between the model building layer and the user interface layer;

[0053] User interface layer: Responsible for presenting the data generated by the model building layer and the functions provided by the service layer.

[0054] In one embodiment, the user behavior data reflects the user's movie-watching habits, preferences, and social behaviors, which are the key to understanding the user's needs and preferences. The movie metadata provides the basic information and market performance of the movie. The social media data is scraped from social platforms to reveal data on public discussions and emotional tendencies towards the movie. The professional movie review data includes the ratings and reviews of professional movie critics.

[0055] In one embodiment, the user preference features are extracted by calculating the user's movie-watching frequency or ratings for different types of movies. The relationships between users and groups are represented using nodes and edges in a social network, and then the group features are extracted by calculating the degree of the nodes. The movie features are extracted from the movie metadata. The multi-dimensional movie features are based on professional movie reviews and user comments, and are extracted through sentiment analysis and topic modeling. Time series analysis is used to extract time features.

[0056] In one embodiment, the specific steps for using Naive Bayes to analyze user comments and quantify positive / negative sentiment are as follows:

[0057] Data preparation: Collect user behavior data, including movie-watching records, search history, comments, likes, shares, etc. Preprocess the text data in the user behavior data, such as removing stop words, punctuation marks, performing stemming or lemmatization, etc. Convert the preprocessed text data into numerical features, usually using methods such as the Bag of Words model, TF-IDF (Term Frequency-Inverse Document Frequency), or word embeddings (such as Word2Vec, BERT);

[0058] Bernoulli Naive Bayes model construction: Assume that the features are independent of each other (the assumption of Naive Bayes), and each feature is binary (Bernoulli distribution), that is, it only takes values of 0 or 1. According to the preprocessed data, calculate the prior probability of each feature, that is, the probability of the feature appearing in all samples. Use Bayes' formula to calculate the posterior probability of each category, that is, the probability of belonging to a certain category given the observed data. The posterior probability P(C∣X) is expressed as Among them, P(C∣X) is the posterior probability of class C given feature X, P(X∣C) is the likelihood probability of feature X given class C, P(C) is the prior probability of class C, and P(X) is the prior probability of feature X. In Bernoulli Naive Bayes, it is assumed that the feature vector X has n binary features, i.e., X = (x1, x2,..., x n ), where x i ∈{0, 1}. Given class C, each feature x i appears independently, i.e., the probability that x i = 1 is θ i,C . The likelihood probability P(X∣C) is expressed as Among them, is the probability that feature x i is 1 under class C, and (1 - θ i,C ) is the probability that feature x i is 0 under class C. The decision rule of the Bernoulli Naive Bayes classifier is expressed as Among them, is the predicted class;

[0059] Sentiment tendency quantification: Input the text data into the Bernoulli Naive Bayes model to obtain the probability that each comment belongs to positive or negative sentiment. Set a threshold according to the probability value to classify the comment as positive or negative sentiment, and summarize the sentiment tendencies of all comments to obtain the overall sentiment tendency of the user towards the movie.

[0060] In one embodiment, when the sentiment tendency distribution of a specific group is unbalanced, the specific process of using focal loss to reduce the weight of easy-to-classify samples is as follows:

[0061] Sentiment tendency distribution analysis: Analyze the sentiment tendency distribution of a specific group to identify whether there is an unbalanced phenomenon (such as the number of positive sentiment samples is much larger than that of negative sentiment samples). When the sentiment tendency distribution of a specific group is unbalanced, that is, the number of positive or negative comments varies greatly, use focal loss to reduce the weight of easy-to-classify samples;

[0062] Definition of focal loss function: For the unbalanced distribution, define the focal loss function. By introducing a modulation factor to reduce the weight of easy-to-classify samples and make the model focus on difficult-to-classify samples, the formula of the focal loss function is FL(p i ) = -α i (1 - p i ) γ log(p i ), where p i is the probability that the sample predicted by the model belongs to the true class (for positive or negative sentiment), α i is the class weight (used to handle class imbalance), and γ is the modulation factor (used to reduce the weight of easy-to-classify samples);

[0063] Model optimization: During the training process of the Bernoulli Naive Bayes model, the focal loss function is used to replace the cross-entropy loss function.

[0064] Suppose in a certain group, positive sentiment comments account for 80% and negative sentiment comments account for 20%. When training the model with the traditional cross-entropy loss function, the model may over-focus on positive sentiment comments, resulting in a decline in the prediction ability for negative sentiment comments. By introducing the focal loss function and appropriately adjusting the values of α i and γ, the model can handle positive and negative sentiment comments more balancedly, thus improving the overall performance.

[0065] In one embodiment, based on historical rating data, multiple XGBoost models are trained using multi-task learning to predict ratings in different dimensions respectively. The Naive Bayes analysis results provide input for sentiment tendencies in different dimensions, including the following steps:

[0066] Data preparation: Collect the features provided by the data storage and processing layer;

[0067] Multi-task learning model construction: Using the multi-task learning framework, an XGBoost model is constructed for each rating dimension. The objective function of the XGBoost model is expressed as where N is the number of users, M is the number of rating dimensions, y ij is the true rating of user i in dimension j, is the predicted rating of the model, w k is the model parameter, and λ is the regularization coefficient;

[0068] Naive Bayes analysis: The sentiment analysis results are input into the multi-task learning model as additional features to help the model more accurately predict ratings;

[0069] Model training and evaluation: Use cross-validation to train the model and evaluate the model performance. Adjust the model parameters and feature set according to the evaluation results to optimize the model performance;

[0070] Prediction and output: Use the trained model to predict the ratings of new movies and display the prediction results to users in an intuitive and easy-to-understand way.

[0071] In one embodiment, the specific steps of adjusting the model prediction results according to the characteristics of the group to which the user belongs to provide personalized ratings for different groups are as follows: Using the K-means clustering algorithm to group users, and then adjusting the ratings according to the group characteristics:

[0072] Data preprocessing: Preprocess the collected data;

[0073] Data integration and feature engineering: Integrate the preprocessed data to generate a unified feature matrix;

[0074] Select the value of K: Preset the value of K in combination with the actual situation;

[0075] Apply the K-means algorithm for clustering: Randomly select K initial clustering centers, assign each data point to the nearest clustering center, update the centroid of the cluster according to the mean of the data points in each cluster, and repeat the assignment and update steps until the clustering centers no longer change or reach the preset maximum number of iterations;

[0076] Analyze the clustering results: Analyze the representativeness of each cluster according to the data characteristics in each cluster, assign a label to each cluster, classify the clusters according to user preferences, and visualize the clustering results using a three-dimensional scatter plot;

[0077] Personalized scoring: Adjust the scoring results according to the classification results. For example, for the user group who likes action movies, the weight of the plot score can be slightly increased; for the user group who likes comedy movies, the weight of the actor performance score can be slightly increased.

[0078] As Figure 2 shown, a movie scoring method based on big data includes:

[0079] S1: Data collection and integration: Conduct data collection, including user behavior data, movie metadata, social media data, and professional movie review data;

[0080] S2: Data storage and processing: Use Hadoop HDFS to store the collected data, perform data cleaning and preprocessing on the collected data, and extract features from the preprocessed data;

[0081] S3: Sentiment analysis: Use the Naive Bayes algorithm to perform sentiment analysis on user behavior data and social media data, quantify the positive / negative sentiment tendency. For the problem of unbalanced sentiment tendency distribution in a specific group, introduce the focal loss function to reduce the weight of easily classified samples;

[0082] S4: Multi-task learning: Based on historical scoring data, use a multi-task learning framework to train multiple XGBoost models, each model is responsible for predicting a dimension of the score (such as overall score, plot score, actor performance score, etc.), and use the sentiment tendency obtained from the Naive Bayes analysis in S3 as a feature input into the XGBoost model to provide reference information on sentiment tendency for different dimensions and improve the accuracy of score prediction;

[0083] S5: Personalized scoring: Combine the characteristics of the user group to which the user belongs to perform personalized adjustment on the model prediction results and provide scores that meet their preferences for different groups.

[0084] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A movie rating system based on big data, characterized in that, It includes a data collection layer, a data storage and processing layer, a model construction layer, a service layer, and a user interface layer: Data collection layer: Conduct data collection, including user behavior data, movie metadata, social media data, and professional movie review data; Data storage and processing layer: Use Hadoop HDFS to store the collected data, and perform data cleaning and preprocessing on the collected data. Extract features from the preprocessed data. The features include user preference features, group features, movie features, multi-dimensional movie features, and time features. The user preference features are extracted from user behavior data, the group features are extracted from user behavior data and social media data, the movie features are extracted from movie metadata, the multi-dimensional movie features are extracted from social media data and professional movie review data, and the time features are extracted from user behavior data and movie metadata; Model construction layer: Responsible for processing and analyzing big data, generating movie ratings. Use Naive Bayes to analyze user preference features and group features, quantify positive / negative sentiment. When the sentiment tendency distribution of a specific group is unbalanced, use focal loss to reduce the weight of easily classified samples. Based on movie features, multi-dimensional features, and time features, adopt multi-task learning to train multiple XGBoost models to predict ratings of different dimensions respectively. The Naive Bayes analysis results provide input of sentiment tendency for different dimensions. At the same time, combined with group features, provide personalized ratings for different groups; Service layer: Responsible for providing API interfaces and real-time processing functions, supporting data interaction and response between the model construction layer and the user interface layer; User interface layer: Responsible for displaying the data generated by the model construction layer and the functions provided by the service layer.

2. The movie rating system based on big data according to claim 1, wherein The user behavior data reflects users' movie-watching habits, preferences, and social behaviors, which is the key to understanding users' needs and preferences. The movie metadata provides basic information and market performance of movies. The social media data is data scraped from social platforms to reveal the public's discussion and sentiment tendency towards movies. The professional movie review data includes ratings and comments from professional movie critics.

3. The movie rating system based on big data according to claim 1, characterized in that, The user preference features are extracted by calculating the movie-watching frequency or ratings of users for different types of movies. Use nodes and edges in the social network to represent the relationship between users and groups, and then extract group features by calculating the degree of nodes. The movie features are extracted from movie metadata. The multi-dimensional movie features are extracted based on professional movie reviews and user comments through sentiment analysis and topic modeling. Use time series analysis to extract time features.

4. A movie rating system based on big data according to claim 1, characterized in that, The specific steps of using Naive Bayes to analyze user comments and quantify positive / negative sentiment are as follows: Data preparation: Collect user behavior data, preprocess the text data in the user behavior data, and convert the preprocessed text data into numerical features; Construction of Bernoulli Naive Bayes Model: Assume that the features are independent of each other (the assumption of Naive Bayes), and each feature is binary (Bernoulli distribution), that is, it only takes values of 0 or 1. According to the preprocessed data, calculate the prior probability of each feature, that is, the probability of the feature appearing in all samples. Use Bayes' formula to calculate the posterior probability of each class, that is, the probability of belonging to a certain class given the observed data. The posterior probability P(C∣X) is expressed as where P(C∣X) is the posterior probability of class C given feature X, P(X∣C) is the likelihood probability of feature X given class C, P(C) is the prior probability of class C, and P(X) is the prior probability of feature X. In Bernoulli Naive Bayes, assume that the feature vector X has n binary features, that is, X = (x1, x2,..., x n ), where x i ∈ {0, 1}. Given class C, each feature x i appears independently, that is, the probability that x i = 1 is θ i,C . The likelihood probability P(X∣C) is expressed as where is the probability that feature x i is 1 under class C, and (1 - θ i,C ) is the probability that feature x i is 0 under class C. The decision rule of the Bernoulli Naive Bayes classifier is expressed as where is the predicted class; Sentiment tendency quantification: Input the text data into the Bernoulli Naive Bayes model to obtain the probability that each comment belongs to positive or negative sentiment. Set a threshold according to the probability value to classify the comment as positive or negative sentiment. Aggregate the sentiment tendencies of all comments to obtain the overall sentiment tendency of users towards the movie.

5. A movie rating system based on big data according to claim 1, characterized in that, When the emotional tendency distribution of a specific group is unbalanced, the specific process of using focal loss to reduce the weight of easily classified samples is as follows: Analysis of emotional tendency distribution: Analyze the emotional tendency distribution of a specific group to identify whether there is an imbalance. When the emotional tendency distribution of a specific group is unbalanced, that is, when the number of positive or negative comments varies greatly, use focal loss to reduce the weight of easily classified samples; Definition of Focal Loss Function: For the imbalanced distribution, the focal loss function is defined. By introducing a modulation factor, the weights of easily classified samples are reduced, enabling the model to focus on difficult-to-classify samples. The formula for the focal loss function is FL(p i ) = -α i (1 - p i ) γ log(p i ), where p i is the probability that the sample predicted by the model belongs to the true class, α i is the class weight, and γ is the modulation factor; Model optimization: During the training process of the Bernoulli Naive Bayes model, use the focal loss function to replace the cross-entropy loss function.

6. The movie rating system based on big data according to claim 1, characterized in that, Based on historical rating data, use multi-task learning to train multiple XGBoost models to predict ratings in different dimensions respectively. The Naive Bayes analysis results provide the input of emotional tendency for different dimensions, including the following steps: Data preparation: Collect the features provided by the data storage and processing layer; Multi-task learning model construction: Using a multi-task learning framework, an XGBoost model is constructed for each rating dimension, and the objective function of the XGBoost model is expressed as where N is the number of users, M is the number of rating dimensions, y ij is the true rating of user i in dimension j, is the predicted rating of the model, w k are the model parameters, and λ is the regularization coefficient; Naive Bayes analysis: The emotional analysis results are input into the multi-task learning model as additional features; Model training and evaluation: Use cross-validation to train the model and evaluate the model performance. Adjust the model parameters and feature set according to the evaluation results to optimize the model performance; Prediction and output: Use the trained model to predict the ratings of new movies.

7. The movie rating system based on big data according to claim 1, characterized in that, The specific steps of adjusting the model prediction results according to the characteristics of the user's group to provide personalized ratings for different groups, using the K-means clustering algorithm to group users, and then adjusting the ratings according to the group characteristics are as follows: Data preprocessing: Preprocess the collected data; Data integration and feature engineering: Integrate the preprocessed data to generate a unified feature matrix; Select the value of K: Preset the value of K in combination with the actual situation; Apply the K-means algorithm for clustering: Randomly select K initial cluster centers, assign each data point to the nearest cluster center, update the centroid of the cluster according to the mean of the data points in each cluster, and repeat the assignment and update steps until the cluster centers no longer change or reach the preset maximum number of iterations; Analyze the clustering results: Analyze the representativeness of each cluster according to the data characteristics in each cluster, assign a label to each cluster, and classify the clusters according to the user preferences; Personalized rating: Adjust the rating results according to the classification results.

8. A method for movie rating based on big data corresponding to the movie rating system based on big data according to claim 1, characterized in that, Including: S1: Data collection and integration: Conduct data collection, including user behavior data, movie metadata, social media data, and professional movie review data; S2: Data storage and processing: Use Hadoop HDFS to store the collected data, and perform data cleaning and preprocessing on the collected data, and extract features from the preprocessed data; S3: Sentiment analysis: Use the Naive Bayes algorithm to perform sentiment analysis on user behavior data and social media data, quantify the positive / negative sentiment tendency, and introduce the focal loss function to reduce the weight of easily classified samples for the problem of unbalanced emotional tendency distribution of a specific group; S4: Multi-task learning: Based on historical rating data, use the multi-task learning framework to train multiple XGBoost models. Each model is responsible for predicting the rating of one dimension. Use the emotional tendency obtained from the Naive Bayes analysis in S3 as a feature and input it into the XGBoost model to provide reference information on emotional tendency for different dimensions; S5: Personalized scoring: Combine the characteristics of the user's group to make personalized adjustments to the model prediction results, and provide scores that meet the preferences of different groups.

Citation Information

Patent Citations

  • Method for predicting movie score categories based on movie structured information and brief introductions

    CN111104552A

  • Film review sentiment analysis method and system based on classifier and feature integration

    CN115795034A

  • Movie score prediction method and device based on multi-modal data

    CN116308567A

  • Prediction of film success-quotient

    US20200372524A1