A movie box office prediction system based on machine learning and big data
By building a movie box office prediction system based on machine learning and big data, integrating multi-source data and simulating the social media information dissemination path, the problem of insufficient accuracy and timeliness of box office prediction in the existing technology is solved, and more accurate market insights and decision-making support are achieved.
Patent Information
- Application Number
- CN202411170523.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-08-26
AI Technical Summary
The existing technology cannot effectively integrate multi-source data, especially the dynamic information dissemination path and node influence in social media, resulting in insufficient accuracy and timeliness of movie box office prediction in complex market environments.
Build a movie box office prediction system based on machine learning and big data, including data collection, processing and feature engineering, social dissemination path prediction, box office prediction, model verification and optimization and result display modules, simulate information dissemination paths and node influence through dynamic graph neural networks, and combine Bayesian optimization neural networks to make box office predictions.
It improves the accuracy and timeliness of movie box office forecasts, provides in-depth communication path analysis and growth forecasts, and helps the film publishers and marketing teams make precise market decisions.
Smart Images

Figure CN119151589B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of movie box office prediction, and particularly to a movie box office prediction system based on machine learning and big data. Background Art
[0002] With the rapid development of the movie industry, movie box office prediction plays a crucial role in movie distribution and marketing strategies. Box office prediction not only affects the choice of movie release time but also has important guiding significance for decisions such as advertising placement and resource allocation. Especially in the context of the rapid development of social media, the influence of audience feedback, word-of-mouth spread, and marketing activities on movie box office has become increasingly obvious. Therefore, how to accurately and timely predict movie box office has become the focus of the movie industry.
[0003] Currently, most traditional movie box office prediction models rely on historical box office data and basic market analysis. This method is inadequate in dealing with market changes and the influence of dynamic social media. Existing technologies often fail to fully integrate multi-source data, especially key features such as the dynamic information dissemination path and node influence in social media. This leads to poor performance of the prediction model in complex and changing market environments, making it difficult to accurately capture the emotional fluctuations of audiences and the real-time impact of information dissemination, thus affecting the accuracy of movie box office prediction and the timeliness of decision-making.
[0004] The present invention aims to solve the above deficiencies in the prior art and proposes a movie box office prediction system based on machine learning and big data, which can provide accurate market insights and decision-making support for movie distributors, optimize marketing strategies, and maximize movie box office. Summary of the Invention
[0005] Based on the above purpose, the present invention provides a movie box office prediction system based on machine learning and big data.
[0006] A movie box office prediction system based on machine learning and big data includes a data collection module, a data processing and feature engineering module, a social dissemination path prediction module, a box office prediction module, a model verification and optimization module, and a result display module, wherein;
[0007] The data collection module collects market data related to movie box office from multiple data sources, including historical box office data, social media feedback data, marketing activity data, and information on competing movies in the same period;
[0008] The data processing and feature engineering module cleans and integrates the collected market data, and constructs a feature set based on factors affecting movie box office, including movie type, the influence of directors and actors, release time, budget scale, and audience emotion data;
[0009] The social propagation path prediction module simulates the information propagation path and speed on social media, and quantifies the influential nodes in the information propagation process, specifically including:
[0010] Social network construction: Construct a social network graph related to movie topics, model users and information on social media platforms as nodes and edges, and form a dynamic social network structure;
[0011] Propagation path tracking: Use a dynamic graph neural network (D-GNN) model to real-time track the propagation path of information in the social network, and identify the starting point, influential nodes and their propagation directions of information propagation;
[0012] Propagation speed analysis: Calculate the propagation speed of information in the social network structure, analyze the time required for information to propagate from one node to another, and evaluate the propagation efficiency in combination with the breadth of the propagation path;
[0013] Node influence quantification: Based on the influential nodes in the information propagation path, quantify the influence of each node in the information diffusion process, and generate a score of node influence in combination with the connection degree and propagation ability of the node, as an input feature for box office prediction;
[0014] The box office prediction module combines the simulated information propagation path and speed, and predicts the movie box office through a movie box office prediction model;
[0015] The model verification and optimization module verifies and optimizes the movie box office prediction model through cross-validation and adjusts the model parameters;
[0016] The result display module presents the predicted movie box office results to users in the form of charts or reports, and provides box office growth predictions based on the social propagation path.
[0017] Optionally, the data collection module includes:
[0018] Historical box office data interface: Automatically obtain historical box office data in each time period through the API connection with the movie box office database, including release time, total box office and number of moviegoers;
[0019] Social media data scraping tool: Use the public API of social media platforms to real-time scrape social media feedback data related to movies, including user comments, number of likes, number of forwards and sentiment analysis results;
[0020] Marketing activity data collection: Obtain the input and effect data of movie marketing activities by docking with the data interfaces of major advertising platforms and marketing channels, including advertising click-through rate, exposure volume and interaction rate;
[0021] Same-period competing movie information collection tool: Automatically collect relevant data of movies released during the same period as the target movie, including the box office performance, marketing strategies, and audience evaluations of competing movies.
[0022] Optionally, the data processing and feature engineering module includes:
[0023] Data cleaning: Use the Z-Score algorithm to clean the collected market data, identify and process abnormal data, missing data, and duplicate data;
[0024] Data integration: Through multi-dimensional data alignment technology, align and integrate market data from different sources in terms of time and space;
[0025] Feature construction: Based on the factors affecting movie box office, construct a feature set, specifically including:
[0026] Movie type feature: Use One-Hot encoding to convert the movie type into multiple binary features to capture the impact of different types of movies on the box office;
[0027] Director and actor influence feature: Based on historical box office data, calculate the influence index I of the director or actor, expressed as:
[0028]
[0029] Where R i represents the box office revenue of the i-th movie, B i represents the budget of the i-th movie, and n is the total number of movies participated by the director or actor;
[0030] Release time feature: Convert the release date into holiday, weekend, or weekday features, and identify the influence trend of the release time on the box office through time series analysis;
[0031] Budget scale feature: Through normalization, convert the budget scale of the movie into a comparable feature value for analyzing the box office performance of movies with different budget scales;
[0032] Audience sentiment feature: Through natural language processing (NLP) technology, perform sentiment analysis on social media feedback data, extract the sentiment tendency score (such as positive, negative, neutral) as the sentiment feature, and calculate the audience sentiment index E by combining sentiment intensity weighting, expressed as:
[0033]
[0034] Where S j is the sentiment score of the j-th feedback, W j is the sentiment intensity weight, and m' is the total number of feedbacks.
[0035] Optionally, the social network construction includes:
[0036] Node definition and identification: Regarding users on the social media platform and movie-related content (such as posts, comments, forwards, etc.) as nodes respectively, where the user node represents a social media user and the information node represents movie-related content;
[0037] Edge definition and generation: Based on the interaction relationships between users (such as likes, comments, forwards) and the associations between users and information nodes (such as publishing, commenting, sharing), generate edges between the nodes to form a preliminary social network graph. The weight w of the edge uv is expressed as:
[0038] w uv = α·F uv + β·I ui + γ·S ii′ ;
[0039] where F uv represents the interaction frequency between user u and user v, I ui represents the association strength between user u and information node i, S ii′ represents the similarity between information node i and information node i′, and α, β, γ are the corresponding weight coefficients;
[0040] Dynamic social network update: Through the dynamic social network evolution model, update the social network graph in real time, capture the changes in the interaction relationships between users over time, and dynamically adjust the weights of the nodes and edges to reflect the real-time social dissemination state, expressed as:
[0041] A t+1 = A t + ΔA t ;
[0042] where A t represents the adjacency matrix at time t, and ΔA t represents the change in the network structure (added or deleted edges) from time t to t + 1;
[0043]
[0044] where represents the probability of generating an edge between nodes u and v at time t + 1, σ is the activation function, is the eigenvector of node u at time t, represents the eigenvector of node v at time t, and A t [u, v] is the element in the adjacency matrix at time t;
[0045] Network Structure Optimization: When constructing a dynamic social network graph, the Louvain algorithm is used to optimize the social network graph, identify the community structure in the network, and improve the resolution and stability of the social network graph by optimizing the distribution of nodes and edges, expressed as:
[0046]
[0047] Among them, Q represents modularity, m is the number of edges in the network, k u and k v are the degrees of nodes u and v respectively, A[u, v] is the value in the adjacency matrix, and δ(c u , c v ) is the indicator function.
[0048] Optionally, the propagation path tracking includes:
[0049] Model Construction: Construct a dynamic graph neural network model to model the information propagation path that changes over time in the social network. Nodes represent users and information, and edges represent the interaction between users or the association between users and information, expressed as:
[0050]
[0051] Among them, represents the hidden state of node u at time t + 1, is the hidden state of node u at time t, is the set of neighbor nodes of node u, and A t [u, v] is the element in the adjacency matrix at time t. W1, W2, and W3 are weight matrices, and σ is the activation function;
[0052] Information Propagation Path Tracking: By dynamically updating the hidden state of nodes, the dynamic graph neural network model can track the information propagation path in the social network in real time, expressed as:
[0053]
[0054] Among them, Active(u) represents the activation state of node u, and θ is the threshold;
[0055]
[0056] Among them, represents the propagation path at time t;
[0057] Identification of the propagation starting point and influential nodes: The dynamic graph neural network model identifies the starting point of information propagation by analyzing the hidden state of the initial node (information publisher). At the same time, through the change amplitude of the node hidden state and the cumulative influence in the propagation path, it quantifies and identifies the influential nodes that contribute to information diffusion during the propagation process, expressed as:
[0058]
[0059] where \(u_0\) is the initial node, \(t\) u is the activation time of node \(u\), is the set of all nodes;
[0060]
[0061] where \(I\) u represents the cumulative influence of node \(u\), \(A\) t [u, v]′ represents the connection state between node \(u\) and node \(v\), and Active(v) represents the activation state of node \(v\);
[0062] Identification of the propagation direction: By analyzing the change trend of the hidden state between nodes, the dynamic graph neural network model infers the propagation direction of information in the social network, expressed as:
[0063]
[0064] where \(D\) uv represents the propagation direction of information from node \(u\) to node \(v\), sign is the sign function, is the hidden state of node \(v\) at time \(t + 1\);
[0065]
[0066] where, represents the overall propagation direction at time \(t\), and \(\epsilon\) is the set of edges.
[0067] Optionally, the propagation speed analysis includes:
[0068] Propagation time calculation: By recording the time \(\Delta t\) required for information to propagate from node \(u\) to node \(v\) uv , and combining with the activation time of the node, calculate the propagation speed \(v\) of information from one node to the next node in the social network uv , expressed as:
[0069]
[0070] where \(d\) uv is the graph distance between node \(u\) and node \(v\);
[0071] Analysis of the breadth of the propagation path: Analyze the breadth B(t) of the propagation path of information in the network, expressed as:
[0072]
[0073] where Active(v) is the activation state of node v at time t;
[0074] Propagation efficiency evaluation: Combine the propagation speed and breadth to evaluate the propagation efficiency of information in the entire social network, expressed as:
[0075]
[0076] where, represents the propagation path at time t, represents the set of propagation paths the number of node pairs in, v uv is the propagation speed of each pair of nodes in the path, and B(t) is the breadth of the propagation path.
[0077] Optionally, the quantification of the node influence power includes:
[0078] Node connectivity calculation: The connectivity k of node u u represents the number of connections of node u with other nodes in the social network, expressed as:
[0079]
[0080] where A[u,v] is the element in the adjacency matrix, is the set of all nodes;
[0081] Node propagation ability calculation: The propagation ability C of node u u is quantified by analyzing the position and influence of the node in the propagation path, expressed as:
[0082]
[0083] where Active(v) represents the activation state of node v;
[0084] Node influence score calculation: The influence score G of node u u is quantified by combining its connectivity and propagation ability, expressed as:
[0085] G u = α·k u + β·C u ;
[0086] where α and β are weight coefficients;
[0087] Application of Node Influence in Box Office Prediction: The generated node influence score G u As a feature of the social propagation path, it is used as an input for box office prediction.
[0088] Optionally, the movie box office prediction model uses a Bayesian optimization neural network, and the Bayesian optimization neural network includes:
[0089] Feature Preprocessing and Fusion: Normalize the social propagation path, speed, and node influence features, and fuse them to generate a feature vector for the input of the neural network, expressed as:
[0090] x fused = α·x path + β·x speed + γ·xi nfluence ;
[0091] where x fused is the feature vector for the input of the fused neural network, x path , x speed , x influence are the social propagation path feature, propagation speed feature, and node influence feature respectively, and α, β, γ are fusion weights;
[0092] Network Structure Design: Introduce a processing mechanism for social network features in the neural network structure, expressed as:
[0093] Input Layer: h (0) = x fused ;
[0094] Branch Hidden Layer:
[0095]
[0096] where h (0) is the fused feature vector of the input layer, are the weight matrices of the social propagation path feature, propagation speed feature, and node influence feature respectively, are the bias terms of the social propagation path feature, propagation speed feature, and node influence feature respectively, are the outputs of the social propagation path feature, propagation speed feature, and node influence feature respectively, σ is the activation function, is the processed fused feature vector;
[0097] Box Office Prediction Output: Input the fused features into the output layer to generate the final movie box office prediction result, expressed as:
[0098]
[0099] where, is the predicted movie box office value, and are the weight matrix and bias term of the output layer, respectively;
[0100] Feature importance analysis and feedback mechanism: Feature importance analysis and feedback mechanism are introduced to enhance the interpretability of the model and make feedback adjustments according to the prediction error, expressed as:
[0101]
[0102] where I i represents the influence degree of the i-th fused feature on the prediction result, is the i-th fused feature, Δw i is the adjustment amount of the i-th feature weight, is the loss function, y is the true box office value, and η is the learning rate.
[0103] Optionally, the model verification and optimization module includes:
[0104] Cross-validation process: The model verification and optimization module verifies the movie box office prediction model through k-fold cross-validation, expressed as:
[0105]
[0106] where CV k is the average error of k-fold cross-validation, is the prediction error in the i-th fold, is the predicted value on the validation set of the k-th fold, y (i) is the true value of the i-th fold;
[0107] Model parameter optimization: After completing cross-validation, the model parameters (such as learning rate, number of hidden layers, number of neurons, etc.) are adjusted by minimizing the average error CV k of cross-validation. The model parameter optimization adopts the Bayesian optimization algorithm, expressed as:
[0108] min θ CV k (θ)+λ·R(θ);
[0109] where θ represents the set of hyperparameters of the model, CV k (θ) is the average cross-validation error under the corresponding parameters, R(θ) is the regularization term of the hyperparameters, and λ is the regularization coefficient;
[0110] Adaptive adjustment: During the model parameter optimization process, the model parameters are adjusted using the gradient descent algorithm, expressed as:
[0111]
[0112] Among them, θ t+1 represents the model parameters after the next iteration, η is the learning rate, is the gradient of the cross-validation error under the current parameter θ t of.
[0113] Optionally, the result display module includes:
[0114] Chart generation: presenting the movie box office results predicted by the movie box office prediction model to the user in the form of charts, including line charts, bar charts, and pie charts;
[0115] Report generation: automatically generating a comprehensive report based on the output data of the movie box office prediction model, including the predicted movie box office results, the performance metrics of the model (such as mean squared error, R 2 value), and relevant analysis and explanations;
[0116] Box office growth prediction based on the social dissemination path: predicting the dynamic changes of the movie box office in combination with the social dissemination path, and providing the box office growth trend within a predetermined future time, expressed as:
[0117]
[0118] Among them, represents the box office growth amount obtained based on the social dissemination path analysis, is the predicted box office value at the current time t, and GrowthPrediction(t) is the box office growth prediction value considering the social dissemination path.
[0119] Advantages of the present invention:
[0120] In the present invention, the market data from multiple data sources, including historical box office data, social media feedback, marketing campaign data, and information on competing movies in the same period, are efficiently and automatically integrated through the data collection module. Through the data processing and feature engineering module, the collected data are comprehensively cleaned, integrated, and feature constructed, ensuring the high quality and multi-dimensionality of the input data, providing a reliable basis for subsequent box office prediction. This modular processing method significantly improves the system's perception ability of market changes and prediction accuracy, enabling movie distributors to obtain more comprehensive market insights.
[0121] In the present invention, by constructing a dynamic social network graph, a dynamic graph neural network model is used to accurately simulate the information dissemination path, speed, and node influence on social media. By quantifying key nodes and path breadth, the information dissemination efficiency is comprehensively evaluated. Combining these dynamic features, the box office prediction module can more accurately reflect the real-time impact of social media on movie box office, providing in-depth dissemination path analysis and box office growth prediction, making the prediction results more accurate and timely, and assisting movie distributors and marketing teams in making precise market decisions.
[0122] In the present invention, through the model verification and optimization module, cross-validation combined with Bayesian optimization and gradient descent method is adopted to dynamically adjust the model parameters to minimize the prediction error, enhancing the generalization ability and robustness of the model. The result display module presents the movie box office prediction results and the growth prediction based on the social dissemination path clearly to users by generating intuitive charts and comprehensive reports, providing comprehensive data support and decision-making basis, enabling movie distributors to better grasp the market trend and optimize marketing strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0123] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0124] Figure 1 It is a schematic diagram of the system function modules of the embodiments of the present invention;
[0125] Figure 2 It is a schematic diagram of the social dissemination path prediction module of the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0126] The present invention will be described in detail below in combination with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.
[0127] It should be noted that in the specification, the mention of "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicates that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when combining embodiments to describe specific features, structures, or characteristics, implementing such features, structures, or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.
[0128] Generally, terms can be understood at least in part from their use in context. For example, at least in part depending on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather can alternatively, at least in part depending on the context, allow for the existence of other factors that may not be explicitly described.
[0129] As Figure 1 - Figure 2 shown, a movie box office prediction system based on machine learning and big data includes a data collection module, a data processing and feature engineering module, a social dissemination path prediction module, a box office prediction module, a model verification and optimization module, and a result display module, where;
[0130] The data collection module collects market data related to movie box office from multiple data sources, including historical box office data, social media feedback data, marketing campaign data, and information on competing movies during the same period;
[0131] The data processing and feature engineering module cleans and integrates the collected market data, and constructs a feature set based on factors affecting movie box office, including movie type, the influence of directors and actors, release time, budget scale, and audience sentiment data;
[0132] The social dissemination path prediction module simulates the information dissemination path and speed on social media, and quantifies the influence nodes in the information dissemination process, specifically including:
[0133] Social network construction: Construct a social network graph related to the movie topic, model users and information on social media platforms as nodes and edges, and form a dynamic social network structure;
[0134] Dissemination path tracking: Use a dynamic graph neural network (D-GNN) model to real-time track the dissemination path of information in the social network, and identify the starting point of information dissemination, influence nodes, and their dissemination directions;
[0135] Propagation speed analysis: Calculate the speed at which information spreads in the social network structure, analyze the time required for information to spread from one node to another, and evaluate the propagation efficiency in combination with the breadth of the propagation path;
[0136] Node influence quantification: Based on the influential nodes in the information propagation path, quantify the influence of each node in the information diffusion process, and combine the connection degree and propagation ability of the nodes to generate a score for node influence, which is used as an input feature for box office prediction;
[0137] The box office prediction module combines the simulated information propagation path and speed, and predicts the movie box office through the movie box office prediction model;
[0138] The model verification and optimization module verifies and optimizes the movie box office prediction model through cross-validation, and adjusts the model parameters to improve the accuracy and robustness of the prediction;
[0139] The result display module presents the predicted movie box office results to the user in the form of charts or reports, and provides box office growth predictions based on the social propagation path, facilitating decision-making support for movie distributors and marketing teams;
[0140] Through the above content, not only can the dynamic impact of social media information dissemination be analyzed in real time, but also the accuracy and adaptability of box office prediction can be significantly improved through multi-dimensional data integration and machine learning algorithms, providing more accurate decision-making support for movie distributors and marketing teams.
[0141] The data collection module includes:
[0142] Historical box office data interface: Automatically obtain historical box office data for each time period by connecting to the API of the movie box office database, including release time, total box office, and number of moviegoers;
[0143] Social media data scraping tool: Use the public API of social media platforms to scrape social media feedback data related to movies in real time, including user comments, likes, shares, and sentiment analysis results;
[0144] Marketing activity data collection: Obtain the input and effect data of movie marketing activities by docking with the data interfaces of major advertising platforms and marketing channels, including advertising click-through rate, exposure volume, and interaction rate;
[0145] Tool for collecting information on competing movies in the same period: Automatically collect relevant data on movies released in the same period as the target movie, including the box office performance, marketing strategies, and audience evaluations of competing movies;
[0146] Through the above, multi-source data can be efficiently and automatically collected, providing accurate and multi-dimensional inputs for subsequent box office prediction, significantly enhancing the system's perception of market changes and prediction accuracy, and providing a solid data foundation for decision support.
[0147] The data processing and feature engineering module includes:
[0148] Data cleaning: The Z-Score algorithm is used to clean the collected market data, identify and process abnormal data, missing data, and duplicate data to ensure data quality and consistency;
[0149] Data integration: Through multi-dimensional data alignment technology, market data from different sources are aligned and integrated in terms of time and space to ensure the consistency and integrity of the dataset;
[0150] Feature construction: Based on the factors affecting movie box office, a feature set is constructed, specifically including:
[0151] Movie type feature: One-Hot encoding is used to convert the movie type into multiple binary features to capture the impact of different types of movies on the box office;
[0152] Director and actor influence feature: Based on historical box office data, the influence index I of the director or actor is calculated, expressed as:
[0153]
[0154] where R i represents the box office revenue of the i-th movie, B i represents the budget of the i-th movie, and n is the total number of movies the director or actor participated in;
[0155] Release time feature: The release date is converted into holiday, weekend, or weekday features, and the influence trend of the release time on the box office is identified through time series analysis, expressed as:
[0156]
[0157] where X t is the box office data at time t, c is a constant, and θ j are the autoregressive and moving average coefficients respectively, ∈ t is the error term, and p and q are the orders of autoregression and moving average respectively;
[0158] Budget scale feature: Through normalization, the budget scale of the movie is converted into a comparable feature value to analyze the box office performance of movies with different budget scales;
[0159] Audience Emotional Characteristics: Through natural language processing (NLP) technology, perform sentiment analysis on social media feedback data, extract sentiment tendency scores (such as positive, negative, neutral) as emotional characteristics, and combine with sentiment intensity weighting to calculate the audience emotional index E, expressed as:
[0160]
[0161] Among them, S j is the sentiment score of the j-th feedback, W j is the sentiment intensity weight, and m′ is the total number of feedbacks;
[0162] Through the above content, the high quality and consistency of the data are ensured, and the key factors affecting the box office are captured, such as the influence of the director and actors, the seasonal effect of the release time, and the audience's emotional response. The finely processed features provide rich and efficient inputs for the box office prediction model, greatly enhancing the prediction ability and adaptability of the model.
[0163] Social network construction includes:
[0164] Node definition and identification: Users on the social media platform and movie-related content (such as posts, comments, forwards, etc.) are respectively used as nodes, where the user node represents the social media user, and the information node represents the movie-related content;
[0165] Edge definition and generation: According to the interaction relationships between users (such as likes, comments, forwards) and the associations between users and information nodes (such as posts, comments, shares), edges are generated between nodes to form a preliminary social network graph. The weight w uv is expressed as:
[0166] w uv =α·F uv +β·I ui +γ·S ii ′;
[0167] Among them, F uv represents the interaction frequency between user u and user v, I ui represents the association intensity between user u and information node i, S ii′ represents the similarity between information node i and information node i′, and α, β, γ are the corresponding weight coefficients;
[0168] Dynamic social network update: Through the dynamic social network evolution model, the social network graph is updated in real time to capture the changes in the interaction relationships between users over time, and the weights of nodes and edges are dynamically adjusted to reflect the real-time social communication state, expressed as:
[0169] A t+1 =A t+ΔA t ;
[0170] where A t represents the adjacency matrix at time t, and ΔA t represents the change in the network structure (added or deleted edges) from time t to t + 1;
[0171]
[0172] where represents the probability of generating an edge between nodes u and v at time t + 1, σ is the activation function, is the feature vector of node u at time t, represents the feature vector of node v at time t, and A t [u, v] is the element in the adjacency matrix at time t, indicating whether there is a direct connection between nodes u and v at time t;
[0173] Network structure optimization: When constructing a dynamic social network graph, the Louvain algorithm is used to optimize the social network graph, identify the community structure in the network, and improve the resolution and stability of the social network graph by optimizing the distribution of nodes and edges, expressed as:
[0174]
[0175] where Q represents modularity, m is the number of edges in the network, k u and k v are the degrees of nodes u and v respectively, A[u, v] is the value in the adjacency matrix, and δ(c u , c v ) is the indicator function, which takes 1 when u and v belong to the same community and 0 otherwise;
[0176] Through the above content, not only the accuracy and adaptability of the network structure are improved, but also a solid data foundation is provided for subsequent propagation path analysis and box office prediction, ensuring the accuracy of the prediction model in a dynamic environment.
[0177] Propagation path tracking includes:
[0178] Model construction: Construct a dynamic graph neural network model to model the information propagation path that changes over time in the social network. Nodes represent users and information, and edges represent the interaction between users or the association between users and information, expressed as:
[0179]
[0180] where represents the hidden state of node u at time t + 1, is the hidden state of node u at time t, is the set of neighbor nodes of node u, A t [u, v] is an element in the adjacency matrix at time t, W1, W2, and W3 are weight matrices, and σ is an activation function;
[0181] Information propagation path tracking: By dynamically updating the hidden state of nodes, the dynamic graph neural network model can track the information propagation path in the social network in real time. The propagation path is characterized by the activation state of nodes. When the hidden state of a node exceeds the threshold, it means the information has reached that node, which is recorded as part of the propagation path, expressed as:
[0182]
[0183] where Active(u) represents the activation state of node u. When its value is 1, it means the information has propagated to node u, and θ is the threshold;
[0184]
[0185] where, represents the propagation path at time t, which is composed of all nodes with an activation state of 1;
[0186] Propagation starting point and influential node identification: The dynamic graph neural network model identifies the starting point of information propagation by analyzing the hidden state of the initial node (information publisher). At the same time, through the change amplitude of the node hidden state and the cumulative influence in the propagation path, it quantifies and identifies the influential nodes that contribute to information diffusion during the propagation process. These nodes play a key role in the propagation path and have an important impact on the subsequent propagation direction and breadth, expressed as:
[0187]
[0188] where u0 is the initial node, t u is the activation time of node u, is the set of all nodes;
[0189]
[0190] where, I u represents the cumulative influence of node u, A t [u, v]′ represents the connection state between node u and node v, and Active(v) represents the activation state of node v;
[0191] Propagation Direction Identification: By analyzing the changing trend of the hidden states between nodes, the dynamic graph neural network model is used to infer the propagation direction of information in the social network. Information propagates from nodes with higher activation states to nodes with lower activation states. The model determines the main direction of information propagation based on the connection weights and changes in hidden states between nodes, expressed as:
[0192]
[0193] where D uv represents the propagation direction of information from node u to node v. sign is the sign function, taking 1 when x > 0, indicating that information propagates from u to v, and taking -1 when x < 0, indicating that information propagates from v to u. is the hidden state of node v at time t + 1;
[0194]
[0195] where represents the overall propagation direction at time t, and ε is the set of edges;
[0196] Through the above content, it is possible to accurately track the diffusion process of information in the social network in real time, identify the starting point of information propagation, key influential nodes, and propagation direction, providing real-time propagation path data for the box office prediction model and improving the accuracy and timeliness of prediction.
[0197] The setting of the threshold θ specifically includes:
[0198] Initial Setting: The initial value of θ can be set based on the historical distribution of the node hidden state and is usually set as the mean of the hidden state plus a multiple of a certain standard deviation, expressed as:
[0199] θ0 = μ h + k·σ h ;
[0200] where μ h is the mean of the hidden state σ h is the standard deviation of the hidden state and k is a constant coefficient used to adjust the sensitivity of the initial threshold;
[0201] Adaptive Adjustment: Over time, the system dynamically adjusts the threshold θ according to the activation situation of nodes and makes adjustments when observing too many or too few activated nodes, expressed as:
[0202] θ t+1 = θ t ± Δθ;
[0203] Among them, Δθ is the adjustment amount, which is determined based on the proportion of activated nodes. Δθ increases when the activation proportion is higher than the preset range, and Δθ decreases when the activation proportion is lower than the preset range;
[0204] Stable convergence: Through iterative adjustment, θ will eventually converge to a stable value, so that the number of node activations is maintained within a reasonable range, ensuring the effective identification of the information propagation path.
[0205] The analysis of the propagation speed includes:
[0206] Calculation of the propagation time: By recording the time Δt required for information to propagate from node u to node v uv and combining with the activation time of the nodes, calculate the propagation speed v of the information from one node to the next node in the social network uv , which is expressed as:
[0207]
[0208] where d uv is the graph distance between node u and node v;
[0209] Analysis of the breadth of the propagation path: Analyze the breadth B(t) of the information propagation path in the network. The breadth is defined as the number of nodes that the information can reach within a given time t, which is expressed as:
[0210]
[0211] where Active(v) is the activation state of node v at time t;
[0212] Evaluation of the propagation efficiency: Combine the propagation speed and breadth to evaluate the propagation efficiency of the information in the entire social network, which is expressed as:
[0213]
[0214] where represents the propagation path at time t, represents the set of propagation paths the number of node pairs in, v uv is the propagation speed of each pair of nodes in the path, and B(t) is the breadth of the propagation path;
[0215] Through the above content, the propagation speed of the information in the social network can be accurately calculated. Combining with the breadth of the propagation path, the efficiency of the information propagation can be comprehensively evaluated, providing more in-depth dynamic propagation characteristics for box office prediction, so that the prediction model can more accurately reflect the impact of social media on the movie box office.
[0216] The quantification of node influence power includes:
[0217] Node connectivity calculation: The connectivity k of node u u represents the number of connections of node u with other nodes in the social network, expressed as:
[0218]
[0219] where A[u, v] is an element in the adjacency matrix, representing the connection status between node u and node v, is the set of all nodes;
[0220] Node propagation ability calculation: The propagation ability C of node u u is quantified by analyzing the position and influence of the node in the propagation path, expressed as:
[0221]
[0222] where Active(v) represents the activation status of node v;
[0223] Node influence score calculation: The influence score G of node u u is quantified by combining its connectivity and propagation ability, expressed as:
[0224] G u = α·k u + β·C u ;
[0225] where α and β are weight coefficients, used to balance the contributions of the node's connectivity and propagation ability in the influence score;
[0226] Application of node influence in box office prediction: The generated node influence score G u is used as a feature of the social propagation path and as an input for box office prediction, to enhance the model's understanding of the dynamics of information diffusion, thereby improving the accuracy of box office prediction;
[0227] Through the above content, it is possible to effectively quantify the influence of each node in the social network during the information diffusion process, and use this dynamic feature as an important input for the box office prediction model, combining the structural attribute (connectivity) and behavioral attribute (propagation ability) of the node, thereby providing more comprehensive social network data support for box office prediction.
[0228] The movie box office prediction model uses a Bayesian optimization neural network, and the Bayesian optimization neural network includes:
[0229] Feature preprocessing and fusion: Normalize the social propagation path, speed, and node influence features, and fuse them to generate a feature vector for input to the neural network, expressed as:
[0230]
[0231] where x i is the i-th feature, and min(x) and max(x) are the minimum and maximum values of the feature respectively;
[0232] x fused = α·x path + β·x speed + γ·xi nfluence ;
[0233] where x fused is the feature vector input to the fused neural network, and x path , x speed , x influence are the social propagation path feature, propagation speed feature, and node influence feature respectively, and α, β, and γ are the fusion weights;
[0234] Network structure design: Introduce a processing mechanism for social network features in the neural network structure to better capture the impact of complex social dynamics on box office, expressed as:
[0235] Input layer: h (0) = x fused ;
[0236] Branch hidden layer:
[0237]
[0238] where h (0) is the fused feature vector of the input layer, are the weight matrices of the social propagation path feature, propagation speed feature, and node influence feature respectively, are the bias terms of the social propagation path feature, propagation speed feature, and node influence feature respectively, are the outputs of the social propagation path feature, propagation speed feature, and node influence feature respectively, σ is the activation function, is the processed fused feature vector;
[0239] Box office prediction output: Input the fused features into the output layer to generate the final movie box office prediction result, expressed as:
[0240]
[0241] where, is the predicted movie box office value, and are the weight matrix and bias term of the output layer respectively;
[0242] Feature Importance Analysis and Feedback Mechanism: Introduce feature importance analysis and feedback mechanism to enhance the interpretability of the model and perform feedback adjustment based on prediction errors, expressed as:
[0243]
[0244] where I i represents the influence degree of the i-th fused feature on the prediction result, is the i-th fused feature, and Δw i is the adjustment amount of the i-th feature weight, is the loss function, y is the true box office value, and η is the learning rate;
[0245] Through the above content, the automatic tuning of model parameters is ensured, the generalization ability and robustness of the overall model are improved, enabling the model to perform excellently in the face of multi-dimensional features and dynamic data, providing strong support for the accurate prediction of movie box office.
[0246] The model verification and optimization module includes:
[0247] Cross-validation process: The model verification and optimization module verifies the movie box office prediction model through k-fold cross-validation. During the cross-validation process, the dataset is divided into k equal subsets. The model is trained on k - 1 subsets and verified on the remaining one subset, alternating k times, with a different subset used as the validation set each time, expressed as:
[0248]
[0249] where CV k is the average error of k-fold cross-validation, is the prediction error in the i-th fold, is the predicted value on the i-th fold validation set, y (i) is the true value in the i-th fold;
[0250] Model parameter optimization: After completing cross-validation, the model parameters (such as learning rate, number of hidden layers, number of neurons, etc.) are adjusted by minimizing the average error CV k of cross-validation. The model parameter optimization adopts the Bayesian optimization algorithm, expressed as:
[0251] min θ CV k (θ)+λ·R(θ);
[0252] where θ represents the set of hyperparameters of the model, CV k (θ) is the average cross-validation error under the corresponding parameters, R(θ) is the regularization term of the hyperparameters, and λ is the regularization coefficient;
[0253] Adaptive adjustment: In the process of model parameter optimization, the model parameters are adjusted based on the gradient descent algorithm to minimize the cross-validation error so that the finally selected parameters can obtain the best prediction performance on different data sets, which is expressed as:
[0254]
[0255] Among them, θ t+1 represents the model parameters after the next iteration, η is the learning rate, is the current parameter θ t The gradient of the cross-validation error under ;
[0256] Through the above content, the generalization ability of the movie box office prediction model can be effectively evaluated, and the model parameters can be automatically adjusted through Bayesian optimization combined with the gradient descent method, so as to maintain high-precision prediction results in different data scenarios.
[0257] The result display module includes:
[0258] Chart generation: The movie box office results predicted by the movie box office prediction model are presented to users in the form of charts, including line charts, bar charts, and pie charts;
[0259] Report generation: Automatically generate a comprehensive report based on the output data of the movie box office prediction model, including the predicted movie box office results, model performance indicators (such as mean square error, R 2 value) and related analysis description;
[0260] Box office growth prediction based on social communication paths: Combine social communication paths to predict the dynamic changes in movie box office and provide box office growth trends within a predetermined period of time in the future, expressed as:
[0261]
[0262] in, represents the box office growth based on social communication path analysis, is the predicted box office value at the current time t, and GrowthPrediction(t) is the predicted box office growth value after considering the social communication path;
[0263] Through the above content, the movie box office forecast results can be presented to users in the form of intuitive charts or detailed reports, and box office growth forecasts based on social communication paths can be provided, so that film distributors and marketing teams can have a more comprehensive understanding of the film’s market performance and future trends, providing data support for decision-making.
[0264] The present invention encompasses any alternatives, modifications, equivalent methods, and solutions that are within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are set forth in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention even without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits, etc. are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0265] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. A movie box office prediction system based on machine learning and big data, characterized in that, It includes a data collection module, a data processing and feature engineering module, a social dissemination path prediction module, a box office prediction module, a model verification and optimization module, and a result display module, where: The data collection module collects market data related to movie box office from multiple data sources, including historical box office data, social media feedback data, marketing campaign data, and information on competing movies in the same period; The data processing and feature engineering module cleans and integrates the collected market data, and constructs a feature set based on factors affecting movie box office, including movie type, influence of the director and actors, release time, budget scale, and audience sentiment data; The social dissemination path prediction module simulates the information dissemination path and speed on social media, and quantifies the influence nodes in the information dissemination process, specifically including: Social network construction: Construct a social network graph for movie-related topics, model users and information on the social media platform as nodes and edges, and form a dynamic social network structure; Dissemination path tracking: Use a dynamic graph neural network model to real-time track the dissemination path of information in the social network, and identify the starting point of information dissemination, influence nodes, and their dissemination directions; The dissemination path tracking includes: Model construction: Construct a dynamic graph neural network model to model the time-varying information dissemination path in the social network, where nodes represent users and information, and edges represent interactions between users or associations between users and information, expressed as: ; Among them, represents the hidden state of the node at time , is the hidden state of the node at time , is the set of neighbor nodes of the node , , and are weight matrices, is an activation function, is the element of the adjacency matrix at time ; Information dissemination path tracking: The dynamic graph neural network model real-time tracks the dissemination path of information in the social network by dynamically updating the hidden state of nodes, expressed as: ; Among them, represents the activation state of the node , and is the threshold value; ; Among them, represents the propagation path at the moment . Identification of the starting point and influence nodes of dissemination: The dynamic graph neural network model identifies the starting point of information dissemination by analyzing the hidden state of the initial node. At the same time, through the change amplitude of the node hidden state and the cumulative influence in the dissemination path, it quantifies and identifies the influence nodes that contribute to information diffusion during the dissemination process, expressed as: ; Among them, is the initial node, is the node activation time, is the set of all nodes; ; Among them, represents the cumulative influence of the node . represents the connection state between node and node represents the activation state of the node . Identification of the dissemination direction: By analyzing the change trend of the hidden state between nodes, use the dynamic graph neural network model to infer the dissemination direction of information in the social network, expressed as: ; Among them, indicates the propagation direction of information from node to node . is the sign function, and is the hidden state of node at time ; Among them, represents the overall propagation direction at a moment, is a set of edges; Analysis of the dissemination speed: Calculate the speed of information dissemination in the social network structure, analyze the time required for information to spread from one node to another, and evaluate the dissemination efficiency in combination with the breadth of the dissemination path; The analysis of the dissemination speed includes: Propagation time calculation: By recording the time required for information to propagate from node to node , and combining with the activation time of the node, calculate the propagation speed of information from one node to the next in the social network , expressed as: ; Among them, is the graph distance between node and node; Analysis of the breadth of the propagation path: the breadth of the propagation path of information in the network is analyzed and expressed as: ; Evaluation of the dissemination efficiency: Combine the dissemination speed and breadth to evaluate the dissemination efficiency of information in the entire social network, expressed as: ; Among them, represents the number of node pairs in the propagation path set , and is the propagation speed of each pair of nodes in the path is the breadth of the propagation path; Quantification of node influence: Based on the influence nodes in the information dissemination path, quantify the influence of each node in the information diffusion process, and generate a score of node influence in combination with the connection degree and dissemination ability of the node, as an input feature for box office prediction; The box office prediction module predicts the movie box office through a movie box office prediction model in combination with the simulated information dissemination path and speed; The model verification and optimization module verifies and optimizes the movie box office prediction model through cross-validation and adjusts the model parameters; The result display module presents the predicted movie box office results to users in the form of charts or reports, and provides box office growth predictions based on the social dissemination path.
2. The movie box office prediction system based on machine learning and big data according to claim 1, characterized in that The data collection module includes: Historical box office data interface: Automatically obtains historical box office data for various time periods, including release time, total box office, and number of moviegoers, by connecting to the API of the movie box office database. Social media data scraping tool: Uses the public API of social media platforms to scrape real-time social media feedback data related to movies, including user comments, likes, shares, and sentiment analysis results. Collection of marketing campaign data: Obtains data on the investment and effectiveness of movie marketing campaigns, including ad click-through rate, exposure, and interaction rate, by docking with the data interfaces of major advertising platforms and marketing channels. Tool for collecting information on competing movies in the same period: Automatically collects relevant data on movies released in the same period as the target movie, including the box office performance, marketing strategies, and audience evaluations of competing movies.
3. A movie box office prediction system based on machine learning and big data according to claim 1, characterized in that, The data processing and feature engineering module includes: Data cleaning: Cleans the collected market data using the Z-Score algorithm to identify and process abnormal data, missing data, and duplicate data. Data integration: Aligns and integrates market data from different sources in terms of time and space through multi-dimensional data alignment technology. Feature construction: Based on the factors affecting movie box office, constructs a feature set, specifically including: Movie type feature: Converts the movie type into multiple binary features using One-Hot encoding to capture the impact of different movie types on box office. Director and actor influence characteristics: Based on historical box office data, calculate the influence index of a director or actor , expressed as: ; Among them, represents the box office revenue of the th movie, represents the budget of the th movie, and is the total number of movies in which the director or actor participated; Release time feature: Converts the release date into holiday, weekend, or weekday features, and identifies the impact trend of release time on box office through time series analysis. Budget scale feature: Converts the budget scale of the movie into a comparable feature value through normalization to analyze the box office performance of movies with different budget scales. Audience emotional characteristics: Through natural language processing technology, perform sentiment analysis on social media feedback data, extract sentiment tendency scores as emotional characteristics, and calculate the audience emotional index by combining with sentiment intensity weighting, which is expressed as: , expressed as: ; Among them, is the sentiment score of the th feedback, is the sentiment intensity weight, is the total number of feedbacks.
4. A movie box office prediction system based on machine learning and big data according to claim 1, characterized in that, The social network construction includes: Node definition and identification: Regards users on social media platforms and movie-related content as nodes respectively, where user nodes represent social media users and information nodes represent movie-related content. Definition and generation of edges: Based on the interaction relationships among users and the associations between users and information nodes, edges are generated between nodes to form a preliminary social network graph, and the weights of the edges are represented as: ; Among them, represents the interaction frequency between users ; represents the association strength between users and information nodes; represents the similarity between information nodes ; , , are the corresponding weight coefficients; Dynamic social network update: Real-time updates the social network graph through a dynamic social network evolution model, captures changes in the interaction relationships between users over time, and dynamically adjusts the weights of nodes and edges to reflect the real-time social dissemination state, expressed as: ; Among them, represents the adjacency matrix at the moment , represents the change in the network structure from the moment to during this period. ; Among them, represents the probability of generating an edge between nodes and at time is the eigenvector of node at time ; represents the eigenvector of node at time ; Network structure optimization: When constructing the dynamic social network graph, uses the Louvain algorithm to optimize the social network graph, identifies the community structure in the network, and improves the resolution and stability of the social network graph by optimizing the distribution of nodes and edges, expressed as: ; where, represents modularity, is the number of edges in the network, and are the degrees of nodes and respectively, is the value in the adjacency matrix, is the indicator function.
5. A movie box office prediction system based on machine learning and big data according to claim 4, characterized in that, The quantification of node influence includes: Node connectivity calculation: The node connectivity represents the number of connections of the node with other nodes in the social network, expressed as: ; Among them, is the set of all nodes; Node propagation ability calculation: Node 's propagation ability is quantified by analyzing the node's position and influence in the propagation path, expressed as: ; Node influence score calculation: The node 's influence score is quantified by comprehensively considering its connectivity and propagation ability, expressed as: ; Among them, and are weight coefficients; Application of Node Influence in Box Office Prediction: The Generated Node Influence Score As a feature of the social propagation path and as an input for box office prediction.
6. A movie box office prediction system based on machine learning and big data according to claim 1, characterized in that, The movie box office prediction model uses a Bayesian optimized neural network, and the Bayesian optimized neural network includes: Feature preprocessing and fusion: Normalizes the social dissemination path, speed, and node influence features, and fuses them to generate a feature vector for input to the neural network, expressed as: ; Among them, is the feature vector input to the fused neural network, , , are the social propagation path feature, the propagation speed feature, and the node influence feature respectively, , , are the fusion weights; Network structure design: Introduces a processing mechanism for social network features into the neural network structure, expressed as: Input layer: ; Branch-type hidden layer: ; ; ; ; Among them, is the fused feature vector of the input layer, , , are the weight matrices of the social propagation path feature, the propagation speed feature, and the node influence feature respectively, , , are the bias terms of the social propagation path feature, the propagation speed feature, and the node influence feature respectively, , , are the outputs of the social propagation path feature, the propagation speed feature, and the node influence feature respectively, is the activation function, is the fused feature vector after processing; Box office prediction output: The fused features are input into the output layer to generate the final movie box office prediction results, expressed as: ; Among them, is the predicted movie box office value, and are the weight matrix and bias term of the output layer respectively; Feature importance analysis and feedback mechanism: Feature importance analysis and feedback mechanism are introduced to enhance the interpretability of the model and make feedback adjustments based on the prediction error, expressed as: ; ; Among them, represents the influence degree of the th fusion feature on the prediction result, is the th fusion feature, is the adjustment amount of the th feature weight, is the loss function, is the true box office value, is the learning rate.
7. A movie box office prediction system based on machine learning and big data according to claim 6, characterized in that, The model verification and optimization module includes: Cross-validation process: The model validation and optimization module validates the movie box office prediction model through k-fold cross-validation, expressed as: ; Among them, is the average error of k-fold cross-validation, is the prediction error in the i-th fold, is the predicted value on the validation set of the i-th fold, is the true value of the i-th fold; Model parameter optimization: After completing cross-validation, adjust the model parameters by minimizing the average error of cross-validation to adjust the model parameters. The model parameter optimization adopts the Bayesian optimization algorithm, which is expressed as: ; Among them, represents the set of hyperparameters of the model, is the average cross-validation error under the corresponding parameters, is the regularization term of the hyperparameters, is the regularization coefficient; Adaptive adjustment: During the model parameter optimization process, the model parameters are adjusted based on the gradient descent algorithm, which is expressed as: ; Among them, represents the model parameters after the next iteration, is the current parameter and is the gradient of the cross-validation error under the current parameter.
8. A movie box office prediction system based on machine learning and big data according to claim 1, characterized in that, The result display module includes: Chart generation: The movie box office results predicted by the movie box office prediction model are presented to users in the form of charts, including line charts, bar charts, and pie charts; Report generation: Automatically generate a comprehensive report based on the output data of the movie box office prediction model, including the predicted movie box office results, the model's performance indicators, and related analysis instructions; Box office growth prediction based on social communication paths: Combine social communication paths to predict the dynamic changes in movie box office and provide box office growth trends within a predetermined period of time in the future, expressed as: ; Among them, represents the box office growth amount obtained based on social dissemination path analysis, is the predicted box office value at the current moment , is the predicted box office growth value after considering the social dissemination path.
Citation Information
Patent Citations
Pre-projection prediction method for movie box office
CN113379448A