Construction method and device of scenic spot passenger flow prediction model
By building a scenic spot passenger flow prediction model and combining a variety of data information and technical means, the problem of inaccurate passenger flow prediction in the existing technology is solved, and accurate prediction and effective management of scenic spot passenger flow are achieved.
Patent Information
- Application Number
- CN202510182792.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-13
AI Technical Summary
The lack of mature tourist flow prediction methods in the existing technology has led to the passenger flow prediction of scenic spots that rely mostly on fuzzy estimation of personal experience or historical data, and cannot effectively deal with the influence of external factors.
By constructing a scenic spot passenger flow prediction model, time information, geographical information, meteorological information, holidays, public opinion information, ticketing data, macroeconomic indicators and historical passenger flow information are collected, data processing and feature screening are carried out, regression models are built for prediction, and prediction accuracy is improved through technical means such as model tuning and seasonal decomposition.
Accurate prediction of the passenger flow of scenic spots is achieved, helping scenic spots to arrange resources reasonably, improve tourists' safety and tourism experience, provide data to support long-term planning, and organize evacuations quickly and effectively in emergencies.
Smart Images

Figure CN120146263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and specifically provides a method and device for constructing a scenic area passenger flow prediction model. Background Art
[0002] Scenic area passenger flow prediction is an important technical means in the fields of tourism management, urban planning, etc., and is of great significance for ensuring tourist safety, improving the tourism experience, optimizing resource allocation, etc. At present, there is no mature technology for scenic area passenger flow prediction, and most scenic areas make fuzzy estimates of the passenger flow based on personal experience or past historical data.
[0003] The tourism industry is sensitive to factors such as location, climate, economy, public opinion, etc., and external stimuli will directly affect the passenger flow, and the role of these factors cannot be ignored. Summary of the Invention
[0004] The present invention aims at the deficiencies of the above-mentioned existing technologies and provides a method for constructing a scenic area passenger flow prediction model with strong practicability.
[0005] A further technical task of the present invention is to provide a device for constructing a scenic area passenger flow prediction model with reasonable design, safety and applicability.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0007] A method for constructing a scenic area passenger flow prediction model has the following steps:
[0008] S1. Data collection;
[0009] S2. Data processing;
[0010] S3. Feature screening and construction of an index system;
[0011] S4. Training a regression model;
[0012] S5. Model tuning;
[0013] S6. Model prediction.
[0014] Further, in step S1, time information, geographical information, meteorological information, holidays, public opinion information, ticket data, macroeconomic indicators, and historical passenger flow information are collected;
[0015] And it is necessary to regularly collect the weather information, holiday information, public opinion information of the scenic area for the next month and the ticket data for the next month reserved online this month.
[0016] Further, in step S2, it includes:
[0017] (1) Filling missing values, removing outliers, standardizing and normalizing structured data;
[0018] (2) Perform feature encoding on categorical features;
[0019] (3) Construct a sentiment analysis sub-model for social media comments;
[0020] (4) Use a seasonal decomposition sub-model or a Prophet sub-model to perform seasonal adjustment on tourism time series data.
[0021] Furthermore, in step (3), obtain comment data from the social media platform, clean and preprocess the obtained comment data, perform word segmentation processing to convert the text into a processable form, extract keywords from the comment text using TF-IDF to represent sentiment, and calculate the sentiment value of the comment according to the sentiment dictionary.
[0022] Furthermore, in step (4), input historical passenger flow data for a certain period of time, preprocess the time series data, select the smoothing coefficient α, trend coefficient β, seasonal coefficient γ, and seasonal cycle degree m, and determine the model parameters through the cross-validation method;
[0023] Fit the model with the training dataset, solve the trend and seasonal components, and output the trend component, seasonal component, and residual data.
[0024] Furthermore, in step S3, perform variance filtering on the features, remove features with low discrimination, remove features with low correlation with the label through correlation filtering, reduce the feature dimension, remove worthless features, construct an index system, and reduce the complexity of model construction.
[0025] Furthermore, in step S4, construct different types of regression prediction models, and select the model with the best performance for prediction according to the model evaluation index;
[0026] (A) Support vector machine sub-model combined with the fireworks algorithm;
[0027] a. Initialize the fireworks population;
[0028] Set the size and number of iterations of the fireworks population, randomly generate an initial fireworks solution vector, calculate its fitness value, and store the optimal solution and the optimal fitness value;
[0029] b. Enter the iterative loop, for each fireworks solution vector;
[0030] Generate a new solution vector: Generate a new solution vector according to the current solution vector and the defined explosion strategy;
[0031] Calculate the fitness value of the new solution vector: Use the support vector machine model to train the new solution vector and calculate its fitness value through appropriate evaluation indexes;
[0032] Determine whether to update the optimal solution: If the fitness value of the new solution vector is better than that of the optimal solution, update the optimal solution and the optimal fitness value;
[0033] c. Return the optimal solution as the result of the combination of the fireworks algorithm and the support vector machine model;
[0034] (B) XGBoost regression sub-model;
[0035] An ensemble algorithm with multiple regression trees as base learners:
[0036] (a) The model first finds an initial predicted value, which is the average value of the target values;
[0037] (b) According to the initial predicted value, calculate the error of each sample, that is, the difference between the actual target value and the predicted value;
[0038] (c) Construct a regression tree, divide the samples according to the feature values, and each division point will generate two child nodes;
[0039] (d) Calculate the new predicted value of each sample;
[0040] (e) Calculate the error of each sample again, adjust the structure of the tree, and build another tree;
[0041] (f) Obtain the prediction results of multiple decision trees, and add up the prediction results of all trees;
[0042] (g) Set the regularization parameter to control the complexity of the tree to prevent overfitting;
[0043] (C) BP neural network;
[0044] (a) Initialize, randomly generate the weight coefficients of the BP neural network, set the learning rates α and β, and the thresholds of the hidden layer and the output layer;
[0045] (b) Input multiple groups of training data;
[0046] (c) Calculate the node data of the hidden layer using the input training data;
[0047] (d) Calculate the initial output variable of the neural network algorithm;
[0048] (e) Update the variable data of the hidden layer;
[0049] (f) Update the variable data of the output layer;
[0050] (g) Update the weight coefficients of the neural network using the method of taking derivatives, and the weight update amplitudes of the input layer and the output layer;
[0051] (h) If the difference between the test data error and the training data error is within a certain range, the model training is completed; otherwise, go to step (c).
[0052] Further, in step S5, it includes:
[0053] (1) First, it is necessary to determine the parameters to be tuned and their possible value ranges;
[0054] (2) By combining the possible values of each parameter, create a grid of parameter combinations;
[0055] (3) According to the algorithm to be tuned, select an appropriate evaluation metric;
[0056] (4) For each parameter combination, use cross-validation to divide the dataset into a training set and a validation set, then use the training set to train the model, and use the validation set to calculate the evaluation metric;
[0057] (5) According to the evaluation metric results of the validation set, select the parameter combination with the best performance as the final model parameters;
[0058] (6) Retrain the model on the entire training set using the best parameter combination;
[0059] (7) Use the test set or cross-validation to evaluate the performance of the model on new data and ensure the generalization ability of the model;
[0060] Repeat the above steps until the best parameter combination is found.
[0061] Further, in step S6, perform feature extraction on the scenic area data of the current month to be predicted, ensure that the features for prediction are consistent with those of the training data, save the regression model with the best evaluation metrics, and use it to predict the scenic area passenger flow, and output the predicted values to be stored in the database.
[0062] An apparatus for constructing a scenic area passenger flow prediction model includes: at least one memory and at least one processor;
[0063] The at least one memory is used to store machine-readable programs;
[0064] The at least one processor is used to call the machine-readable programs and execute a method for constructing a scenic area passenger flow prediction model.
[0065] Compared with the prior art, a method and apparatus for constructing a scenic area passenger flow prediction model of the present invention have the following outstanding beneficial effects:
[0066] The present invention can assist in aspects such as scenic area tourism management and public safety. Through prediction, it can help scenic areas reasonably arrange resources, such as personnel scheduling, facility maintenance, etc., to cope with different passenger flow situations. Passenger flow data can provide a basis for the long-term planning of scenic areas, such as whether to expand infrastructure, add service facilities, etc. It helps scenic areas identify potential peak hours in advance, take timely measures to avoid congestion, and ensure the safety of tourists. In case of emergencies, accurate passenger flow prediction helps to quickly and effectively organize evacuations. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0068] Attached Figure 1 is a XGBoost calculation flow chart in a method for constructing a scenic area passenger flow prediction model;
[0069] Attached Figure 2 is a BP neural network calculation flow chart in a method for constructing a scenic area passenger flow prediction model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the following will further elaborate on the present invention in combination with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0071] The following gives an optimal embodiment:
[0072] In a method for constructing a scenic area passenger flow prediction model in this embodiment, relevant data of the scenic area is collected, the feature system of the model is constructed after processing different types of data, and the passenger flow of the scenic area is used as the label. Historical data is used as the training set to train and optimize the model, and the future relevant feature data of the scenic area is used as the test set to input the prediction model. The model can predict the corresponding output value according to the learned model parameters and output the predicted passenger flow. The regression model is a machine learning model used to predict the value of a continuous output variable. The regression model analyzes the relationship between features (independent variables) and labels (dependent variables) to establish a function to predict the value of the target variable.
[0073] For example, the prediction target: the passenger flow of Zhangjiajie Scenic Area in January 2024;
[0074] During the initial training, collect the daily meteorological information, holiday information, daily Baidu Index votes, daily ticketing data, monthly consumer price index of each scenic area from 2021 to 2023, and the passenger flow data of each year, month, week, and day from 2021 to 2023.
[0075] Take the daily meteorological forecast information, holiday information, daily Baidu Index votes, daily ticketing data, and monthly consumer price index of the scenic area in month T as features, and the corresponding daily passenger flow data in month T as labels to construct training data. Divide the training data into N parts using N-fold cross-validation, with N - 1 parts as the training set and 1 part as the test set for the construction and optimization of the regression model. Input the daily features of the month to be predicted into the trained model, output the predicted values of its labels, that is, the predicted daily passenger flow, and sum them up to obtain the passenger flow of the scenic area at the monthly granularity.
[0076] To enable the model to learn the latest knowledge, when predicting the passenger flow in month T subsequently, in month T - 1, add the data of month T - 2 for incremental training of the model. And input the trained model into the subsequent prediction of the passenger flow in month T.
[0077] The specific steps are as follows:
[0078] S1. Data collection;
[0079] Time information: time units such as year, season, month, day of the week, date, etc.;
[0080] Geographical information: different regions or region groups: cities, provinces, scenic areas;
[0081] Meteorological information: temperature (highest temperature, lowest temperature, average temperature), humidity, rainfall, wind force;
[0082] Holidays: whether it includes New Year's Day, Spring Festival, Tomb-Sweeping Day, Labor Day, Dragon Boat Festival, Mid-Autumn Festival, National Day;
[0083] Public opinion: Baidu Index, social media comments: number of likes, number of comments, number of forwards, comment content (sentiment analysis), generally there are anti-crawling strategies;
[0084] Ticketing data: historical online and offline ticket booking data;
[0085] Macroeconomic indicators: consumer price index CPI;
[0086] Historical passenger flow information: passenger flow data of each year, month, week, and day for at least one year in history;
[0087] Regularly collect the weather information (China Weather Network), holiday information (whether it includes holidays), public opinion information (Baidu), and the ticketing data for the next month reserved online this month of the scenic area next month.
[0088] S2. Data processing;
[0089] including:
[0090] (1) Filling missing values, removing outliers, standardizing, and normalizing structured data;
[0091] (2) Performing feature encoding on categorical features (such as seasons, holidays, etc.);
[0092] (3) Constructing a sentiment analysis sub - model for social media comments;
[0093] Obtain comment data from social media platforms, and data can be crawled through API interfaces, crawlers, etc. Clean and pre - process the obtained comment data, including removing special characters, punctuation marks, stop words, etc., and perform word segmentation to convert the text into a processable form. TF - IDF extracts keywords from the comment text to represent sentiment, and calculates the sentiment value of the comment according to the sentiment dictionary.
[0094] (4) Using the seasonal decomposition sub - model Holt - Winters or Prophet sub - model to perform seasonal adjustment on tourism time - series data;
[0095] Input historical passenger flow data for one year, and pre - process the time - series data, including stationarization, removing trends, removing seasonality, etc. Select model parameters including smoothing coefficient α, trend coefficient β, seasonal coefficient γ, and seasonal cycle degree m, etc., and determine the model parameters through methods such as cross - validation.
[0096] Use the training data set to fit the model, solve the trend and seasonal components, and output the trend component, seasonal component, and residual data.
[0097] S3. Feature screening and constructing an index system;
[0098] Perform feature screening such as variance filtering to remove features with small discrimination, and correlation filtering to remove features with low correlation with the label, etc.: reduce the feature dimension, remove worthless features, construct an index system, and reduce the complexity of model construction.
[0099] S4. Training a regression model;
[0100] Construct different types of regression prediction models, and select the model with the best performance for prediction according to the model evaluation index;
[0101] (A) Support vector machine sub - model combined with the fireworks algorithm;
[0102] a. Initialize the fireworks population;
[0103] Set the size and number of iterations of the fireworks population, randomly generate an initial fireworks solution vector, calculate its fitness value, and store the optimal solution and the optimal fitness value;
[0104] b. Enter the iterative loop. For each fireworks solution vector;
[0105] Generate a new solution vector: Generate a new solution vector according to the current solution vector and the defined explosion strategy;
[0106] Calculate the fitness value of the new solution vector: Use the support vector machine model to train the new solution vector and calculate its fitness value through appropriate evaluation metrics;
[0107] Determine whether to update the optimal solution: If the fitness value of the new solution vector is better than that of the optimal solution, update the optimal solution and the optimal fitness value;
[0108] c. Return the optimal solution as the result of the combination of the fireworks algorithm and the support vector machine model;
[0109] (B) XGBoost regression sub-model;
[0110] As Figure 1 shown, an ensemble algorithm with multiple regression trees as base learners:
[0111] (a) The model first finds an initial predicted value, the average value of the target value;
[0112] (b) According to the initial predicted value, calculate the error of each sample, that is, the difference between the actual target value and the predicted value;
[0113] (c) Build a regression tree, divide the samples according to the feature values, and each division point will generate two child nodes;
[0114] (d) Calculate the new predicted value of each sample;
[0115] (e) Calculate the error of each sample again, adjust the structure of the tree, and build another tree;
[0116] (f) Obtain the prediction results of multiple decision trees, and add up the prediction results of all the trees;
[0117] (g) Set the regularization parameter to control the complexity of the tree to prevent overfitting;
[0118] (C) As Figure 2 shown, BP neural network;
[0119] (a) Initialize, randomly generate the weight coefficients of the BP neural network, set the learning rates α and β, and the thresholds of the hidden layer and the output layer;
[0120] (b) Input multiple groups of training data;
[0121] (c) Calculate the node data of the hidden layer using the input training data;
[0122] (d) Calculate the initial output variables of the neural network algorithm;
[0123] (e) Update the variable data of the hidden layer;
[0124] (f) Update the variable data of the output layer;
[0125] (g) Update the weight coefficients of the neural network using the method of taking derivatives, and update the weight amplitudes of the input layer and the output layer;
[0126] (h) If the difference between the test data error and the training data error is within a certain range, complete the model training; otherwise, go back to step (c) of this process.
[0127] S5. Tuning of the model;
[0128] The performance of the model can be further optimized by methods such as adjusting parameters and cross-validation.
[0129] Grid search:
[0130] (1) Determine the parameter range to be tuned: First, it is necessary to determine the parameters to be tuned and their possible value ranges;
[0131] (2) Create a grid of parameter combinations: By combining the possible values of each parameter, create a grid of parameter combinations;
[0132] (3) Construct the model and evaluation metrics: According to the algorithm to be tuned, select an appropriate evaluation metric, such as the mean squared error (MSE) (representing the average degree of difference between the true value and the predicted value), the mean absolute error (MAE) (representing the average absolute degree of difference between the true value and the predicted value), etc. to measure the performance of the model;
[0133] (4) Conduct grid search: For each parameter combination, use cross-validation to divide the dataset into a training set and a validation set, then use the training set to train the model, and use the validation set to calculate the evaluation metrics;
[0134] (5) Select the best parameter combination: According to the evaluation metric results of the validation set, select the parameter combination with the best performance as the final model parameters;
[0135] (6) Train the model with the best parameters: Retrain the model on the entire training set using the best parameter combination;
[0136] (7) Evaluate the model performance: Use the test set or cross-validation to evaluate the performance of the model on new data to ensure the generalization ability of the model;
[0137] Repeat the above steps until the best parameter combination is found.
[0138] S6. Prediction of the model;
[0139] Extract features from the scenic area data of the month to be predicted to ensure that the features for prediction are consistent with those of the training data. Save the regression model with the best evaluation index to predict the passenger flow of the scenic area, and output the predicted values and store them in the database.
[0140] Based on the above method, in this embodiment, a device for constructing a scenic area passenger flow prediction model includes: at least one memory and at least one processor;
[0141] The at least one memory is used to store machine-readable programs;
[0142] The at least one processor is used to call the machine-readable programs and execute a method for constructing a scenic area passenger flow prediction model.
[0143] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any technical solution that conforms to the technical solutions described in the above specific embodiments of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art shall fall within the patent protection scope of the present invention.
[0144] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a scenic spot passenger flow prediction model, characterized in that: The steps are as follows: S1, data collection; S2, data processing; S3, feature screening and index system construction; S4, training regression model; S5. Model tuning; S6. Model predictions.
2. The method for constructing a scenic spot passenger flow prediction model according to claim 1, characterized in that: In step S1, time information, geographic information, weather information, holidays, public opinion information, ticketing data, macroeconomic indicators and historical passenger flow information are collected; It is also necessary to regularly collect weather information, holiday information, public opinion information of the scenic spot for the next month and ticket data for next month booked online this month.
3. The method for constructing a scenic spot passenger flow prediction model according to claim 2, characterized in that: In step S2, it includes: (1) Fill missing values, remove outliers, standardize and normalize structured data; (2) Feature encoding of categorical features; (3) Construct a sentiment analysis sub-model for social media comments; (4) Use the seasonal decomposition sub-model or the Prophet sub-model to seasonally adjust the tourism time series data.
4. The method for constructing a scenic spot passenger flow prediction model according to claim 3, characterized in that: In step (3), the comment data is obtained from the social media platform, cleaned and preprocessed, and word segmentation is performed to convert the text into a processable form. TF-IDF extracts keywords from the comment text to express emotions, and the sentiment value of the comment is calculated based on the sentiment dictionary.
5. The method for constructing a scenic spot passenger flow prediction model according to claim 4, characterized in that: In step (4), historical passenger flow data of a certain period of time is input, the time series data is preprocessed, the smoothing coefficient α, the trend coefficient β, the seasonal coefficient γ and the seasonal cycle degree m are selected, and the model parameters are determined by the cross-validation method; The training data set is used to fit the model, solve the trend and seasonal components, and output the trend component, seasonal component, and residual data.
6. The method for constructing a scenic spot passenger flow prediction model according to claim 5, characterized in that: In step S3, variance filtering is performed on the features to remove features with low discrimination, correlation filtering is performed to remove features with low correlation with labels, the feature latitude is reduced, worthless features are removed, an indicator system is constructed, and the complexity of model construction is reduced.
7. The method for constructing a scenic spot passenger flow prediction model according to claim 6, characterized in that: In step S4, different types of regression prediction models are constructed, and the model with the best performance is selected for prediction according to the model evaluation index; (A) Support vector machine model combined with the fireworks algorithm; a. Initialize the fireworks group; Set the size and number of iterations of the fireworks group, randomly generate the initial fireworks solution vector, calculate its fitness value, and store the optimal solution and optimal fitness value; b. Enter the iteration loop for each firework solution vector; Generate new solution vector: Generate a new solution vector based on the current solution vector and the defined explosion strategy; Calculate the fitness value of the new solution vector: Use the support vector machine model to train the new solution vector and calculate its fitness value through appropriate evaluation indicators; Determine whether to update the optimal solution: If the fitness value of the new solution vector is better than the fitness value of the optimal solution, update the optimal solution and the optimal fitness value; c. Return the optimal solution as the result of combining the Fireworks algorithm with the support vector machine model; (B) XGBoost regression sub-model; An ensemble algorithm that uses multiple regression trees as base learners: (a) The model first finds an initial prediction value, the average of the target value; (b) Based on the initial predicted value, calculate the error of each sample, that is, the difference between the actual target value and the predicted value; (c) Build a regression tree and divide the samples according to the feature values. Each division point will generate two child nodes. (d) Calculate the new predicted value for each sample; (e) Calculate the error of each sample again, adjust the tree structure, and build another tree; (f) Get the prediction results of multiple decision trees and add the prediction results of all trees together; (g) Set regularization parameters to control the complexity of the tree to prevent overfitting; (C) BP neural network; (a) Initialization: randomly generate the weight coefficients of the BP neural network, set the learning rates α and β, and the thresholds of the hidden layer and the output layer; (b) Input multiple sets of training data; (c) Calculate the node data of the hidden layer using the input training data; (d) calculating the initial output variables of the neural network algorithm; (e) Update the variable data of the hidden layer; (f) Update the variable data of the output layer; (g) Using the derivation method to update the weight coefficients of the neural network, the input layer and the output layer update the weight amplitude; (h) If the difference between the test data error and the training data error is within a certain range, the model training is completed; otherwise, go to step (c).
8. The method for constructing a scenic spot passenger flow prediction model according to claim 7, characterized in that: In step S5, it includes: (1) First, it is necessary to determine the parameters to be tuned and their possible value ranges; (2) Create a grid of parameter combinations by combining the possible values of each parameter; (3) Select an appropriate evaluation metric based on the algorithm to be tuned; (4) For each parameter combination, use cross-validation to divide the dataset into a training set and a validation set, then use the training set to train the model and use the validation set to calculate the evaluation metrics; (5) Based on the evaluation index results of the validation set, select the parameter combination with the best performance as the final model parameters; (6) Retrain the model on the entire training set using the best parameter combination; (7) Use a test set or cross-validation to evaluate the performance of the model on new data to ensure the generalization ability of the model; Repeat the above steps until the best parameter combination is found.
9. The method for constructing a scenic spot passenger flow prediction model according to claim 8, characterized in that: In step S6, feature extraction is performed on the scenic spot data to be predicted for the current month to ensure that the predicted features are consistent with the features of the training data, and the regression model with the best evaluation index is saved to predict the scenic spot passenger flow, and the output prediction value is stored in the database.
10. A device for constructing a scenic spot passenger flow prediction model, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 9.