An AI large language model-based scenic spot passenger flow trend prediction analysis method
By using an AI-driven large language model to drive a multi-agent system, integrating multi-source data to build a knowledge base and fine-tuning the model, the problem of a single data source in traditional scenic area visitor flow management is solved, enabling more accurate visitor flow trend prediction and intelligent management.
Patent Information
- Application Number
- CN202411339977.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-09-25
AI Technical Summary
Traditional tourist flow management relies on a single data source, which lacks comprehensiveness and depth, leading to inaccurate predictions of tourist flow trends.
The system employs an AI-driven large language model to drive a multi-agent system, integrating LBS data, scenic area data, holiday information, and weather information. It uses RAG technology to build a multi-dimensional knowledge base, constructs a general model, and fine-tunes it to enable multi-agent collaborative work for passenger flow trend prediction.
It improves the accuracy and intelligence of scenic area visitor flow trend prediction, provides prediction results in natural language form, facilitates scenic area management decisions, and enhances the efficiency and intelligence of scenic area management.
Smart Images

Figure CN119398214B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and big data analysis technology, specifically involving a method for predicting and analyzing tourist flow trends in scenic areas based on an AI large language model. Background Technology
[0002] In traditional scenic area visitor flow management, data collection and analysis often rely on a single data source, such as ticket sales data or online visitor behavior data. While this method can reflect visitor flow trends to some extent, it lacks comprehensiveness and depth. With the development of big data technology, more and more data sources can be utilized, such as social media data, meteorological data, and traffic flow data. The integration and analysis of these data can more accurately predict visitor flow trends in scenic areas. Summary of the Invention
[0003] To address the aforementioned technical problem—namely, how to more accurately predict visitor flow trends in scenic areas—this invention provides the following technical solution:
[0004] A method for predicting and analyzing tourist flow trends in scenic areas based on AI large language models includes the following steps:
[0005] S1. Full Data Collection and Processing: Collect data from multiple sources related to visitor flow in the scenic area, including LBS data, scenic area data, holiday information and weather information, and third-party map data. Clean, integrate and standardize this data.
[0006] S2. Building a knowledge base based on multi-dimensional data: Standardized scenic area data and operator big data are input into the big data platform, and the processed data is integrated into the knowledge base to build a comprehensive industry and scenario knowledge base to store and organize various types of information; RAG technology is used to enable the knowledge base to search and generate content.
[0007] S3. Capability to Build General Models: Utilizing machine learning techniques, general models are pre-trained based on multi-dimensional data in a knowledge base to identify patterns and relationships in the data; general models include scenic area visitor flow prediction models, travel chain analysis models, and visitor profile models; simultaneously, based on pre-training, fine-tuning techniques are used to fine-tune the general models using multiple sources of data related to scenic area visitor flow;
[0008] S4. Construct a multi-agent system based on general model capabilities: Construct a multi-agent system including a travel group monitoring agent, a full-chain travel trajectory query agent, and a scenic spot visitor flow query agent based on the general model obtained through pre-training and fine-tuning.
[0009] S5. Predictive analysis based on LLM-driven multi-agent model: By integrating the AI large language model LLM, multiple agents are driven to work collaboratively to perform intelligent predictive analysis of tourist flow trends in scenic areas.
[0010] Furthermore, LBS data includes tourists' location information and movement trajectories. LBS data includes: user data, location data, timestamps, movement trajectories, geospatial data, points of interest data, user behavior data, and social network data.
[0011] Furthermore, in step S1, data cleaning includes removing irrelevant characters, removing duplicates, removing noise, filling missing values, validating data types and formats, and checking for logical errors. Noise removal includes identifying and cleaning non-numeric values in the ticket sales data of the scenic area data. Logical error checking includes ensuring that the prices in the ticket sales data of the scenic area data do not exceed the set upper limit for ticket sales prices.
[0012] Further, step S2 includes,
[0013] S21. Data Integration: Integrate the processed data, including LBS data, scenic spot data, holiday information, weather information, and third-party map data.
[0014] S22. Feature Extraction: Using model predictive analysis, key features and patterns are extracted from the integrated data. Key features include tourist behavior patterns, peak tourist flow periods, and the impact of weather on tourist flow. Clustering algorithms are used to identify tourist behavior patterns, time series analysis is used to predict peak tourist flow periods, and regression models are used to assess the potential impact of weather changes on the number of tourists.
[0015] S23. Rule and Pattern Recognition: Identify rules and patterns in the data through model analysis, including using hypothesis testing to determine the correlation between holidays and increased tourist traffic, and applying time series analysis to identify the impact of weather changes on tourist traffic.
[0016] S24. Use RAG technology to build industry and scenario knowledge bases: integrate the extracted features, identified rules and patterns, and generated conclusions into a structured database to build a knowledge base.
[0017] Furthermore, step S21 specifically includes:
[0018] S211. Read LBS data, holiday information, weather information, and third-party map data;
[0019] S212. Perform data information and transformation operations, including removing invalid records from LBS data and converting timestamp formats;
[0020] S213. Using timestamps and location information as key fields, connect multiple datasets of LBS data, holiday information, weather information, and third-party map data to form a unified data view;
[0021] S214. Select the columns required for analysis from the concatenated dataset. These columns include timestamps, location information, holiday information, weather information, and third-party map data. Rename the selected columns to unify the format of the data view.
[0022] S215. Store the formatted data in the knowledge base.
[0023] Furthermore, in step S22, the use of clustering algorithms to identify tourist behavior patterns specifically includes: selecting location data, timestamps, and weather information as features, applying the K-means algorithm to cluster tourist behavior, and identifying different tourist behavior patterns.
[0024] Furthermore, in step S23, hypothesis testing is used to determine the correlation between holidays and increased tourist traffic, including:
[0025] Add holiday labels to the dates in the integrated dataset;
[0026] Construct a linear regression model to assess the impact of holidays on tourist numbers;
[0027] Output a statistical summary of the linear regression model.
[0028] Further, step S3 includes:
[0029] S31. Model pre-training: Using machine learning techniques, a general model is pre-trained based on multidimensional data in a knowledge base;
[0030] S32. Model fine-tuning: Using multiple sources of data related to tourist flow in scenic areas, fine-tune the pre-trained model, including adjusting the output layer of the model to adapt to the output format of the scenic area tourist flow prediction model, travel chain analysis model, and tourist profile model.
[0031] Further, step S4 includes:
[0032] S41. Design of the intelligent agent architecture: This includes defining the role and responsibilities of each intelligent agent and the interaction methods between intelligent agents; the intelligent agents include the travel group monitoring intelligent agent, the full-chain travel trajectory query intelligent agent, and the scenic spot visitor flow query intelligent agent;
[0033] S42. General Model Integration: Based on the general model obtained through pre-training and fine-tuning, the general model is integrated into various intelligent agents. This includes integrating a deep learning model into the travel group monitoring intelligent agent to analyze tourist behavior patterns, integrating a sequence prediction model into the full-chain travel trajectory query intelligent agent to track tourist movement paths, and integrating a time series analysis model into the scenic area visitor flow query intelligent agent to provide visitor flow data.
[0034] S43. Function Implementation and Optimization: The intelligent agent for monitoring travel groups includes an event detection system for identifying and responding to abnormal patterns in tourist behavior; the intelligent agent for querying the entire travel trajectory includes a path tracking algorithm for processing and analyzing mobile data; the intelligent agent for querying scenic area visitor flow includes a query optimization system for responding to visitor flow data query requests.
[0035] S44. Collaborative Working Mechanism: Intelligent agents communicate with each other through well-defined communication protocols and data sharing mechanisms.
[0036] Further, step S5 includes,
[0037] S51, Language Input and Intent Parsing, including:
[0038] S511. The scenic area administrator inputs queries or instructions into the system via natural language. The input includes requests for visitor flow prediction for specific dates, time periods, or events. The LLM first receives the natural language input from the scenic area administrator. The LLM automatically recognizes and parses the administrator's intent to determine the type of task to be performed.
[0039] After receiving the input, the S512 and LLM use their Natural Language Understanding (NLU) technology to analyze the administrator's query intent.
[0040] S513. Based on the analysis results, LLM determines the specific task types that need to be performed, including short-term passenger flow forecasting, long-term trend analysis, and real-time data monitoring and abnormal traffic warning.
[0041] S52. Agent Invocation and Analysis Execution: After determining the task type, based on the parsing results, the LLM drives the multi-agent system to work collaboratively and invoke the corresponding sub-agents; the agents work collaboratively, share information and resources, select the best database path, access and analyze historical data, real-time data and other relevant information stored in the knowledge base, and select the corresponding prediction model to perform scenic area visitor flow prediction analysis.
[0042] S53. Result Feedback and Text Output: After completing the analysis, the sub-agent will feed back the prediction results to the LLM, and the LLM will provide the prediction analysis results to the scenic area administrator in natural language.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] (1) This invention integrates an AI large language model to drive a multi-agent system for predicting and analyzing tourist flow trends in scenic areas, achieving collaborative work and decision support for the multi-agent system. This method not only improves the efficiency and effectiveness of the overall system and more accurately predicts tourist flow trends in scenic areas, but also returns the prediction results in natural language, enabling scenic area managers to intuitively understand and use these prediction results. This solves the problem of complex and difficult-to-understand results presentation in traditional systems, making decision-making more convenient and providing a more efficient, intelligent, and user-friendly solution for scenic area tourist flow management.
[0045] (2) By using the scenic area visitor flow trend prediction and analysis method, scenic area managers can more accurately grasp and predict visitor flow trends and take corresponding countermeasures in advance. This method not only improves the accuracy of prediction, but also enhances the level of intelligent management of scenic areas, providing technical support for the sustainable development of scenic areas.
[0046] (3) By integrating LBS location data, scenic area data, holiday information, third-party maps, and weather data, this invention constructs a multi-dimensional knowledge base and utilizes RAG technology to enhance the search and generation capabilities of the knowledge base and external data. Through model pre-training and fine-tuning, general model capabilities including scenic area visitor flow prediction, travel chain analysis, and visitor profiling have been developed. The application of these technologies enables this invention to provide scenic area managers with visitor flow trend prediction and analysis based on natural language input, realize the collaborative work and decision support of multi-agent systems, and improve the intelligence level and response efficiency of scenic area management. Attached Figure Description
[0047] Figure 1 This is a flowchart of the scenic area visitor flow trend prediction and analysis method of the present invention;
[0048] Figure 2 This is a data processing architecture diagram for the scenic area visitor flow trend prediction and analysis method of the present invention. Detailed Implementation
[0049] The technical solution of the present invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are not all embodiments of the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0050] like Figure 1As shown, this invention provides a method for predicting and analyzing tourist flow trends in scenic areas based on an AI Large Language Model (LLM). This method utilizes big data and artificial intelligence technologies, comprehensively applying multi-source data such as LBS data, scenic area data, holiday information, third-party maps, and weather information to construct a comprehensive multi-dimensional knowledge base. Simultaneously, it leverages RAG technology to enhance knowledge search and generation capabilities. Based on this, through model pre-training and fine-tuning, general model capabilities are developed, and a multi-agent system is further constructed. Then, based on the natural language processing capabilities of the AI Large Language Model (LLM), the LLM can intelligently parse the text input of scenic area administrators, driving collaborative work among multiple agents to achieve intelligent prediction and analysis of tourist flow trends. This method aims to improve the scientific rigor and foresight of scenic area management, optimize the tourist experience, and provide data support for scenic area operational decisions.
[0051] like Figure 1 As shown, the method for predicting and analyzing visitor flow trends in scenic areas specifically includes the following steps:
[0052] S1. Full Data Acquisition and Processing:
[0053] Data related to visitor flow in scenic areas is collected from multiple sources, including LBS data, scenic area data, holiday and weather information, and third-party map data. This data undergoes a series of processes such as cleaning, integration, and standardization to prepare for subsequent analysis and model training. This enables better prediction of future visitor flow trends and peaks, helping the scenic area to formulate reasonable operational plans. Step S1 includes:
[0054] S11. Full Data Acquisition:
[0055] Full data collection is a fundamental step in the scenic area visitor flow trend prediction and analysis method. To obtain more accurate prediction and analysis results, this method requires a large amount of raw data, more diverse data dimensions, and more complex annotation methods and evaluation criteria.
[0056] Data related to visitor flow in scenic areas is collected from multiple sources, including LBS data, scenic area data, holiday and weather information, and third-party map data. Data collection is automated through API interfaces for data exchange and manual upload of data files via FTP. This process utilizes automated tools and APIs to scrape data from different sources, ensuring the comprehensiveness and diversity of the data. Simultaneously, the acquired visitor data is used to create user tags and profiles through tagging. This data provides rich background information and predictive basis for scenic area visitor flow forecasting models.
[0057] 1. LBS (Location Based Services) data:
[0058] Real-time visitor flow analysis is crucial for effective management and planning in scenic areas, requiring accurate location information and real-time data updates to predict and provide early warnings of visitor flow trends. This method requires collecting LBS data, including visitor location information and movement trajectories. Simultaneously, third-party map data is utilized to enhance the accuracy and coverage of geographic locations.
[0059] LBS (Location-Based Services) technology uses mobile wireless networks and mobile positioning sensors to obtain a user's real-time geographical location information. This technology combines positioning technology and the mobile internet, not only accurately tracking the user's location but also providing various location-based services and functions. LBS data aggregates several key pieces of information, including:
[0060] User data: users' basic information and preferences;
[0061] Location data: Core data that records the user's specific location;
[0062] Timestamp: Records the specific time when location information was obtained;
[0063] Movement trajectory: The path a user moves over a period of time;
[0064] Geospatial data: Geographic area information relating to the user's location;
[0065] POI (Point of Interest) data: Information about points of interest visited by users;
[0066] User behavior data: User activity and behavior patterns in specific locations;
[0067] Social network data: User activity and interactions on social networks.
[0068] This data can be acquired through various technologies, such as outdoor positioning technologies like GPS and BeiDou, and indoor positioning technologies like WiFi and Bluetooth. Location-based services (LBS) technology can collect location information from a large number of mobile phone users at a single point in time, which is particularly important in China where mobile internet is highly prevalent.
[0069] In terms of tourist flow forecasting, LBS data offers particularly significant advantages. Its high spatiotemporal accuracy, massive data capacity, and time-series characteristics provide a solid data foundation for tourist flow forecasting and analysis. This data not only helps managers monitor current tourist flow in real time but also allows for the prediction of tourist flow trends through analytical forecasting models, thereby providing decision support for the operation, management, and planning of the scenic area.
[0070] 2. Scenic Area Data:
[0071] Collect and summarize scenic area data, including information on attractions, visitor capacity, facility distribution, ticket sales data, current visitor flow data, historical visitor flow data, and visitor behavior data.
[0072] (1) Ticket sales data:
[0073] Ticket sales data is a crucial data source for predicting visitor flow trends at tourist attractions, reflecting the total number of tickets sold within a specific time period. It can be statistically analyzed by day, week, month, or specific event period, helping managers understand the scale of visitor traffic. Data from online platforms, such as ticket bookings, hotel reservations, and tourist information inquiries, can serve as a reference for predicting future visitor flow.
[0074] (2) Real-time visitor flow data of the scenic area:
[0075] The scenic area visitor flow analysis and statistics system collects data from multiple sources, including entrance and exit gates, ticketing data, checkpoint visitor flow data, and Wi-Fi sniffing data from the core scenic area, to achieve real-time statistics on the number of visitors. Employing technologies such as sensors and cameras, the system accurately records information such as the number, time, and location of visitors entering the scenic area. Through machine learning technology, the collected data is processed, analyzed, and visualized, allowing scenic area management departments to understand visitor flow, density, and distribution in real time.
[0076] Based on real-time visitor flow data, scenic areas can formulate reasonable visitor control strategies to optimize visitor experience and improve management efficiency. When visitor flow is too high, the system can implement flow control measures or use LED displays to inform visitors of the current visitor capacity, thereby avoiding overcrowding. Furthermore, the system can also regulate visitor numbers in real time by extending wait times, limiting the number of visitors allowed in, and adjusting tour routes to ensure the safety and order of the scenic area.
[0077] (3) Historical passenger flow data:
[0078] Historical visitor flow data can help identify cyclical patterns and trends in tourist traffic, such as seasonal fluctuations and holiday peaks. It can also be used to analyze the impact of various factors (such as weather, economic conditions, and special events) on visitor flow. By analyzing past visitor flow data, predictive models can be trained and validated, improving the accuracy of future predictions.
[0079] Methods for obtaining historical passenger flow data include:
[0080] Internal data collection: Data such as the number of tourists and the time of ticket purchase are obtained through the scenic area's ticket sales system, as well as the scenic area's entry records. The number of people entering the park each day is recorded using the counter at the entrance or the electronic ticketing system.
[0081] External data acquisition: On the one hand, we obtain historical tourism statistics from tourism management departments; on the other hand, we obtain passenger flow data of transportation tools (such as buses, subways, trains, etc.) to scenic spots from transportation management departments.
[0082] Social media and online platforms: By analyzing discussions and comments on social media, we can obtain tourist feedback and behavioral patterns. We can also obtain tourist booking data and review information from travel booking websites and travel review platforms.
[0083] Historical event records: Obtain records of historical events in the scenic area (such as major events, natural disasters, etc.) and analyze the impact of these events on visitor flow.
[0084] (4) Tourist behavior data:
[0085] Visitor behavior patterns within scenic areas, such as routes, duration of stay, and spending habits, help in understanding visitor preferences and optimizing services. By analyzing visitor routes and stay times, it's possible to predict visitor flow at different times and attractions, helping scenic area managers allocate resources and personnel more effectively.
[0086] 3. Holiday Information:
[0087] Integrate holiday information to identify the impact of special dates on passenger flow.
[0088] Statutory holidays, school holidays, and special events or festivals held by scenic spots typically lead to significant changes in visitor traffic. Information on holidays and special events can help identify peak visitor periods. By analyzing visitor flow data during these times, future peak periods can be predicted, allowing for advance preparation. For example, statutory holidays such as Spring Festival, Labor Day, and National Day usually attract large numbers of tourists, leading to a surge in visitor traffic. Optimizing resource allocation in advance and understanding visitor flow trends during holidays and special events can help scenic spot managers allocate resources more rationally, such as increasing staff, adjusting opening hours, and optimizing transportation arrangements, to cope with peak visitor flow and improve the visitor experience.
[0089] 4. Weather Information:
[0090] Collect and analyze weather information to assess its potential impact on tourist behavior.
[0091] Weather conditions (such as temperature, precipitation, wind speed, and humidity) have a significant impact on tourists' travel decisions. Sunny weather tends to attract more tourists, while inclement weather may dampen tourist traffic.
[0092] S12. Data Cleaning and Preprocessing:
[0093] like Figure 2As shown, the collected data undergoes cleaning and preprocessing such as deduplication and standardization, including noise removal, missing value filling, and standardization, in preparation for subsequent analysis and model training.
[0094] 1. Data cleaning:
[0095] Data cleaning is a crucial step in ensuring data quality and accuracy. The interface server, as the first point of contact for incoming data, receives raw data from various sources. After receiving the raw data, the interface server performs basic cleaning operations, including removing irrelevant characters, removing duplicates, removing noise, filling in missing values, validating data types and formats, and checking for logical errors, ensuring the data conforms to preset standards for subsequent processing. The main tasks in this process are as follows:
[0096] Noise Removal: Identify and remove outliers or erroneous data from the data. This includes formatting errors, data type errors, or logical errors. For example, non-numeric values may appear in ticket sales data and need to be identified and cleaned.
[0097] Deduplication: Duplicate data can affect the accuracy of analytical results, especially when performing statistical analysis or training machine learning models. Deduplication identifies and removes duplicate records from a dataset. This includes not only identical records but also those that repeat on key fields.
[0098] Imputing missing values: Dealing with missing values in data. Common methods include estimating missing values using the mean, median, mode, or a predictive model. The appropriate imputation method depends on the nature of the data and the needs of the analysis.
[0099] Data type and format validation: Ensure that the data conforms to the predefined data type and format requirements, such as converting strings to date formats or numbers to floating-point numbers.
[0100] Logical error checking. This checks whether the data conforms to business logic, such as ensuring that ticket sales prices do not exceed a certain upper limit.
[0101] Through this series of cleaning steps, the collected raw data will be transformed into a clean, neat, and consistent format, providing a data foundation for subsequent data analysis, statistics, and model training.
[0102] 2. Data standardization processing:
[0103] Data standardization ensures data consistency and usability. The following is a detailed description of the data standardization process:
[0104] (1) Format conversion. For date and time data, all data that does not conform to the standard format are uniformly converted to yyyy-mm-dd format to ensure the accuracy of time series analysis and the comparability of time data.
[0105] (2) Assign a default value. Check each field in the dataset. For cases where there are null or missing values, assign a suitable default value based on business logic and data characteristics, such as 0, an empty string, or a statistical value (mean, median, etc.).
[0106] (3) Type conversion: unify the data types in different data sources. For example, convert the Number type in the database to the varchar2 type suitable for SQL operations to avoid type mismatch errors.
[0107] (4) Length transformation. For character data, adjust the length of the field as needed, such as expanding varchar2(10) to varchar2(20) to accommodate longer data content.
[0108] (5) Code conversion. For code that has changed due to system upgrades, perform corresponding updates and mappings to ensure data consistency, such as converting old code values to new code values.
[0109] (6) Remove spaces. Clean up character data by removing spaces before and after fields to ensure data cleanliness.
[0110] (7) Specific character conversion. For data containing special characters, replace or remove them according to business rules, such as replacing symbols such as +, -, *, / in calculated fields with appropriate representations.
[0111] (8) Consistent data types. Ensure that the data types in all datasets are consistent. For example, unify all date fields to date type and numeric fields to numeric type.
[0112] (9) Consistent data format. Format the text data, such as converting all text to lowercase and removing unnecessary spaces and punctuation marks, to achieve consistency in the text data.
[0113] (10) Data encoding. Encoding categorical data, such as converting weather conditions (sunny, cloudy, rainy) or holiday types into numerical labels, helps machine learning models better handle categorical features.
[0114] S2. Building a knowledge base based on multi-dimensional data:
[0115] The cleaned and standardized data, including the scenic area's own data and operator big data, is fed into a big data platform. The data is stored in HBase and Hive respectively. HBase is used for real-time query services, while Hive is primarily used for data analysis and mining services. Using Hive as a data warehouse tool, it can organize the data in HDFS, supports a SQL-like query language, and facilitates data management and analysis. Finally, the processed and analyzed data is integrated into a knowledge base, providing decision-makers with insights and support to help formulate data-driven strategies.
[0116] Using the processed data, a comprehensive industry and scenario-based knowledge base is built to store and organize various types of information. Simultaneously, RAG technology is employed to enhance the knowledge base's search and content generation capabilities with external search engines, enabling it to quickly respond to queries and provide relevant information.
[0117] Multimodal data processing is an advanced technique that enhances the predictive power and applicability of models by combining information from data of different sources and types. This approach offers significant advantages in improving the accuracy and scope of predictive models.
[0118] With the development of large-scale neural network models, multimodal large models can simultaneously process and understand multiple data types, including text, audio, images, and video, thereby obtaining a more comprehensive data perspective and exhibiting higher robustness in different contexts, thus enabling more accurate scenic area visitor flow trend prediction and analysis applications.
[0119] In the context of passenger flow trend forecasting, the application details of multimodal data processing can include the following steps:
[0120] The first step is data fusion, which involves integrating data from different modalities, including tourist reviews, visual content of the scenic area such as photos and videos, and related content on social media. This data comes from diverse platforms, such as online travel review websites, social media platforms, and the scenic area's monitoring systems.
[0121] In the feature extraction phase, each data modality can potentially offer unique insights. For example, text analysis can reveal tourists' sentimental tendencies and specific evaluations of the scenic area; image and video analysis can provide visual information about real-time crowd density and activity characteristics at the scenic area. Deep learning techniques are used to extract key features from these different modalities of data, such as sentiment in text, crowd flow in images, and activity recognition.
[0122] By applying deep learning models, such as convolutional neural networks (CNNs) to process image data or recurrent neural networks (RNNs) to process sequential data, in-depth analysis of these multimodal data can be performed. These models can capture the complex relationships between different data modalities, thereby improving the accuracy and reliability of passenger flow forecasting.
[0123] The advantage of multimodal data processing lies not only in improving prediction accuracy but also in significantly expanding the application scope of the model. Besides predicting visitor flow trends, this method can also be applied to various fields such as scenic area safety management, optimizing visitor experience, and developing marketing strategies. By comprehensively analyzing image and video content on social media, we can gain a more realistic understanding of visitors' feelings and needs, thereby providing more customized and targeted services. Step S2 includes...
[0124] S21. Data Integration:
[0125] Data integration involves combining processed data, including LBS data, scenic spot data, holiday and weather information, and third-party map data, to form a unified data view.
[0126] In the data integration phase of building the scenic area knowledge base, the first step is to collect and process data from multiple sources. This process involves acquiring data from different data sources, such as obtaining real-time tourist location data from LBS services via API, extracting historical scenic area visit records from databases, obtaining weather information from meteorological service interfaces, and integrating geospatial data from third-party map services. This data integration can be achieved through SQL queries, ETL (Extract, Transform, Load) jobs, or by using data processing frameworks such as Apache Spark.
[0127] S22, Feature Extraction:
[0128] By using model predictive analytics, key features and patterns are extracted from the integrated data. These features include tourist behavior patterns, peak tourist seasons, and the impact of weather on tourist flow.
[0129] In the feature extraction stage of the scenic area knowledge base construction process, advanced data analysis and machine learning techniques are used to identify and extract key features from the data. This step is achieved by analyzing multi-dimensional information such as tourists' location data, visit time, and weather conditions within the scenic area. For example, clustering algorithms, such as K-means or DBSCAN, are used to identify tourist behavior patterns, or time series analysis is used to predict peak visitor periods. Furthermore, regression models are used to assess the potential impact of weather changes on tourist numbers.
[0130] By selecting location, timestamp, and weather conditions as features, the K-means algorithm is applied to cluster tourist behavior, thereby identifying different tourist behavior patterns. These patterns can provide scenic area managers with insights into tourist preferences and behavioral habits, helping to optimize scenic area services and the tourist experience.
[0131] The extracted features and patterns will be analyzed in depth to formulate rules and conclusions that provide practical guidance for the operation of the scenic area. For example, the analysis may reveal a significant decrease in visitor numbers under certain weather conditions or a surge in visitor flow within specific time periods. These analytical results will be integrated into a knowledge base to provide data support for visitor flow management and resource allocation in the scenic area.
[0132] S23, Rule and Pattern Recognition:
[0133] Through model analysis, rules and patterns in the data can be identified, such as the positive impact of holidays on tourist traffic or the changing trends of tourist traffic under specific weather conditions.
[0134] In the rule and pattern recognition phase of constructing the scenic area knowledge base, statistical analysis and machine learning models are used to discover potential patterns in the data. Hypothesis testing is used to determine the correlation between holidays and increased tourist traffic, or time series analysis is applied to identify the impact of weather changes on tourist traffic.
[0135] First, a holiday label was added to the dates in the dataset. Then, a linear regression model was built to assess the impact of holidays on tourist numbers. The statistical summary of the linear regression model provides the significance level and magnitude of the holiday effect.
[0136] Furthermore, sequence analysis was used to identify patterns in how weather conditions affect tourist traffic. An ARIMA model was constructed to predict and analyze the impact of weather changes on tourist traffic.
[0137] S24. Utilize RAG technology to build industry and scenario knowledge bases:
[0138] The extracted features, identified rules and patterns, and generated conclusions are integrated into a structured database to build a knowledge base. The knowledge base is designed to facilitate information retrieval and updates. In this process, RAG technology is used to build industry-specific and scenario-specific knowledge bases.
[0139] RAG (Retrieval-Augmented Generation) technology is a deep learning model that combines retrieval and generation to improve the performance of natural language processing tasks. RAG models enhance traditional generative models by retrieving relevant information, enabling them to leverage knowledge bases to generate more accurate and richer responses.
[0140] 1. RAG Model Design:
[0141] The RAG model is designed, including a retrieval component and a generation component. The retrieval component is responsible for retrieving information related to the current task from the knowledge base. The generation component generates prediction results based on the retrieved information and the input context.
[0142] 2. Model Training:
[0143] The RAG model was trained using historical passenger flow data and relevant influencing factors.
[0144] During training, the model learns how to retrieve relevant information from the knowledge base and combine this information to predict passenger flow trends.
[0145] S3. Ability to build general models:
[0146] General-purpose model capability refers to the cross-domain and cross-task processing ability exhibited by a general-purpose model. This capability enables the general-purpose model to perform well in different application scenarios without the need for extensive customized development for each specific task. The level of general-purpose model capability directly determines the model's effectiveness and reliability in practical applications.
[0147] General models include scenic area visitor flow prediction models, travel chain analysis models, and tourist profile models. Machine learning techniques are used to pre-train these general models on multi-dimensional data from a knowledge base to identify patterns and relationships within the data. Simultaneously, based on the pre-trained models, fine-tuning techniques are employed using scenic area-related data to optimize model performance for specific tasks (such as scenic area visitor flow prediction, travel chain analysis, and tourist profile construction), enhancing their intelligent prediction and analysis capabilities for scenic area visitor flow trends. The trained models are better suited for visitor flow prediction tasks.
[0148] S31. Model pre-training:
[0149] Machine learning techniques are used to pre-train a general model based on multidimensional data in a knowledge base to identify patterns and relationships in the data.
[0150] The pre-training process begins with in-depth analysis and understanding of the multidimensional data in the knowledge base. This step involves identifying different types and structures in the data, such as text, images, and time series. After data preparation, the next step is to select a suitable machine learning model architecture for pre-training. This is a complex decision-making process that requires consideration of data characteristics, task requirements, and computational resource limitations, as well as the development of training strategies, including selecting appropriate loss functions, optimization algorithms, and learning rate scheduling, to ensure that the model can effectively learn from the data.
[0151] The final stage of pre-training is model training. In this stage, the model is trained on a large amount of data, learning patterns and associations within the data. This process typically involves significant computational resources and time, as the model needs to undergo multiple iterations on large datasets to converge. During training, the model gradually adjusts its parameters to minimize prediction error. As training progresses, the model begins to recognize high-level features in the data, which can capture complex data structures and underlying semantic information. After pre-training, the model is able to extract rich feature representations, which can serve as a starting point for downstream tasks, providing a solid foundation for fine-tuning specific tasks.
[0152] S32, Model Fine-tuning:
[0153] Fine-tuning is a crucial step in the model training process. In this step, the pre-trained model is fine-tuned using scenic area-specific data. This includes adjusting the model's output layer to adapt to the output format of a specific task, or further training the model to learn task-specific features. Simultaneously, the fine-tuned model is trained using both training and validation sets, and its performance on the validation set is monitored to better adapt it to the visitor flow prediction task.
[0154] Fine-tuning is the process of adapting a pre-trained model to a specific task. Fine-tuning techniques are a common strategy in machine learning and deep learning, allowing a model pre-trained on a large dataset to be adjusted and optimized to suit a specific task. The key to this approach is that the pre-trained model has already learned a wide range of languages and patterns, and the fine-tuning process further refines this learned knowledge to better fit the specific application scenario. In the scenario of tourist flow prediction in scenic areas, fine-tuning enables the model to more accurately understand and predict tourist behavior and flow patterns.
[0155] The method employed is Parameter-Efficient Fine-tuning (PEFT). Language models and fine-tuning are powerful tools in the field of natural language processing. Combining PEFT with parameter-efficient strategies such as LoRA and quantization can effectively leverage these models.
[0156] Parameter-Efficient Fine-tuning (PEFT) is an improved form of fine-tuning technique that aims to reduce the number of parameters that need to be adjusted during the fine-tuning process. This approach is particularly important for large models, which typically contain hundreds of millions of parameters. Completely fine-tuning all these parameters is not only computationally expensive but can also lead to overfitting. PEFT allows the model to quickly adapt to new tasks by adjusting only a small subset of the model's parameters, while retaining the extensive knowledge learned during pre-training.
[0157] Low-Rank Adaptation (LoRA) is a typical PEFT method. Methods such as P-Tuning, P-Tuningv2, AdapterTuning and its variants, LoRA, AdaLoRA, QLoRA, MAMAdapter, and UniPELT are also important components of efficient parameter fine-tuning techniques. These methods each have their own characteristics and can be selected and adjusted according to specific tasks and datasets. The core idea of LoRA is to use a low-rank matrix to approximate the update of model weights. By introducing a low-rank matrix to approximate the update of model weights, it reduces the number of parameters. This method is particularly suitable for large models because these models contain a huge number of parameters. Fully fine-tuning all parameters is not only computationally expensive but also carries the risk of overfitting. LoRA, by adjusting only a small subset of the model's parameters, allows the model to quickly adapt to new tasks while retaining the rich knowledge learned during pre-training.
[0158] Efficient parameter fine-tuning techniques are one of the important research directions in the field of deep learning. By selecting appropriate fine-tuning methods and optimizing them in conjunction with specific tasks and datasets, the computational and storage costs of models can be greatly reduced while maintaining model performance.
[0159] S4. Building a multi-agent system based on general model capabilities:
[0160] Based on the general model obtained through pre-training and fine-tuning, a multi-agent system is further constructed, including agents for monitoring travel groups, querying full-chain travel trajectories, and querying tourist flow in scenic areas. These agents collaborate, sharing information and resources to optimize prediction results. A multi-agent system is a computational system composed of multiple interacting agents, designed to solve problems that are difficult for a single agent to handle through collaborative work. Multi-agent systems are an important branch of distributed artificial intelligence and represent a cutting-edge field of artificial intelligence internationally from the late 20th to the early 21st century.
[0161] Step S4 includes:
[0162] S41, Agent Architecture Design:
[0163] Before building a multi-agent system, the agent architecture must first be designed. This includes defining the role, responsibilities, and interaction methods of each agent. The agent architecture adopts a modular design for easy management and expansion. Each agent is designed to be autonomous, social, reactive, and proactive, capable of independently performing tasks and collaborating with other agents to provide more comprehensive services. Agents include agents for monitoring travel groups, querying full-chain travel trajectories, and querying tourist flow at scenic spots.
[0164] Each agent is responsible for different functions:
[0165] Intelligent agents for monitoring travel groups: monitoring tourists' travel patterns and behaviors.
[0166] Intelligent agent for tracking the entire travel trajectory: Tracking tourists' complete travel routes.
[0167] Intelligent tourist flow query tool: Provides real-time or historical tourist flow data query services.
[0168] S42, General Model Integration:
[0169] Based on the general models obtained through pre-training and fine-tuning, these models are integrated into various intelligent agents as their decision-making and inference engines. For example, a travel group monitoring agent may integrate a deep learning model to analyze tourist behavior patterns; a full-chain travel trajectory query agent may use a sequence prediction model to track tourist movement paths; and a scenic spot visitor flow query agent may integrate a time series analysis model to provide visitor flow data.
[0170] S43. Function Implementation and Optimization:
[0171] The implementation of each intelligent agent requires meticulous algorithm development and optimization. For example, a travel group monitoring agent might need to implement a complex event detection system to identify and respond to abnormal patterns in tourist behavior. A full-chain travel trajectory query agent needs to develop efficient path tracking algorithms to process and analyze large amounts of mobile data. A scenic area visitor flow query agent needs to implement a query optimization system to quickly respond to visitor flow data query requests.
[0172] S44. Collaborative Working Mechanism:
[0173] Collaborative work between intelligent agents is achieved through well-defined communication protocols and data sharing mechanisms. This involves using message queues, RESTful APIs, or other middleware technologies to facilitate information exchange between agents. Furthermore, a coordinator or scheduler is required to manage task allocation and resource scheduling for the agents, ensuring the overall efficiency and responsiveness of the system.
[0174] S5. Predictive analytics based on LLM-driven multi-agent model:
[0175] AI Large Language Models (LLMs) possess powerful natural language processing capabilities, enabling them to understand and handle the complexity of human language. In this system, the LLM acts as the brain, responsible for receiving input, interpreting intent, guiding decision-making, and integrating text for output. First, a large amount of historical visitor flow data is collected, including factors such as visitor numbers, time of day, weather conditions, and holidays. Then, through natural language processing techniques, this data is transformed into a machine-understandable form and input into a deep learning model. Through deep learning algorithms, the model learns from the historical behavior data of visitors to the scenic area, predicting potential visitor flow at different times and during different time periods, thus achieving trend analysis and prediction of visitor flow in the scenic area.
[0176] By integrating AI Large Language Model (LLM), LLM can understand and process natural language input, such as queries or instructions from scenic area administrators, driving multiple agents to achieve intelligent prediction and analysis of scenic area visitor flow trends.
[0177] Using Large Language Model (LLM) natural language processing technology, LLM can parse the natural language input of scenic area managers. Based on the parsing results, LLM drives collaborative work among multiple agents to achieve predictive analysis of visitor flow trends in the scenic area. LLM then outputs the predictive analysis results to the scenic area managers in an easy-to-understand text format, helping them make more accurate decisions.
[0178] Natural Language Processing (NLP) is a key technology in the field of artificial intelligence, enabling machines to understand and process human language. In this invention, NLP technology plays a crucial role. Through deep learning algorithms, these models can parse and understand large amounts of textual data, such as tourist reviews, social media posts, and news reports. This textual data contains rich information, such as tourist satisfaction, expectations of the scenic area, and their behavioral patterns. By analyzing this information, AI models can identify key factors influencing visitor flow, such as the attractiveness of the scenic area, service quality, and accessibility.
[0179] The following is the working principle and execution process of this method:
[0180] S51. Language Input and Intent Parsing:
[0181] Scenic area managers input queries or instructions into the system using natural language (such as text or voice). These inputs include requests for visitor flow forecasts for specific dates, time periods, or events. The LLM first receives the natural language input from the scenic area manager, and then automatically recognizes and parses the manager's intent to determine the type of task that needs to be performed.
[0182] Upon receiving input, LLM uses its advanced Natural Language Understanding (NLU) technology to analyze the administrator's query intent. The system identifies keywords, phrases, and contextual information to determine the administrator's specific needs, such as whether short-term forecasting, long-term trend analysis, or other specific types of analysis are required.
[0183] Based on the analysis results, LLM determines the specific types of tasks that need to be performed. These include short-term passenger flow forecasting (such as during holidays or special events), long-term trend analysis (such as seasonal changes or annual growth forecasting), and real-time data monitoring and abnormal traffic alerts.
[0184] S52, Agent Invocation and Analysis Execution:
[0185] Once the task type is determined, the LLM will drive the multi-agent system to work collaboratively based on the analysis results, invoking the corresponding sub-agents. Each sub-agent is specifically designed to handle a particular type of analysis task.
[0186] Intelligent agents collaborate, share information and resources, select the best database path, access and analyze historical data, real-time data and other relevant information stored in the knowledge base, and select appropriate prediction models to conduct tourist flow prediction analysis for scenic spots.
[0187] S53. Result Feedback and Text Output:
[0188] After completing the analysis, the sub-agent feeds back the prediction results to the LLM. The LLM then provides the predictive analysis results to the scenic area managers in natural language, enabling them to quickly grasp visitor flow trends and make corresponding plans and adjustments.
[0189] Through this LLM-based multi-agent approach, scenic area managers can more effectively respond to changes in visitor flow, optimize resource allocation, make advance plans and adjustments, enhance the visitor experience, and ultimately achieve efficient and sustainable development of scenic area operations.
[0190] Sub-agents typically refer to agents in a multi-agent system that undertake specific tasks or sub-tasks. They are the basic units constituting a multi-agent system and possess characteristics similar to the system itself, such as autonomy, responsiveness, and sociality. However, compared to a multi-agent system, the functions and task scope of sub-agents are more specific and limited. In a multi-agent system, sub-agents collaborate and coordinate to jointly accomplish complex tasks.
[0191] By following these steps, scenic area managers can more accurately predict visitor flow trends and take corresponding countermeasures in advance. This method not only improves the accuracy of predictions but also enhances the intelligence level of scenic area management, providing technical support for the sustainable development of the scenic area.
[0192] The AI-based large language model-based method for predicting and analyzing visitor flow trends in scenic areas effectively improves the intelligence level of scenic area management and provides a scientific basis for scenic area operation by comprehensively utilizing various advanced data analysis techniques and machine learning algorithms. In practical applications, this method can be effectively used to monitor visitor flow in real time and predict future visitor flow trends through predictive models. It can achieve the following benefits:
[0193] 1. Optimized visitor flow management. By utilizing AI-powered large language models, scenic area managers can monitor and predict visitor flow trends in real time, thereby enabling them to manage visitor flow in advance during holidays or special events, optimize resource allocation, effectively reduce congestion, and improve the visitor experience.
[0194] 2. More accurate resource allocation. By gaining a deeper understanding of tourist behavior patterns and preferences, scenic area managers can allocate resources such as security, services, and cleaning more precisely, ensuring sufficient human and material support at critical moments and improving the operational efficiency of the scenic area.
[0195] 3. Personalized Marketing Strategies. Based on the analysis of tourist data, scenic area managers can develop more targeted marketing strategies to attract more target tourists, enhance the scenic area's visibility and attractiveness, and strengthen its market competitiveness.
[0196] 4. Timely Risk Response. In the face of emergencies or unforeseen events, assist scenic area managers in adjusting management strategies in a timely manner, reducing risks, and ensuring tourist safety and order within the scenic area.
[0197] 5. Data-driven decision support. The constructed multi-dimensional knowledge base and intelligent toolset provide rich data support for scenic area managers, making the decision-making process more scientific and accurate, promoting the development of scenic area management towards a data-driven direction, and improving the overall management level.
[0198] By using AI-powered large language models to predict visitor flow trends, scenic area managers can forecast visitor volume during holidays or special events, optimizing resource allocation and reducing congestion and waste. For example, during holidays or special events, managers can use forecasts to manage visitor flow in advance, optimize resource allocation, and improve the visitor experience. This method also helps managers better understand visitor behavior patterns and preferences, enabling them to develop more targeted marketing strategies and improvement measures, thereby enhancing the overall attractiveness and competitiveness of the scenic area.
[0199] The above technical features constitute the preferred embodiment of the present invention, which has strong adaptability and optimal implementation effect. Non-essential technical features can be added or removed according to actual needs to meet the needs of different situations.
[0200] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention do not depart from the essence and scope of the technical solution of the present invention.
Claims
1. A method for predicting and analyzing tourist flow trends in scenic areas based on an AI-powered large language model, characterized in that, Includes the following steps: S1. Full Data Collection and Processing: Collect data from multiple sources related to visitor flow in the scenic area, including LBS data, scenic area data, holiday information, weather information, and third-party map data. Clean, integrate, and standardize this data. S2 includes: S21. Data Integration: Integrate the processed data, including LBS data, scenic spot data, holiday information, weather information, and third-party map data; S22. Feature Extraction: Extract key features and tourist behavior patterns from the integrated data. Key features include tourist behavior patterns, peak tourist flow periods, and the impact of weather on tourist flow. Use clustering algorithms to identify tourist behavior patterns, use time series analysis to predict peak tourist flow periods, and use regression models to assess the potential impact of weather changes on the number of tourists. S23. Rule and Pattern Recognition: Identify rules and patterns in data, including using hypothesis testing to determine the correlation between holidays and increased tourist traffic, and applying time series analysis to identify the impact of weather changes on tourist traffic; S24. Utilize RAG technology to build industry and scenario knowledge bases: Integrate extracted features, identified rules and patterns, and generated conclusions into a structured database to build a knowledge base; S3 includes: S31. Model pre-training: Using machine learning techniques, a general model is pre-trained based on multidimensional data in a knowledge base; S32. Model fine-tuning: Use multiple sources of data related to tourist flow in scenic areas to fine-tune the pre-trained model, including adjusting the output layer of the model to adapt to the output format of the scenic area tourist flow prediction model, travel chain analysis model, and tourist profile model. S4 includes: S41. Design of the intelligent agent architecture: This includes defining the role and responsibilities of each intelligent agent and the interaction methods between intelligent agents; the intelligent agents include a travel group monitoring intelligent agent, a full-chain travel trajectory query intelligent agent, and a scenic spot visitor flow query intelligent agent; S42. General Model Integration: Based on the general model obtained through pre-training and fine-tuning, the general model is integrated into various intelligent agents. This includes integrating a deep learning model into the travel group monitoring agent to analyze tourist behavior patterns; integrating a sequence prediction model into the full-chain travel trajectory query agent to track tourist movement paths; and integrating a time series analysis model into the scenic area visitor flow query agent to provide visitor flow data. S43. Function Implementation and Optimization: The intelligent agent for monitoring travel groups includes an event detection system for identifying and responding to abnormal patterns in tourist behavior; the intelligent agent for querying the entire travel trajectory includes a path tracking algorithm for processing and analyzing mobile data; the intelligent agent for querying scenic area visitor flow includes a query optimization system for responding to visitor flow data query requests. S44. Collaborative Working Mechanism: Intelligent agents communicate with each other through clearly defined communication protocols and data sharing mechanisms; S5 includes: S51, Language Input and Intent Resolution, including: S511. The scenic area administrator inputs queries or instructions into the system via natural language. The input includes requests for visitor flow prediction for specific dates, time periods, or events. The LLM first receives the natural language input from the scenic area administrator. The LLM automatically recognizes and parses the administrator's intent to determine the type of task to be performed. After receiving the input, the S512 and LLM use their Natural Language Understanding (NLU) technology to analyze the administrator's query intent. S513. Based on the analysis results, LLM determines the specific task types that need to be performed, including short-term passenger flow forecasting, long-term trend analysis, and real-time data monitoring and abnormal traffic warning. S52. Agent Invocation and Analysis Execution: After determining the task type, based on the parsing results, the LLM drives the multi-agent system to work collaboratively and invoke the corresponding sub-agents; the agents work collaboratively, share information and resources, select the best database path, access and analyze historical data, real-time data and other relevant information stored in the knowledge base, and select the corresponding prediction model to perform scenic area visitor flow prediction analysis. S53. Result Feedback and Text Output: After completing the analysis, the sub-agent will feed back the prediction results to the LLM, and the LLM will provide the prediction analysis results to the scenic area administrator in natural language.
2. The scenic area visitor flow trend prediction and analysis method according to claim 1, characterized in that, LBS data includes: user data, location data, timestamps, movement trajectories, geospatial data, points of interest data, user behavior data, and social network data.
3. The scenic area visitor flow trend prediction and analysis method according to claim 1, characterized in that, In step S1, data cleaning includes removing irrelevant characters, removing duplicates, removing noise, filling missing values, validating data types and formats, and checking for logical errors. Noise removal includes: identifying and cleaning up non-numeric values in the ticket sales data of the scenic area; logical error checking includes ensuring that the prices in the ticket sales data of the scenic area do not exceed the set upper limit for ticket sales prices.
4. The scenic area visitor flow trend prediction and analysis method according to claim 1, characterized in that, Step S21 specifically includes: S211. Read LBS data, holiday information, weather information, and third-party map data; S212. Perform data information and transformation operations, including removing invalid records from LBS data and converting timestamp formats; S213. Using timestamps and location information as key fields, connect multiple datasets of LBS data, holiday information, weather information, and third-party map data to form a unified data view; S214. Select the columns required for analysis from the concatenated dataset. These columns include timestamps, location information, holiday information, weather information, and third-party map data. Rename the selected columns to unify the format of the data view. S215. Store the formatted data in the knowledge base.
5. The scenic area visitor flow trend prediction and analysis method according to claim 1, characterized in that, In step S22, the use of clustering algorithms to identify tourist behavior patterns specifically includes: selecting location data, timestamps, and weather information as features, applying the K-means algorithm to cluster tourist behavior, and identifying different tourist behavior patterns.
6. The scenic area visitor flow trend prediction and analysis method according to claim 1, characterized in that, In step S23, hypothesis testing is used to determine the correlation between holidays and increased tourist traffic, including: Add holiday labels to the dates in the integrated dataset; Construct a linear regression model to assess the impact of holidays on tourist numbers; Output a statistical summary of the linear regression model.
Citation Information
Patent Citations
Artificial intelligence enterprise customer risk information analysis and evaluation method and system based on large language model
CN117726166A