Tourism comprehensive statistics big data platform based on fused multi-channel data

By constructing a comprehensive tourism statistics big data platform with multi-channel data, and combining tourist behavior hypergraph, inverse reinforcement learning, and bee swarm clustering algorithm, individualized travel paths and preference patterns are identified. This solves the problem that existing technologies cannot accurately capture tourists' personalized needs, and improves the efficiency of tourism plan optimization and resource allocation.

CN121119331AInactive Publication Date: 2025-12-12SANYA UNIVERSITY +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511272987.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-12-12
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN121119331A_ABST
    Figure CN121119331A_ABST
Patent Text Reader

Abstract

The invention discloses a tourism comprehensive statistics big data platform based on fused multi-channel data, and relates to the technical field of data analysis, and the platform comprises a tourism data collection unit which is used for extracting the behavior characteristics of tourists from multi-channel tourism data, and constructing a tourism behavior hypergraph through employing a hypergraph clustering technology; the tourism preference identification unit is used for introducing dynamic factors into the tourist behavior hypergraph to form an individualized tourism path set, identifying high-frequency route segments in the individualized tourism path set, and clustering tourist preference behavior modes based on an identification result; and the tourism income statistical unit is used for predicting tourism income in a future time period based on the tourist preference behavior mode. According to the method, the high-frequency route segment which is high in access frequency and meets the bearing condition is extracted, unloading optimization is carried out on the overload route in combination with the bee colony clustering algorithm, and the repeatability and distribution rationality of the travel route are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis technology, and more specifically, to a comprehensive tourism statistics big data platform based on the integration of multi-channel data. Background Technology

[0002] Tourism big data refers to a vast and complex collection of information gathered through multiple channels, encompassing tourist behavior, tourism projects, consumption, traffic flow, weather, and other data. It covers multiple dimensions of tourist data, including travel behavior, preferences, and spending. The main purpose of statistical analysis of tourism big data is to gain a deeper understanding of tourist behavior patterns, optimize resource allocation, improve operational efficiency and personalized services, thereby providing support for decision-making by scenic spots and related industries.

[0003] If tourism data statistics cannot incorporate traveler behavior for preference path analysis, it will be impossible to accurately capture tourists' personalized needs and preferences, overlooking their actual behavioral characteristics when choosing travel routes. Consequently, it will be impossible to provide tourists with more suitable recommendations and optimized travel plans. Furthermore, if tourism project cost forecasting cannot be combined with tourist behavior, it will not only lead to biased cost estimates, making it difficult to accurately reflect the actual spending and needs of different tourist groups, but also affect the budget allocation and resource scheduling of tourism projects, resulting in operational inefficiency or resource waste.

[0004] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention

[0005] To address the problems in related technologies, this invention proposes a comprehensive tourism statistics big data platform based on the integration of multi-channel data, in order to overcome the aforementioned technical problems existing in the current related technologies.

[0006] Therefore, the specific technical solution adopted by the present invention is as follows: A comprehensive tourism statistics big data platform based on the integration of multi-channel data, the platform includes: The tourism data collection unit is used to collect and integrate tourism data from multiple channels, extract tourists' behavioral characteristics from the multi-channel tourism data, and construct a tourist behavior hypergraph using hypergraph clustering technology. The tourism preference identification unit is used to introduce dynamic factors into the tourist behavior hypergraph to form an individualized tourism route set, identify high-frequency route segments in the individualized tourism route set, and cluster tourist preference behavior patterns based on the identification results. The tourism revenue statistics unit is used to predict tourism revenue for future periods based on tourist preference and behavior patterns. The tourism project renewal unit is used to compare and analyze future tourism revenue with historical investment costs, and to formulate tourism project renewal measures based on the results of the comparative analysis.

[0007] Preferably, the travel preference identification unit includes: The tourism route generation module takes the tourist behavior hypergraph as input, uses inverse reinforcement learning techniques to train and generate individualized tourism routes that meet environmental constraints, and summarizes them to obtain an individualized tourism route set. The high-frequency route analysis module is used to perform sequence analysis on all routes in the individualized tourism route set and to use the route connection algorithm to obtain the high-frequency route segments that frequently recur in the individualized tourism route set. The preference behavior clustering module is used to label high-frequency route segments as structural features in the individualized tourist route set, thereby obtaining an updated individualized tourist route set. It also clusters and updates high-frequency route segments with similar preferences in the individualized tourist route set, generating tourist preference behavior patterns.

[0008] Preferably, the nodes in the tourist behavior hypergraph represent tourist attraction points, and the hyperedges represent combinations of tourist behaviors. Environmental factors and congestion factors are added as additional attributes to obtain the tourist behavior hypergraph.

[0009] Preferably, the high-frequency route analysis module, when performing sequence analysis on all routes in the individualized tourism route set and using a route connection algorithm to obtain frequently recurring high-frequency route segments in the individualized tourism route set, includes: Calculate the minimum number of paths for the individualized tourism route set based on the frequency of visitor visits and the maximum carrying capacity of the tourist attractions. Calculate the savings value between any pair of tourist attractions based on the minimum number of paths, and select tourist attractions in descending order of the number of paths to connect them. If the selected tourist attraction is not currently connected, the selected tourist attraction will be updated. At the same time, during the path connection process, it will be checked whether the total number of tourists on the current path exceeds the preset capacity. If it does not exceed the preset capacity, the current connected path will be regarded as a frequently recurring high-frequency route segment. Otherwise, the next step will be executed. The bee colony clustering algorithm is used to unload the overloaded paths where the total number of tourists exceeds the preset capacity, and the unloaded paths are taken as high-frequency route segments that frequently recur.

[0010] Preferably, the overloaded paths where the total number of tourists exceeds the preset capacity are processed using a bee colony clustering algorithm, and the paths after the capacity unloading process are identified as frequently recurring high-frequency route segments, including: Individualized tourist routes are used as hives, tourist attractions are used as nectar sources, tourist visit frequency is used as nectar quantity, overloaded routes are marked as crowded hives, and bee colonies are assigned, with bee colony types including worker bees, scout bees, and queen bees. Worker bees release pheromones in the overloaded path. After the scout bees detect the pheromones, they form a temporary path cluster at the bottleneck point of the overloaded path. The scout bees enter the temporary path cluster first and begin to explore available alternative paths. Calculate the path space degree on available alternative routes, divert tourists to available alternative routes with low carrying capacity based on the path space degree, and recalculate the carrying capacity of available alternative routes. Based on the carrying capacity calculation results, the bee colony will release pheromones again, triggering a new round of exploration of available alternative paths, until the overloaded path is successfully unloaded and the iteration stops. After the iteration is completed, the available alternative paths that have been repeatedly explored by the scout bees are marked as high-frequency route segments.

[0011] Preferably, clustering updates of individualized travel routes focus on high-frequency route segments with similar preferences, generating tourist preference behavior patterns including: Each path in the individualized tourism route set is mapped to a discrete spatial point, and a complex topology is constructed to extract the continuous homology results of each path in spatial structure; The continuous homology result of the path in the spatial structure is mapped to the vectorized representation of the function space. The complex topology is hierarchically clustered using a hierarchical clustering algorithm to obtain the tensor factor representation of the high-frequency route segments of the hierarchical clustering result. Causal discovery algorithms are used to identify causal chains between high-frequency route segments and combinations of tourist behaviors. The results of hierarchical clustering, tensor factor representation and causal chain identification are integrated to generate tourist preference behavior patterns corresponding to the hierarchical clustering results.

[0012] Preferably, identifying the causal chain between high-frequency route segments and combinations of tourist behaviors using a causal discovery algorithm includes: From the combinations of tourist behaviors, we gradually search for tourist behaviors that have statistical dependence on high-frequency route segments, and construct a candidate set of potential nodes by taking tourist behaviors as potential nodes; Each potential node in the potential node candidate set is examined, and potential nodes that reduce overall dependency are removed. The remaining potential nodes in the potential node candidate set are taken as visitor preference behavior nodes. Construct directed edges from tourist preference behavior nodes to high-frequency route segments to form a causal chain between high-frequency route segments and tourist preference behavior nodes.

[0013] Preferably, the tourism revenue statistics unit includes: The prediction model building module is used to build tourism cost prediction models based on tourist preference behavior patterns. The sample transfer learning module is used to transfer the tourism cost prediction model to a new data environment using a time-domain adaptive algorithm, so as to obtain an updated tourism cost prediction model. The time series analysis module is used to analyze the time-series characteristics of tourism data from multiple channels and to predict tourism revenue for future periods using an updated tourism cost prediction model.

[0014] Preferably, the sample transfer learning module, when using a time-domain adaptive algorithm to transfer the tourism cost prediction model to a new data environment to obtain an updated tourism cost prediction model, includes: Multi-channel tourism data is divided into source domain samples and target domain samples, and a value function is constructed to match the verified source domain samples and target domain samples. Based on the sample matching structure, the maximum kernel mean difference index is formulated to select the optimal source domain sample as the migration candidate. The source domain sample and the target domain sample are weighted according to the distribution similarity between the optimal source domain sample and the target domain sample. We use weighted source domain samples and target domain samples to jointly train the tourism cost prediction model, and introduce a domain discriminator to calibrate the sample distribution. Compare the migration benefits of the tourism cost prediction models before and after migration in the target domain, adjust the parameters of the tourism cost prediction models based on the migration benefits, and retrain them.

[0015] Preferably, the formula for calculating the savings between the tourism project pairs is as follows: ; In the formula, S ij Indicates tourist attraction sites i The savings between tourism projects and tourist attractions; d 0i Indicates the route from the entrance to the tourist attractions. i The Euclidean distance; d 0j These represent the distance from the entrance to the tourist attractions. j The Euclidean distance; d ij Indicates tourist attraction sites i and tourist attractions j The direct distance between them; β Indicates the path reduction factor; f i Indicates tourist attraction sites i The average daily visitor frequency; f j Indicates tourist attraction sites j The average daily visitor frequency; max ( f k This indicates the highest frequency of visits among all tourist attractions. c This represents the traffic weighting coefficient; t iThis indicates that tourists are at tourist attractions. i Average stay time; t j This indicates that tourists are at tourist attractions. j Average stay time; T max This indicates the threshold for the maximum permissible difference in stay time within a scenic area; d Indicates the time coordination coefficient; α This indicates the weight of the basic savings item.

[0016] The beneficial effects of this invention are as follows: 1. This invention comprises a tourism data collection unit that integrates multi-channel data and constructs a tourist behavior hypergraph to capture the dynamic behavioral structure of tourists; a tourism preference identification unit that further introduces dynamic factors to generate individualized paths and identify high-frequency route segments, thereby clustering tourist preference patterns; a tourism revenue statistics unit that inputs the preference patterns into a tourism cost prediction model and outputs revenue prediction results for future time periods; and a tourism project update unit that intelligently identifies popular and unpopular project points and formulates optimization strategies based on a comparison of predicted revenue and historical costs. This not only improves prediction accuracy and resource allocation efficiency but also provides strong support for personalized services and maximizing project revenue.

[0017] 2. This invention combines inverse reinforcement learning algorithms to generate individualized tourism route sets that conform to environmental constraints. A high-frequency route analysis module performs sequence analysis on the route set, extracts high-frequency route segments with high access frequency and meeting carrying capacity conditions through a route connection algorithm, and optimizes overloaded routes by combining bee colony clustering algorithm to improve the repeatability and distribution rationality of tourism routes. Furthermore, it integrates clustering results, tensor factors, and causal structures to generate stable and highly interpretable tourist preference behavior patterns, thereby providing accurate population behavior profiles and high-value route basis for tourism recommendation, service scheduling, and project decision-making.

[0018] 3. This invention combines an updated model to predict future tourism revenue, improving the accuracy of tourism revenue forecasting. It also utilizes the maximum kernel mean difference index and weighted processing to match source and target domain samples, ensuring consistency in feature distribution between the source and target domains, thereby reducing potential errors during the migration process. Furthermore, it introduces a domain discriminator to further calibrate the sample distribution, enhancing the model's adaptability in the target domain. Ultimately, accurate revenue forecasting provides a reliable basis for optimizing scenic area management, resource allocation, and marketing strategies. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a principle block diagram of a tourism comprehensive statistical big data platform based on the integration of multi-channel data according to an embodiment of the present invention.

[0021] In the picture: 1. Tourism data collection unit; 2. Tourism preference identification unit; 201. Tourism route generation module; 202. High-frequency route analysis module; 203. Preference behavior clustering module; 3. Tourism revenue statistics unit; 301. Predictive model construction module; 302. Sample transfer learning module; 303. Time series analysis module; 4. Tourism project update unit; 5. Display unit. Detailed Implementation

[0022] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are generally used to represent similar components.

[0023] According to an embodiment of the present invention, a comprehensive tourism statistics big data platform based on the integration of multi-channel data is provided.

[0024] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, according to an embodiment of the present invention, a comprehensive tourism statistical big data platform based on the fusion of multi-channel data includes: Tourism data collection unit 1 is used to collect and integrate tourism data from multiple channels, extract tourists' behavioral characteristics from the multi-channel tourism data, and construct a tourist behavior hypergraph using hypergraph clustering technology.

[0025] In the tourist behavior hypergraph, nodes represent tourist attractions, and hyperedges represent combinations of tourist behaviors. Environmental factors and congestion factors are added as additional attributes to obtain the tourist behavior hypergraph.

[0026] It should be noted that the process of collecting and integrating tourism data from multiple channels, extracting tourist behavioral characteristics from this data, and constructing a tourist behavior hypergraph using hypergraph clustering technology includes: Data was collected simultaneously from multiple channels (scenic area or city management systems, consumption systems, statistical platforms, etc.). The main raw data obtained included: visitor volume, consumption, project access volume, and project construction cost for each tourist attraction. The different data sources were cleaned, time-aligned, and uniformly encoded in sequence. Based on timestamps, spatial locations, and project IDs, the data were merged into a multidimensional table to obtain a time-series indicator matrix for each tourist attraction.

[0027] Tourist behavioral features are extracted from a time-series index matrix to identify combinations of tourist behaviors, such as the attractions visited, the time spent at each attraction, and the amount spent, as well as the frequency of these behavioral combinations, during a single trip. Based on the idea of ​​trajectory reconstruction, the complete visit sequence for each tourist is extracted, resulting in a high-dimensional feature vector of tourist behavior.

[0028] After the behavioral features are prepared, the hypergraph construction phase begins. Each node represents a tourist attraction, and these nodes are not simply connected one-to-one, but rather through higher-order relationships formed by "traveler behavior combinations." Each tourist's behavior combination is represented by a hyperedge, connecting all the attractions visited by that tourist. For example, if tourist A visits nodes 1, 3, and 5 during a trip, a hyperedge connecting these nodes is generated in the hypergraph. For each hyperedge, not only is the set of nodes it connects recorded, but two key contextual attributes are also attached: environmental factors (such as the weather conditions and air quality index of the day) and congestion factors (such as the tourist density and traffic congestion level of each attraction during that time period). The final constructed tourist behavior hypergraph consists of nodes representing tourist attractions, hyperedges representing higher-order combinations of tourist behavior, and hyperedges carrying environmental and congestion attributes as contextual information.

[0029] The tourism preference identification unit 2 is used to introduce dynamic factors into the tourist behavior hypergraph to form an individualized tourism route set, identify high-frequency route segments in the individualized tourism route set, and cluster tourist preference behavior patterns based on the identification results.

[0030] Among them, the travel preference identification unit 2 includes: The tourism route generation module 201 is used to take the tourist behavior hypergraph as input, use inverse reinforcement learning technology to train and generate individualized tourism routes that meet environmental constraints, and summarize them to obtain an individualized tourism route set.

[0031] It should be noted that, using the tourist behavior hypergraph as input, inverse reinforcement learning techniques are used to train and generate individualized tourism routes that satisfy environmental constraints. The resulting set of individualized tourism routes includes: Each tourist's historical path is extracted as an expert trajectory, which is considered the optimal behavior driven by certain latent preferences. Each edge in the hypergraph is mapped to a sequence of state-action pairs. When constructing the reinforcement learning environment, the state is set to spatial location, time period, current attraction crowd level, etc., and the action pairs are the edges reachable from the current node. Using these expert trajectories as training input, an inverse reinforcement learning algorithm is used to infer the tourist's implicit reward function, enabling the learned policy to reproduce the expert path behavior in the environment.

[0032] Once the reward function and strategy are obtained, positive strategies can be applied to generate new personalized travel routes for different types of tourists (with different preference attributes) under the same environmental constraints (such as attraction capacity, time budget, and weather conditions). These routes simultaneously satisfy individual preferences and realistic constraints, and are efficient and feasible paths under the target environment. The personalized routes generated for all tourist types are aggregated to form a set of personalized travel routes.

[0033] Inverse reinforcement learning (IRL) is a learning method that infers the underlying reward function from expert behavior. Unlike traditional reinforcement learning's direct learning strategy, IRL focuses on finding the implicit motivations that explain expert behavior. In applications that use tourist behavior hypergraphs as input, IRL estimates the preference drivers for each behavior by learning historical tourist paths and generates feasible travel routes that conform to individual preferences, taking into account environmental constraints (such as time limits, crowding, and resource scarcity).

[0034] The high-frequency route analysis module 202 is used to perform sequence analysis on all routes in the individualized tourism route set and to use the route connection algorithm to obtain the high-frequency route segments that frequently recur in the individualized tourism route set.

[0035] It should be noted that by performing sequence analysis on individualized tourist route sets and combining it with path connection algorithms to extract frequently repeated high-frequency route segments, efficient and highly reusable connection patterns in tourist routes can be systematically identified. By considering tourist visit frequency and the maximum capacity of tourist attractions, the minimum number of paths required to meet load balancing conditions is first determined, providing a theoretical upper limit for subsequent path connections. Based on this, the savings value of all attraction pairs is calculated, and priority connection pairs are selected from high to low, gradually expanding the path structure. If the total number of tourists on the connected path does not exceed the capacity limit, it is marked as a high-frequency route segment, forming a high-value connection template; otherwise, it enters the swarm clustering stage, unloading the overloaded paths and reducing pressure through clustering and recombination, while retaining reusable path patterns.

[0036] When performing sequence analysis on individualized tourist route sets and using path connection algorithms to extract frequently recurring high-frequency route segments, these segments represent route combinations that tourists repeatedly choose and frequently visit among multiple individual routes, thus exhibiting high representativeness and stability. Identifying these segments helps reveal tourists' collective preference patterns, typical tour routes, and route traffic hotspots.

[0037] The high-frequency route analysis module 202 performs sequence analysis on all routes in the individualized tourism route set and uses a route connection algorithm to obtain frequently recurring high-frequency route segments in the individualized tourism route set, including: Calculate the minimum number of paths for the individualized tourism route set based on the frequency of visitor visits and the maximum carrying capacity of the tourist attractions.

[0038] It should be noted that the total number of visitors to each attraction is calculated, and then divided by the maximum capacity of that attraction to obtain the minimum number of paths required for each attraction. The minimum number of paths for the entire personalized tour route set is obtained by summing the required path numbers for all attractions.

[0039] Calculate the savings value between any pair of tourist attractions based on the minimum number of paths, and select tourist attractions in descending order of the number of paths to connect them. It should be noted that the formula for calculating the savings between pairs of tourist attractions is as follows: ; In the formula, S ij Indicates tourist attraction sites i The savings between tourism projects and tourist attractions; d 0i , d 0j These represent the distance from the entrance to the tourist attractions. i and tourist attractions j The Euclidean distance; d ij Indicates tourist attraction sites i and tourist attractions j The direct distance between them; β Indicates the path reduction factor; f i , f j These represent tourist attractions. i and tourist attractions j The average daily visitor frequency; max ( f k This indicates the highest frequency of visits among all tourist attractions. c This represents the traffic weighting coefficient;t i , t j These respectively represent tourists at tourist attraction sites. i and tourist attractions j Average stay time; T max This indicates the threshold for the maximum permissible difference in stay time within a scenic area; d Indicates the time coordination coefficient; α This indicates the weight of the basic savings item.

[0040] If the selected tourist attraction is not currently connected, the selected tourist attraction will be updated. At the same time, during the path connection process, it will be checked whether the total number of tourists on the current path exceeds the preset capacity. If it does not exceed the preset capacity, the current connected path will be regarded as a frequently recurring high-frequency route segment. Otherwise, the next step will be executed. The bee colony clustering algorithm is used to unload the overloaded paths where the total number of tourists exceeds the preset capacity, and the unloaded paths are taken as high-frequency route segments that frequently recur.

[0041] It should be noted that by using the bee colony clustering algorithm to unload the load on overloaded paths where the total number of tourists exceeds the preset capacity, the distribution of tourists on the path can be dynamically adjusted, alleviating the pressure on popular paths. Worker bees in the colony release pheromones to identify path bottlenecks, scout bees actively explore alternative paths and realize intelligent tourist diversion based on path spatial degree, and the queen bee regulates the overall path structure and marks the optimal alternative route as a high-frequency route segment. This achieves effective unloading of tourist load and optimization of path structure, improves the carrying capacity rationality and reuse value of the overall path set, avoids resource overload, improves the operational efficiency of the scenic area, and enhances the tourist experience.

[0042] Among these methods, the bee colony clustering algorithm is used to unload overloaded routes where the total number of tourists exceeds the preset capacity. These unloaded routes are then identified as frequently recurring high-frequency route segments, including: Individualized tourist routes are used as hives, tourist attractions are used as nectar sources, tourist visit frequency is used as nectar quantity, overloaded routes are marked as crowded hives, and bee colonies are assigned, with bee colony types including worker bees, scout bees, and queen bees. Worker bees release pheromones in the overloaded path. After the scout bees detect the pheromones, they form a temporary path cluster at the bottleneck point of the overloaded path. The scout bees enter the temporary path cluster first and begin to explore available alternative paths. Calculate the path space degree on available alternative routes, divert tourists to available alternative routes with low carrying capacity based on the path space degree, and recalculate the carrying capacity of available alternative routes. Based on the carrying capacity calculation results, the bee colony will release pheromones again, triggering a new round of exploration of available alternative paths, until the overloaded path is successfully unloaded and the iteration stops. After the iteration is completed, the available alternative paths that have been repeatedly explored by the scout bees are marked as high-frequency route segments.

[0043] It should be noted that in the bee colony type, the role of worker bees is to release pheromones at crowded locations in overloaded paths, identify potential bottleneck areas, and guide scout bees to pay attention to these hotspots. Based on the distribution of pheromones, scout bees actively search for feasible alternative paths around bottleneck points and establish temporary path clusters, while also probing the spatial degree of paths to determine whether diversion conditions are available. The queen bee is responsible for global scheduling and path update control, recording high-value paths that are repeatedly explored and selected during the continuous iteration of the bee colony, and driving the bee colony to converge toward the optimal alternative path during the iteration process.

[0044] The preference behavior clustering module 203 is used to label high-frequency route segments as structural features in the individualized tourist route set to obtain an updated individualized tourist route set. It then clusters and updates the high-frequency route segments with similar preferences in the individualized tourist route set to generate tourist preference behavior patterns.

[0045] Among these, clustering updates of high-frequency route segments with similar preferences for individualized travel routes generate tourist preference behavior patterns, including: Each path in the individualized tourism route set is mapped to a discrete spatial point, and a complex topology is constructed to extract the continuous homology results of each path in spatial structure; The continuous homology result of the path in the spatial structure is mapped to a vectorized representation of the function space. The complex topology is then hierarchically clustered using a hierarchical clustering algorithm to obtain the tensor factor representation of the high-frequency route segments in the hierarchical clustering result.

[0046] It should be noted that hierarchical clustering is an unsupervised learning method that progressively aggregates or partitions data based on the distance between samples. Its core idea is to construct a tree-like aggregation structure (such as a dendrogram) among the data. It is divided into two categories: agglomerative (bottom-up) and divisive (top-down), specifically including: Step 1: After mapping the persistent cohomology results of the path in the spatial structure to a vectorized representation of the function space, the cohomology features of each path are transformed into vector embeddings, such as by using persistent graphs or persistent images to represent them as tensors. Then, a similarity matrix is ​​constructed for these path tensors using distance metric functions (such as Euclidean distance, relevance distance, etc.). Step 2: Merge path tensor vectors layer by layer from bottom to top to gradually form a clustered tree structure between paths, and prune at different levels of the tree to extract path clusters that appear frequently in high-density areas as high-frequency route segments; Step 3: Extract the tensor embedding form of the frequent path clusters to form the tensor factor representation of the high-frequency path segments.

[0047] Causal discovery algorithms are used to identify causal chains between high-frequency route segments and combinations of tourist behaviors. The results of hierarchical clustering, tensor factor representation and causal chain identification are integrated to generate tourist preference behavior patterns corresponding to the hierarchical clustering results.

[0048] It should be noted that by using causal discovery algorithms to identify causal chains between high-frequency route segments and combinations of tourist behavior, and integrating hierarchical clustering results, tensor factor representations, and causal chain identification results, it is possible to reveal tourists' behavioral patterns in specific route selection and their causal relationships with high-frequency route segments. This integration effectively identifies the preferences of different tourist groups and provides scenic area managers with data-driven decision-making support to optimize tourist experience and route selection.

[0049] The process requires mapping each path cluster in the path clustering results to its corresponding tensor factor, and matching the tourist behavior features contained in these path clusters with the behavior nodes in the causal chain. This structural alignment establishes a causal correspondence between path clusters and tourist behavior, thereby analyzing the tourist preference behavior patterns reflected behind each path cluster. The integration process relies on similarity matching and causal node alignment, requiring unified encoding of structural features in different information spaces. For example, tensor factors can be dimensionality reduced or a common embedding space can be used to represent causal variables and path embeddings to avoid dimensionality conflicts or information incompatibility issues. At the same time, fusion criteria (such as minimum distance and maximum information entropy matching) can be set to ensure the robustness of the integration. The effect of the integration is that it can clearly define the preference drivers of tourist behavior at the path level, thereby supporting accurate segmentation of behavioral groups, preference prediction, and personalized recommendations. This makes high-frequency route segments not only reflect structural commonalities but also carry causal preference labels, improving behavioral interpretability and service decision-making efficiency.

[0050] Among them, the use of causal discovery algorithms to identify the causal chain between high-frequency route segments and combinations of tourist behavior includes: From the combinations of tourist behaviors, we gradually search for tourist behaviors that have statistical dependence on high-frequency route segments, and construct a candidate set of potential nodes by taking tourist behaviors as potential nodes; Each potential node in the potential node candidate set is examined, and potential nodes that reduce overall dependency are removed. The remaining potential nodes in the potential node candidate set are taken as visitor preference behavior nodes. Construct directed edges from tourist preference behavior nodes to high-frequency route segments to form a causal chain between high-frequency route segments and tourist preference behavior nodes.

[0051] It should be noted that by progressively identifying behavioral features that have statistical dependencies on high-frequency route segments from tourist behavior combinations and constructing a potential node candidate set, behaviors that can significantly improve overall dependency are selected as tourist preference behavior nodes. Then, by constructing directed edges from these preference behavior nodes to high-frequency route segments to form causal chains, causal modeling between tourist behavior and route selection can be achieved. This process avoids irrelevant or interfering behaviors from misleading the model structure, improves the explanatory power and causal validity of the model, and complements the hierarchical clustering results and tensor factor representation. It not only maintains the consistency between the path structure and spatial representation, but also enhances the causal traceability of the behavioral layer.

[0052] Tourism revenue statistics unit 3 is used to predict tourism revenue for future periods based on tourist preference behavior patterns.

[0053] Among them, tourism revenue statistics unit 3 includes: The prediction model building module 301 is used to build a tourism cost prediction model based on tourist preference behavior patterns.

[0054] It should be noted that the network architectures used to build tourism cost prediction models based on tourist preference behavior patterns include fully connected neural networks, convolutional neural networks, and recurrent neural networks. Fully connected neural networks perform non-linear mapping of input data through multiple fully connected layers, making them suitable for processing structured data and capturing potential relationships in tourist behavior patterns. Convolutional neural networks automatically extract local features from data through convolutional layers, making them suitable for processing image or similar grid-like data, and capable of extracting spatial or local features from tourist behavior patterns. Recurrent neural networks and their variant LSTM are suitable for processing time-series data and can predict future tourism costs by capturing the historical evolution of tourist preferences.

[0055] The sample transfer learning module 302 is used to transfer the tourism cost prediction model to a new data environment using a time-domain adaptive algorithm to obtain an updated tourism cost prediction model.

[0056] The sample transfer learning module 302, when using a time-domain adaptive algorithm to transfer the tourism cost prediction model to a new data environment and obtain an updated tourism cost prediction model, includes: Multi-channel tourism data is divided into source domain samples and target domain samples, and a value function is constructed to match the verified source domain samples and target domain samples.

[0057] It should be noted that multi-channel tourism data is divided according to region, time window, or tourist type label. Regions with abundant data, strong stability, and complete tourist behavior and income information are used as source domain samples, while regions with less data and requiring transfer learning support are used as target domain samples. The maximum mean difference is used to perform feature distribution alignment pre-validation on the source and target domain samples to ensure the transferability of statistical characteristics between domains. A value function is constructed to evaluate the matching quality and its contribution to the effectiveness of transfer prediction. The value function is used to select pairs in the source and target domain samples for subsequent transfer prediction training. The expression of the value function is: ; In the formula, V This represents the total value function value, used to evaluate the overall matching effect between samples from the source and target domains. The larger the value, the higher the matching quality. N Indicates the number of samples in the source domain; M Indicates the number of samples in the target domain; w ij Represents source domain samples x i With target domain samples y j The matching weights are assigned based on distributional similarity or the inverse of spatial distance; s ( x i , y j ) represents the source domain sample x i With target domain samples y j Similarity metrics (such as Euclidean similarity, cosine similarity, or a custom distance kernel function) represent the similarity between two samples in the tourist behavior feature space, path structure tensor representation, and income-related labels. x i Vectorized feature representation of source domain samples; y j This represents the vectorized feature representation of the target domain sample.

[0058] Based on the sample matching structure, the maximum kernel mean difference index is formulated to select the optimal source domain sample as the migration candidate. The source domain sample and the target domain sample are weighted according to the distribution similarity between the optimal source domain sample and the target domain sample.

[0059] It should be noted that the source domain samples and target domain samples are represented as feature vector sets respectively. The maximum kernel mean difference between each source domain sample subset and the overall target domain samples is calculated. By calculating the maximum kernel mean difference between multiple source domain subsets and the target domain, the source domain subset that minimizes the maximum kernel mean difference is selected as the optimal migration candidate. Weighting factors are assigned to each sample based on the similarity in distribution between the optimal source domain subset and the target domain. Each pair of samples is weighted and scored, and samples with high similarity are given higher weights.

[0060] We use weighted source domain samples and target domain samples to jointly train the tourism cost prediction model, and introduce a domain discriminator to calibrate the sample distribution. Compare the migration benefits of the tourism cost prediction models before and after migration in the target domain, adjust the parameters of the tourism cost prediction models based on the migration benefits, and retrain them.

[0061] The time series analysis module 303 is used to analyze the time series characteristics of multi-channel tourism data and use the updated tourism cost prediction model to predict tourism revenue for future periods.

[0062] Tourism Project Update Unit 4 is used to compare and analyze future tourism revenue with historical investment costs, and to formulate tourism project update measures based on the results of the comparative analysis.

[0063] It should be noted that comparing future tourism revenue with historical investment costs helps identify which tourism projects have achieved a high cost-return ratio, thus prioritizing their retention or optimization as popular projects. Projects with low cost-return ratios or persistent losses are identified as less popular and considered for elimination, downsizing, or renovation. Measures that can be taken in this process include: For popular attractions, increase publicity and resource investment, expand service content to enhance added value, and optimize queuing and service processes to improve reception capacity. For less popular attractions, functional transformation, introduction of special activities to activate traffic, and temporary closure of inefficient areas can save operating costs, ultimately achieving dynamic optimization of scenic area resources and maximizing revenue.

[0064] In addition, the present invention also includes a display unit 5. The function of the display unit 5 is to display the analysis results, prediction data and optimization suggestions output by each unit to users or managers in a visual form, so that they can intuitively understand tourist preference behavior patterns, high-frequency route segments, tourism revenue prediction and project update decision content, thereby assisting scenic area managers in real-time monitoring, scientific judgment and strategy adjustment, and improving decision-making efficiency and system interactivity.

[0065] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A comprehensive tourism statistical big data platform based on the integration of multi-channel data, characterized in that, The platform includes: The tourism data collection unit is used to collect and integrate tourism data from multiple channels, extract tourists' behavioral characteristics from the multi-channel tourism data, and construct a tourist behavior hypergraph using hypergraph clustering technology. In the tourist behavior hypergraph, the nodes represent tourist attractions, the hyperedges represent combinations of tourist behaviors, and environmental factors and congestion factors are added as additional attributes to obtain the tourist behavior hypergraph. The tourism preference identification unit is used to introduce dynamic factors into the tourist behavior hypergraph to form an individualized tourism route set, identify high-frequency route segments in the individualized tourism route set, and cluster tourist preference behavior patterns based on the identification results. The tourism revenue statistics unit is used to predict tourism revenue for future periods based on tourist preference and behavior patterns. The tourism project renewal unit is used to compare and analyze future tourism revenue with historical investment costs, and to formulate tourism project renewal measures based on the results of the comparative analysis. The travel preference identification unit includes: The tourism route generation module takes the tourist behavior hypergraph as input, uses inverse reinforcement learning techniques to train and generate individualized tourism routes that meet environmental constraints, and summarizes them to obtain an individualized tourism route set. The high-frequency route analysis module is used to perform sequence analysis on all routes in the individualized tourism route set and to use the route connection algorithm to obtain the high-frequency route segments that frequently recur in the individualized tourism route set.

2. The tourism comprehensive statistical big data platform based on the integration of multi-channel data as described in claim 1, characterized in that, The travel preference identification unit also includes: The preference behavior clustering module is used to label high-frequency route segments as structural features in the individualized tourist route set, thereby obtaining an updated individualized tourist route set. It also clusters and updates high-frequency route segments with similar preferences in the individualized tourist route set, generating tourist preference behavior patterns.

3. The tourism comprehensive statistical big data platform based on the integration of multi-channel data as described in claim 1, characterized in that, The tourism route generation module, when taking a tourist behavior hypergraph as input and using inverse reinforcement learning techniques to train and generate individualized tourism routes that meet environmental constraints, and then summarizing them to obtain a set of individualized tourism routes, includes: Transform each hyperedge in the tourist behavior hypergraph into a sequence of state-action pairs; Define the state space and action space of the inverse reinforcement learning algorithm, and use the state-action pair sequence as the input of the inverse reinforcement learning algorithm to output the reward function; By combining the reward function with predefined environmental constraints, a tourist preference execution strategy is generated. This strategy is then used to generate several individualized travel paths for tourists. Finally, these individualized travel paths are aggregated to obtain an individualized travel path set.

4. The tourism comprehensive statistical big data platform based on the integration of multi-channel data as described in claim 1, characterized in that, The high-frequency route analysis module performs sequence analysis on all routes in the individualized travel route set and uses a route connection algorithm to obtain frequently recurring high-frequency route segments in the individualized travel route set, including: Calculate the minimum number of paths for the individualized tourism route set based on the frequency of visitor visits and the maximum carrying capacity of the tourist attractions. Calculate the savings value between any pair of tourist attractions based on the minimum number of paths, and select tourist attractions in descending order of the number of paths to connect them. If the selected tourist attraction is not currently connected, the selected tourist attraction will be updated. At the same time, during the path connection process, it will be checked whether the total number of tourists on the current path exceeds the preset capacity. If it does not exceed the preset capacity, the current connected path will be regarded as a frequently recurring high-frequency route segment. Otherwise, the next step will be executed. The bee colony clustering algorithm is used to unload the overloaded paths where the total number of tourists exceeds the preset capacity, and the unloaded paths are taken as high-frequency route segments that frequently recur.

5. A tourism comprehensive statistical big data platform based on the integration of multi-channel data as described in claim 4, characterized in that, The method of using bee colony clustering algorithm to unload overloaded paths where the total number of tourists exceeds the preset capacity, and then designating the unloaded paths as frequently recurring high-frequency route segments, includes: Individualized tourist routes are used as hives, tourist attractions are used as nectar sources, tourist visit frequency is used as nectar quantity, overloaded routes are marked as crowded hives, and bee colonies are assigned, with bee colony types including worker bees, scout bees, and queen bees. Worker bees release pheromones in the overloaded path. After the scout bees detect the pheromones, they form a temporary path cluster at the bottleneck point of the overloaded path. The scout bees enter the temporary path cluster first and begin to explore available alternative paths. Calculate the path space degree on available alternative routes, divert tourists to available alternative routes with low carrying capacity based on the path space degree, and recalculate the carrying capacity of available alternative routes. Based on the carrying capacity calculation results, the bee colony will release pheromones again, triggering a new round of exploration of available alternative paths, until the overloaded path is successfully unloaded and the iteration stops. After the iteration is completed, the available alternative paths that have been repeatedly explored by the scout bees are marked as high-frequency route segments.

6. A tourism comprehensive statistical big data platform based on the integration of multi-channel data as described in claim 2, characterized in that, The clustering update of individualized travel routes focuses on high-frequency route segments with similar preferences, generating tourist preference behavior patterns including: Each path in the individualized tourism route set is mapped to a discrete spatial point, and a complex topology is constructed to extract the continuous homology results of each path in spatial structure; The continuous homology result of the path in the spatial structure is mapped to the vectorized representation of the function space. The complex topology is hierarchically clustered using a hierarchical clustering algorithm to obtain the tensor factor representation of the high-frequency route segments of the hierarchical clustering result. Causal discovery algorithms are used to identify causal chains between high-frequency route segments and combinations of tourist behaviors. The results of hierarchical clustering, tensor factor representation and causal chain identification are integrated to generate tourist preference behavior patterns corresponding to the hierarchical clustering results.

7. A tourism comprehensive statistical big data platform based on the integration of multi-channel data as described in claim 6, characterized in that, The method of using causal discovery algorithms to identify the causal chain between high-frequency route segments and combinations of tourist behavior includes: From the combinations of tourist behaviors, we gradually search for tourist behaviors that have statistical dependence on high-frequency route segments, and construct a candidate set of potential nodes by taking tourist behaviors as potential nodes; Each potential node in the potential node candidate set is examined, and potential nodes that reduce overall dependency are removed. The remaining potential nodes in the potential node candidate set are taken as visitor preference behavior nodes. Construct directed edges from tourist preference behavior nodes to high-frequency route segments to form a causal chain between high-frequency route segments and tourist preference behavior nodes.

8. A tourism comprehensive statistical big data platform based on the integration of multi-channel data as described in claim 1, characterized in that, The tourism revenue statistics unit includes: The prediction model building module is used to build tourism cost prediction models based on tourist preference behavior patterns. The sample transfer learning module is used to transfer the tourism cost prediction model to a new data environment using a time-domain adaptive algorithm, so as to obtain an updated tourism cost prediction model. The time series analysis module is used to analyze the time-series characteristics of tourism data from multiple channels and to predict tourism revenue for future periods using an updated tourism cost prediction model.

9. A tourism comprehensive statistical big data platform based on the integration of multi-channel data as described in claim 8, characterized in that, The sample transfer learning module, when using a time-domain adaptive algorithm to transfer the tourism cost prediction model to a new data environment to obtain an updated tourism cost prediction model, includes: Multi-channel tourism data is divided into source domain samples and target domain samples, and a value function is constructed to match the verified source domain samples and target domain samples. Based on the sample matching structure, the maximum kernel mean difference index is formulated to select the optimal source domain sample as the migration candidate. The source domain sample and the target domain sample are weighted according to the distribution similarity between the optimal source domain sample and the target domain sample. We use weighted source domain samples and target domain samples to jointly train the tourism cost prediction model, and introduce a domain discriminator to calibrate the sample distribution. Compare the migration benefits of the tourism cost prediction models before and after migration in the target domain, adjust the parameters of the tourism cost prediction models based on the migration benefits, and retrain them.

10. A tourism comprehensive statistical big data platform based on the fusion of multi-channel data as described in claim 4, characterized in that, The formula for calculating the savings between the tourism project pairs is as follows: ; In the formula, S ij Indicates tourist attraction sites i The savings between tourism projects and tourist attractions; d 0i Indicates the route from the entrance to the tourist attractions. i The Euclidean distance; d 0j These represent the distance from the entrance to the tourist attractions. j The Euclidean distance; d ij Indicates tourist attraction sites i and tourist attractions j The direct distance between them; β Indicates the path reduction factor; f i Indicates tourist attraction sites i The average daily visitor frequency; f j Indicates tourist attraction sites j The average daily visitor frequency; max ( f k This indicates the highest frequency of visits among all tourist attractions. γ This represents the traffic weighting coefficient; t i This indicates that tourists are at tourist attractions. i Average stay time; t j This indicates that tourists are at tourist attractions. j Average stay time; T max This indicates the threshold for the maximum permissible difference in stay time within a scenic area; δ Indicates the time coordination coefficient; α This indicates the weight of the basic savings item.