Big Data-Driven Full-Process Tourism Industry Data Center Management System
Patent Information
- Application Number
- CN202610799576.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-01
AI Technical Summary
现有旅游数据管理体系缺乏统一的数据汇聚机制,无法同步整合多业态业务系统产生的原始数据
[0048]This method employs dynamic time-warping distance to calculate the similarity between different tourist behavior sequences, changing the calculation logic of traditional fixed-measure methods and adapting to temporal misalignment and duration differences in behavior sequences. Based on tourist identification fields, it performs behavior trajectory aggregation processing, merging and integrating similar behaviors according to the inherent variation characteristics of the sequences. This ensures that the aggregated collection of individual tourist behavior trajectory chains possesses complete temporal logic, achieving the systematic aggregation of scattered behavior records and providing well-organized trajectory data support for subsequent multi-dimensional analysis.
Smart Images

Figure CN122674984A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tourism data management technology, and in particular to a data center management system for the entire tourism industry based on big data. Background Technology
[0002] The tourism industry encompasses multiple business segments, including scenic spots, hotels, transportation, catering, and travel agencies. Each segment's business systems are deployed independently, with data scattered across different platforms. The existing tourism data management system lacks a unified data aggregation mechanism, making it impossible to synchronously integrate raw data generated by multiple business systems. The various data types are heterogeneous, with inconsistent formats and field definitions. Without systematic cleaning and standardization, it is difficult to build a unified and usable basic data warehouse for the entire tourism industry.
[0003] Current methods for aggregating tourist behavior mostly employ conventional clustering algorithms, using fixed metrics to calculate behavioral sequence similarity. These methods fail to adapt to the temporal fluctuations in tourist behavior, making it difficult to accurately aggregate and categorize behavioral trajectories based on identity identifiers, and unable to generate complete and continuous chains of individual tourist behavior trajectories. Furthermore, tourism industry data analysis often adopts a single-dimensional statistical model, lacking a multi-granular, hierarchical data organization architecture, and thus failing to integrate and correlate tourist trajectory information with various industry-specific dimensional fields.
[0004] Traditional data analysis architectures lack multi-dimensional online analytical processing capabilities, only outputting basic statistical results. They cannot break down tourism business indicators from multiple perspectives, nor can they generate standardized, comprehensive data management views. This fragmented and isolated data processing and analysis model makes it difficult to achieve unified management and control of tourism data across all sectors, in-depth aggregation of tourist behavior, and multi-dimensional indicator linkage analysis. It also fails to meet the application requirements of intensive, systematic, and multi-dimensional data center management for tourism in a big data environment. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a big data-based data center management system for the entire tourism industry.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a big data-based full-process tourism industry data center management system, comprising:
[0007] The data aggregation module acquires a collection of original data from multiple sources across the entire tourism industry. This collection includes data from scenic spot ticketing systems, hotel booking systems, transportation ticketing systems, catering consumption systems, and travel agency group systems.
[0008] The data processing module performs heterogeneous data cleaning and standardization processing on the multi-source tourism industry raw data set to generate a standardized tourism industry basic data warehouse.
[0009] The trajectory generation module calls an improved clustering algorithm to aggregate tourist behavior trajectories in the tourist identity field of the tourism industry basic data warehouse, generating a set of individual tourist behavior trajectory chains. The improved clustering algorithm calculates the similarity between different tourist behavior sequences based on dynamic time curvature distance.
[0010] The model building module performs multi-granularity data cube construction processing based on the set of individual tourist behavior trajectory chains and the industry dimension fields of the basic data warehouse of the entire tourism industry, generating a multi-dimensional data model of the entire tourism industry.
[0011] The analysis and presentation module performs online analysis and processing of all-area tourism indicators based on the multi-dimensional data model of the entire tourism industry, and generates a data management view of the entire tourism industry.
[0012] As a further aspect of the present invention, the step of performing heterogeneous data cleaning and standardization processing on the multi-source tourism industry-wide raw data set to generate a standardized tourism industry-wide basic data warehouse specifically includes:
[0013] For each data source in the multi-source tourism industry original data set, perform field semantic parsing processing to identify the tourist identity field, timestamp field, geographical location field, and consumption amount field in each data source;
[0014] Based on the tourist identity identifier field, cross-data source association matching is performed on the records in the multi-source tourism industry original data set, and records from different data sources belonging to the same tourist are marked as the same tourist identity group.
[0015] For all records within each tourist identity group, a unified time format conversion process is performed on the timestamp field to convert the time representation format of different data sources into a preset unified time encoding format;
[0016] For all records within each tourist identity group, perform geocoding standardization processing on the geographic location field to convert geographic location descriptions in different formats into a unified geographic coordinate code;
[0017] For all records within each tourist identity group, currency unit normalization and exchange rate conversion are performed on the consumption amount field to generate a standard amount value;
[0018] All records, after undergoing unified time format conversion, geocoding standardization, and currency unit normalization, are grouped and merged according to tourist identity for storage, generating the standardized basic data warehouse for the entire tourism industry.
[0019] As a further aspect of the present invention, the geocoding standardization process calls a third-party geographic information service to convert the geographic location described in the text into latitude and longitude coordinates.
[0020] As a further aspect of the present invention, the step of using an improved clustering algorithm to aggregate tourist behavior trajectories in the tourist identity field of the tourism industry-wide basic data warehouse to generate a set of individual tourist behavior trajectory chains specifically includes:
[0021] Extract all records corresponding to each tourist identification field from the basic data warehouse of the entire tourism industry, sort and concatenate the geographical location fields of each record according to the order of the timestamp fields in each record, and generate the original behavior location sequence corresponding to each tourist identification field.
[0022] For each of the original behavioral location sequences, a stop point identification process is performed. Adjacent records whose geographical location changes are less than a preset distance threshold within a continuous time window are merged into a stop point entity to obtain the stop point sequence corresponding to each tourist identity field.
[0023] Calculate the dynamic time warp distance between the stop point sequences corresponding to every two tourist identification fields, and use the dynamic time warp distance as the behavioral trajectory similarity between the two tourist identification fields.
[0024] Based on the behavioral trajectory similarity, a tourist similarity matrix is constructed, and spectral clustering segmentation is performed on the tourist similarity matrix to obtain multiple tourist behavioral trajectory clusters;
[0025] The sequence alignment and fusion of the stop point sequences corresponding to all tourist identity fields within each tourist behavior trajectory cluster are performed to generate a representative behavior trajectory chain for each tourist behavior trajectory cluster. The set of all representative behavior trajectory chains is taken as the set of individual tourist behavior trajectory chains.
[0026] As a further aspect of the present invention, the preset distance threshold used in the stop point identification process is adaptively adjusted according to the geographic spatial scale of the specific scenic area.
[0027] As a further aspect of the present invention, the improved clustering algorithm calculates the similarity between different tourist behavior sequences based on dynamic time curvature distance. The specific working steps of the improved clustering algorithm include:
[0028] Obtain the first stop sequence of the first tourist and the second stop sequence of the second tourist, wherein the first stop sequence contains a first sequence length of stop entities and the second stop sequence contains a second sequence length of stop entities;
[0029] Construct an initial distance matrix between the first stop point sequence and the second stop point sequence, wherein the number of rows in the initial distance matrix is the length of the first sequence and the number of columns in the initial distance matrix is the length of the second sequence;
[0030] Traverse each matrix position in the initial distance matrix, calculate the spatial Euclidean distance between the stop point entities in the first stop point sequence and the stop point entities in the second stop point sequence corresponding to the matrix position, and fill the spatial Euclidean distance into the corresponding position of the initial distance matrix to obtain the filled distance matrix;
[0031] Perform dynamic programming cumulative distance calculation on the filled distance matrix. Starting from the top left corner of the filled distance matrix, calculate the cumulative distance value row by row and column by column. The cumulative distance value is calculated by adding the minimum value among the spatial Euclidean distance of the current position, the left position, the top position, and the top left corner position to obtain the cumulative distance matrix.
[0032] The cumulative distance value at the bottom right corner of the cumulative distance matrix is used as the dynamic time curvature distance between the first stop sequence and the second stop sequence, and this dynamic time curvature distance is output as a similarity metric between the two tourist behavior sequences.
[0033] As a further aspect of the present invention, the coordinates used to calculate the spatial Euclidean distance are geographic latitude and longitude coordinates based on the WGS-84 coordinate system.
[0034] As a further aspect of the present invention, the step of constructing a multi-granularity data cube based on the set of individual tourist behavior trajectory chains and the industry-specific dimension field of the basic data warehouse for the entire tourism industry, to generate a multi-dimensional data model of the entire tourism industry, specifically includes:
[0035] Extract a set of industry-specific dimension fields from the basic data warehouse of the entire tourism industry. The set of industry-specific dimension fields includes time dimension fields, geographic region dimension fields, tourist profile dimension fields, tourism industry type dimension fields, and consumption level dimension fields.
[0036] Each behavior trajectory chain in the set of individual tourist behavior trajectory chains is used as a fact record. The fact record is associated and bound with each dimension field in the set of business format dimension fields to obtain the fact record value corresponding to each dimension field.
[0037] A time hierarchy structure is established based on the time dimension field, which includes year, quarter, month, day, and hour levels.
[0038] A spatial hierarchy structure is established based on the geographic region dimension field. The spatial hierarchy structure includes national level, provincial level, city level, scenic area level, and attraction level.
[0039] The time hierarchy, spatial hierarchy, tourist profile dimension field, tourism industry type dimension field, and consumption level dimension field are used as the dimension axes of the multidimensional data model, and the fact records are used as the fact measurement values of the multidimensional data model to construct the full-process tourism full-industry multidimensional data model.
[0040] As a further aspect of the present invention, the step of performing online analysis and processing of all-area tourism indicators based on the multi-dimensional data model of the entire tourism industry to generate a data management view of the entire tourism industry specifically includes:
[0041] The system receives a combination of analytical dimensions and a type of analytical indicator specified by the user. The combination of analytical dimensions includes a time dimension level selected from the time hierarchy, a spatial dimension level selected from the spatial hierarchy, and a profile slice condition selected from the visitor profile dimension field.
[0042] Based on the combination of analytical dimensions, perform a dimension roll-up or dimension drill-down operation on the multi-dimensional data model of the entire tourism industry to generate an aggregated data cube slice corresponding to the combination of analytical dimensions.
[0043] Based on the analysis indicator type, the aggregated data cube slices are subjected to indicator calculation operations, including tourist visit frequency statistics, average tourist stay duration calculation, tourist spending summary, and tourism industry conversion rate calculation.
[0044] The results of the indicator calculation operations are sorted and grouped according to the dimension hierarchy in the analysis dimension combination to generate a structured analysis result table.
[0045] The structured analysis results table is converted into a visual chart object, and all the visual chart objects are arranged according to the preset management view template to generate the tourism industry data management view.
[0046] As a further embodiment of the present invention, the visualization chart objects include heatmaps, trajectory flow charts, stacked bar charts, and trend line charts.
[0047] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0048] This method employs dynamic time-warping distance to calculate the similarity between different tourist behavior sequences, changing the calculation logic of traditional fixed-measure methods and adapting to temporal misalignment and duration differences in behavior sequences. Based on tourist identification fields, it performs behavior trajectory aggregation processing, merging and integrating similar behaviors according to the inherent variation characteristics of the sequences. This ensures that the aggregated collection of individual tourist behavior trajectory chains possesses complete temporal logic, achieving the systematic aggregation of scattered behavior records and providing well-organized trajectory data support for subsequent multi-dimensional analysis.
[0049] Based on the collection of individual tourist behavior trajectory chains and the industry-specific dimension fields in the basic data warehouse of the entire tourism industry, a multi-granularity data cube is constructed. Tourism data is then hierarchically divided and structured according to different levels and dimensions. A correlation is established between tourist trajectory information and the business attribute fields of each industry, building a multi-level, multi-dimensional data organization architecture to form a complete, multi-dimensional data model of the entire tourism industry.
[0050] Based on the established multidimensional data model, online analysis and processing of tourism indicators across the entire region are carried out. This supports the layer-by-layer decomposition and correlation query of data at different dimensions and levels, and establishes the inherent relationships between data from multiple business sectors. Various statistical indicators are integrated according to business analysis needs to generate a standardized tourism industry-wide data management view, achieving integrated presentation of multi-business data and meeting the needs of the tourism industry for multi-level, multi-perspective data query, indicator statistics, and comprehensive management. Attached Figure Description
[0051] Figure 1 This is a sequence diagram of the big data-based full-process tourism industry data center management system described in this invention;
[0052] Figure 2 A workflow diagram for cleaning and standardizing heterogeneous data to generate a basic data warehouse for the entire tourism industry;
[0053] Figure 3 A flowchart illustrating the workflow for generating individual behavior trajectory chains by aggregating tourist behavior trajectories based on an improved clustering algorithm. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0055] See Figure 1 This invention provides a big data-based data center management system for the entire tourism industry, specifically including:
[0056] The data aggregation module acquires a multi-source raw data set covering the entire tourism industry, including data from scenic spot ticketing systems, hotel booking systems, transportation ticketing systems, catering consumption systems, and travel agency group systems. The data processing module performs heterogeneous data cleaning and standardization on this raw data set, generating a standardized basic data warehouse for the entire tourism industry. The trajectory generation module uses an improved clustering algorithm to aggregate tourist behavior trajectories from the tourist identification field in the basic data warehouse, generating a set of individual tourist behavior trajectory chains. The improved clustering algorithm calculates the similarity between different tourist behavior sequences based on dynamic time curvature distance. The model building module constructs multi-granularity data cubes based on the individual tourist behavior trajectory chains and the industry dimension fields of the basic data warehouse, generating a multi-dimensional data model for the entire tourism industry. The analysis and presentation module performs online analysis of all-area tourism indicators based on the multi-dimensional data model, generating a data management view for the entire tourism industry.
[0057] In one embodiment of the present invention, see [reference] Figure 2 The data processing module performs semantic parsing on each data source in the multi-source tourism industry's original data set, identifying the tourist identity field, timestamp field, geographic location field, and consumption amount field in each data source. Based on the tourist identity field, it performs cross-data source association matching on the records in the multi-source tourism industry's original data set, marking records belonging to the same tourist from different data sources as the same tourist identity group. For all records within each tourist identity group, it performs a unified time format conversion on the timestamp field, converting the time representation formats of different data sources into a preset unified time encoding format. For all records within each tourist identity group, it performs geocoding standardization on the geographic location field, converting different formats of geographic location descriptions into a unified geographic coordinate code. During this geocoding standardization process, it calls a third-party geographic information service to convert the textual description of the geographic location into latitude and longitude coordinates. For all records within each tourist identity group, it performs currency unit normalization and exchange rate conversion on the consumption amount field, generating a standard amount value. All records after time format normalization, geocoding standardization, and currency unit normalization are merged and stored according to tourist identity groups, generating a standardized basic data warehouse for the entire tourism industry.
[0058] In practical implementation, the data aggregation module acquires a multi-source, full-format tourism raw data set for a provincial-level tourist destination. This set covers data from scenic spot ticketing systems, hotel booking systems, transportation ticketing systems, catering consumption systems, and travel agency group systems. The data processing module performs semantic parsing on each data source in the multi-source tourism raw data set, identifying the tourist identification field, timestamp field, geographic location field, and consumption amount field in each data source. In the scenic spot ticketing system data, the tourist identification field is identified as the ID card number field, the timestamp field as the entry time field, the geographic location field as the scenic spot name field, and the consumption amount field as the ticket price field. In the hotel booking system data, the tourist identification field is identified as the guest ID number field. The following data was misinterpreted: timestamp field was identified as check-in date field; geolocation field was identified as hotel address field; and consumption amount field was identified as room fee field. In the transportation ticketing system data, the tourist identity field was identified as passenger ID card field; timestamp field was identified as ticket purchase time field; geolocation field was identified as departure station field; and consumption amount field was identified as ticket price field. In the catering consumption system data, the tourist identity field was identified as member identity field; timestamp field was identified as consumption time field; geolocation field was identified as restaurant name field; and consumption amount field was identified as actual payment amount field. In the travel agency group system data, the tourist identity field was identified as group member ID number field; timestamp field was identified as itinerary date field; geolocation field was identified as tourist attraction field; and consumption amount field was identified as tour fee field.
[0059] In practice, the data processing module performs cross-data source association and matching on records in the multi-source tourism industry's original data set based on the tourist identification field. Records with the same tourist identification field value from different data sources are marked as the same tourist identity group. For example, a tourist with ID number "3401021990xxxx" has records in the scenic spot ticketing system, hotel booking system, and transportation ticketing system; these three records are grouped into the same tourist identity group. For all records within each tourist identity group, the data processing module processes the timestamp... The time format of the segment is uniformly converted. The timestamp values "2026-05-1409:30:25" in the scenic spot ticketing system data, "05 / 14 / 202614:00" in the hotel booking system data, and "20260514T08:15" in the transportation ticketing system data are uniformly converted into a preset unified time encoding format. The unified time encoding format is the Coordinated Universal Time string in ISO8601 extended format. The converted timestamp value is "2026-05-14T09:30:25+08:00".
[0060] In practice, the data processing module performs geocoding standardization on the geographic location field for all records within each tourist identity group. It calls third-party geographic information services to convert the scenic spot name text from the scenic spot ticketing system, the hotel address text from the hotel booking system, the departure station text from the transportation ticketing system, and the restaurant name text from the catering system into geographic latitude and longitude coordinates based on the WGS-84 coordinate system. For example, "Guangmingding" in the scenic spot ticketing system data is converted to longitude 118.182456 and latitude 30.141789, and so on. The coordinates of "No. 191 Huangshan Road, Shushan District, Hefei City" were converted to longitude 117.230456 and latitude 31.825678. The coordinates of "Hefei South Railway Station" in the transportation ticketing system were converted to longitude 117.286530 and latitude 31.815892. The coordinates of "Huishang Hometown Restaurant" in the catering consumption system were converted to longitude 117.250321 and latitude 31.861234. The coordinates of "Huangshan Scenic Area" in the travel agency group system were converted to longitude 118.167324 and latitude 30.126789. The converted longitude and latitude coordinates were stored as standard geographic coordinate codes.
[0061] In practice, the data processing module performs currency unit normalization and exchange rate conversion on the consumption amount field for all records within each tourist identity group. It identifies the original currency code of each record. The consumption amount field of the scenic spot ticketing system data is in RMB, the consumption amount field of the hotel booking system data includes records priced in USD, the consumption amount field of the transportation ticketing system data is in RMB, the consumption amount field of the catering system data is in RMB, and the consumption amount field of the travel agency group system data includes records priced in either RMB or USD. The data processing module obtains the real-time exchange rate from the exchange rate interface provided by the cooperating financial institution, multiplies the USD price by the real-time exchange rate to convert it into RMB, and converts all consumption amounts into RMB standard amount values, which are retained to two decimal places.
[0062] In practice, the data processing module merges and stores all records after time format conversion, geocoding standardization, and currency unit normalization according to tourist identity. All records under each tourist identity group are aggregated into a group dataset. Each record in the group dataset contains a tourist identity identifier field, a unified time code field, a standard geographic coordinate code field, a standard monetary value field, an original data source identifier field, and a tourism industry type field. After merging and storing, a standardized tourism industry basic data warehouse is generated. In the standardized tourism industry basic data warehouse, a tourist's scenic spot visits, hotel accommodations, transportation, catering consumption, and travel agency participation records are linked and organized in time series form, completely recording a tourist's multi-industry behavior trajectory.
[0063] In one embodiment of the present invention, see [reference] Figure 3 The trajectory generation module extracts all records corresponding to each tourist identification field from the basic data warehouse of the entire tourism industry. It sorts and concatenates the geographic location fields of each record according to the order of the timestamp fields, generating the original behavioral location sequence corresponding to each tourist identification field. For each original behavioral location sequence, it performs stop point identification processing, merging adjacent records whose geographic location changes within a continuous time window are less than a preset distance threshold into a single stop point entity, resulting in a stop point sequence corresponding to each tourist identification field. The preset distance threshold used in the stop point identification processing is adaptively adjusted according to the geographic spatial scale of the specific scenic area. It calculates the dynamic time curvature distance between stop point sequences corresponding to every two tourist identification fields, using this dynamic time curvature distance as the behavioral trajectory similarity between the two tourist identification fields. Based on the behavioral trajectory similarity, it constructs a tourist similarity matrix and performs spectral clustering segmentation on the tourist similarity matrix to obtain multiple tourist behavioral trajectory clusters. It then performs sequence alignment and fusion processing on all stop point sequences corresponding to all tourist identification fields within each tourist behavioral trajectory cluster, generating a representative behavioral trajectory chain for each tourist behavioral trajectory cluster. The set of all representative behavioral trajectory chains is then used as the set of individual tourist behavioral trajectory chains.
[0064] In its implementation, the trajectory generation module extracts all records corresponding to each tourist's identity field from the tourism industry's basic data warehouse. Taking a tourist with ID number "34010219900101xxxx" as an example, the tourism industry's basic data warehouse stores the tourist's ticketing records, hotel accommodation records, transportation records, catering consumption records, and travel agency group records generated in and around Huangshan Scenic Area. The trajectory generation module sorts and concatenates the standard geographic coordinate codes in the geographic location field of each record according to the order of the unified time codes contained in the timestamp field of each record, generating the original behavioral location sequence corresponding to the tourist's identity field. The original behavioral location sequence is a set of latitude and longitude coordinates arranged by time. The coordinates are in the following order: Hefei South Railway Station exit coordinates, Huangshan Scenic Area transfer center coordinates, Guangmingding scenic spot coordinates, Paiyunting scenic spot coordinates, Baiyun Hotel coordinates, Huishang Hometown Restaurant coordinates, and return transportation station coordinates.
[0065] In practice, the trajectory generation module performs stop point identification processing on each original behavioral location sequence. It merges adjacent records whose geographical location changes within a continuous time window are less than a preset distance threshold into a single stop point entity. The time window is set to 30 minutes. For the original behavioral location sequence of the aforementioned tourist, within a continuous 40-minute period starting from the coordinates of the Bright Summit scenic spot, the spatial Euclidean distance between the subsequent four coordinate points and the coordinates of the Bright Summit scenic spot is within the preset distance threshold. The trajectory generation module merges these four coordinate points along with the coordinates of the Bright Summit scenic spot into a single stop point entity. The representative coordinates of the stop point entity are taken as the geometric center of these coordinate points. The stop point entity records both the entry stop time and the departure stop time. The entry stop time is the timestamp recorded by the coordinates of the Bright Summit scenic spot, and the departure stop time is the timestamp of the last coordinate point within the distance threshold. After performing stop point identification processing on all coordinate points in the original behavioral location sequence, the trajectory generation module obtains the stop point sequence corresponding to the tourist's identity field. The stop point sequence includes the Huangshan Scenic Area Transfer Center stop point, the Bright Summit scenic area stop point, the Baiyun Hotel stop point, and the Huishang Hometown Restaurant stop point.
[0066] In some embodiments, the preset distance threshold used in the stop point identification process is adaptively adjusted according to the geospatial scale of the specific scenic area. The trajectory generation module obtains the geospatial scale parameter of each scenic area from the geographic region dimension field of the tourism industry basic data warehouse. The geospatial scale parameter of Huangshan Scenic Area is the diagonal length of the scenic area boundary rectangle. The preset distance threshold is set to one-thousandth of the diagonal length of the scenic area boundary rectangle. When the diagonal length of the scenic area boundary rectangle is 30,000 meters, the preset distance threshold is 30 meters. When the scenic area is a city park and the diagonal length of the boundary rectangle is 800 meters, the preset distance threshold is 0.8 meters. Before performing stop point identification processing for each scenic area, the trajectory generation module reads the geospatial scale parameter of the corresponding scenic area in real time and automatically calculates the applicable preset distance threshold to achieve adaptive matching of distance thresholds between different scenic areas.
[0067] In practice, the trajectory generation module calculates the dynamic time curvature distance between the stop point sequences corresponding to every two tourist identity fields. One tourist's stop point sequence A includes the Huangshan Scenic Area Transfer Center, Guangmingding Scenic Area, Baiyun Hotel, and Huishang Hometown Restaurant. Another tourist's stop point sequence B includes the Huangshan Scenic Area Transfer Center, Yingkesong Scenic Area, Baiyun Hotel, and Huishang Hometown Restaurant. Both sequences A and B have a length of 4 stop point entities. The trajectory generation module calculates the dynamic time curvature distance between stop point sequences A and B, and uses the calculated dynamic time curvature distance as... The similarity of the behavioral trajectories of these two tourist identification fields is calculated. After calculating the similarity of the behavioral trajectories of all tourists pairwise, the trajectory generation module constructs a tourist similarity matrix based on the behavioral trajectory similarity. The tourist similarity matrix is a symmetric square matrix, with the number of rows and columns equal to the total number of tourists. The element in the i-th row and j-th column of the matrix is the similarity of the behavioral trajectory between the i-th tourist and the j-th tourist. The trajectory generation module performs spectral clustering segmentation on the tourist similarity matrix, and performs dimensionality reduction and clustering based on the feature vector of the tourist similarity matrix to obtain multiple tourist behavioral trajectory clusters. Each tourist behavioral trajectory cluster contains the identification fields of several tourists with similar behavioral trajectories.
[0068] In its implementation, the trajectory generation module performs sequence alignment and fusion processing on the stop point sequences corresponding to all tourist identification fields within a tourist behavior trajectory cluster. A dynamic time warp alignment algorithm is used to align all stop point sequences within the cluster along the time dimension, identifying the high-frequency stop point locations that appear in all stop point sequences. The mean of the spatial coordinates of these high-frequency stop point locations is then calculated to generate a representative behavior trajectory chain for the tourist behavior trajectory cluster. This representative behavior trajectory chain consists of several sequentially arranged representative stop points, each containing fused latitude and longitude coordinates and an average stay duration. The trajectory generation module uses the set of representative behavior trajectory chains from all tourist behavior trajectory clusters as the set of individual tourist behavior trajectory chains. This set encompasses all major tour trajectory patterns within the Huangshan Scenic Area, such as sightseeing route trajectory chains, in-depth hiking route trajectory chains, and fast tour route trajectory chains.
[0069] In one embodiment of the present invention, when calculating the similarity between different tourist behavior sequences, the improved clustering algorithm obtains the first stop point sequence of the first tourist and the second stop point sequence of the second tourist. The first stop point sequence contains a first sequence length of stop point entities, and the second stop point sequence contains a second sequence length of stop point entities. An initial distance matrix is constructed between the first stop point sequence and the second stop point sequence, wherein the number of rows in the initial distance matrix is the length of the first sequence, and the number of columns is the length of the second sequence. Each matrix position in the initial distance matrix is traversed, and the spatial Euclidean distance between the stop point entities in the first stop point sequence and the stop point entities in the second stop point sequence corresponding to that matrix position is calculated. This spatial Euclidean distance is then filled into the matrix. The corresponding positions of the initial distance matrix are used to obtain the filled distance matrix, where the coordinates used to calculate the spatial Euclidean distance are geographic latitude and longitude coordinates based on the WGS-84 coordinate system. Dynamic programming is performed on the filled distance matrix to calculate the cumulative distance value row by row and column by column, starting from the upper left corner of the filled distance matrix. The cumulative distance value is calculated by adding the minimum value among the spatial Euclidean distance of the current position, the left position, the upper position, and the upper left corner position to obtain the cumulative distance matrix. The cumulative distance value at the lower right corner of the cumulative distance matrix is used as the dynamic time curvature distance between the first stop point sequence and the second stop point sequence, and this dynamic time curvature distance is output as the similarity measure between the two tourist behavior sequences.
[0070] In its implementation, the improved clustering algorithm obtains the first stop sequence of the first tourist and the second stop sequence of the second tourist. The first stop sequence is generated by a tourist visiting Huangshan Scenic Area. The first stop sequence contains four stop entities, namely, Huangshan Scenic Area Transfer Center, Bright Summit Scenic Area, Baiyun Hotel, and Huishang Hometown Restaurant. The length of the first sequence is 4. The second stop sequence is generated by another tourist visiting Huangshan Scenic Area. The second stop sequence contains four stop entities, namely, Huangshan Scenic Area Transfer Center, Welcoming Pine Scenic Area, Baiyun Hotel, and Huishang Hometown Restaurant. The length of the second sequence is 4.
[0071] In its implementation, the improved clustering algorithm constructs an initial distance matrix between the first stop point sequence and the second stop point sequence. The initial distance matrix has 4 rows, which is the length of the first sequence, and 4 columns, which is the length of the second sequence. The initial distance matrix is a 4x4 two-dimensional matrix structure. Each matrix position corresponds to a stop point entity in the first stop point sequence and a stop point entity in the second stop point sequence. The matrix position in the i-th row and j-th column corresponds to the i-th stop point entity in the first stop point sequence and the j-th stop point entity in the second stop point sequence. The improved clustering algorithm traverses each matrix position in the initial distance matrix, calculates the spatial Euclidean distance between the stop point entity in the first stop point sequence and the stop point entity in the second stop point sequence corresponding to this matrix position, and fills the calculated spatial Euclidean distance into the corresponding position in the initial distance matrix to obtain the filled distance matrix.
[0072] In its implementation, the improved clustering algorithm, when calculating the spatial Euclidean distance, extracts the geographic latitude and longitude coordinates of the current stop point entity based on the WGS-84 coordinate system from the first stop point sequence, and extracts the geographic latitude and longitude coordinates of the corresponding stop point entity based on the WGS-84 coordinate system from the second stop point sequence. In the first stop point sequence, the WGS-84 geographic latitude and longitude coordinates of the Guangmingding Scenic Area stop point are 30.141789 degrees North latitude and 118.182456 degrees East longitude. In the second stop point sequence, the WGS-84 geographic latitude and longitude coordinates of the Yingkesong Scenic Area stop point are... The coordinates are 30.136524 degrees North latitude and 118.168792 degrees East longitude. The improved clustering algorithm converts the latitude and longitude coordinates of the stop points from degrees to radians. The latitude radian value of the stop point at Guangmingding Scenic Area is 0.526122 and the longitude radian value is 2.062173, while the latitude radian value of the stop point at Yingkesong Scenic Area is 0.526030 and the longitude radian value is 2.061934. The improved clustering algorithm calculates the spatial Euclidean distance between the two stop points based on the converted radian coordinates. The formula for calculating the spatial Euclidean distance is:
[0073]
[0074] in: This represents the spatial Euclidean distance between two points of contact. This represents the Earth's average radius constant, with a value of 6,371,000 meters. This represents the latitude in radians of the entities that are the points of interest in the first stop sequence. This represents the latitude in radians of the stop point entities in the second stop point sequence. This represents the longitude in radians of the entities that are the points of interest in the first stop sequence. This represents the longitude in radians of the stop point entities in the second stop point sequence. Let cosine function be used. Substituting the latitude radian value of 0.526122 and longitude radian value of 2.062173 of the Guangmingding Scenic Area stop point, and the latitude radian value of 0.526030 and longitude radian value of 2.061934 of the Yingkesong Scenic Area stop point into the spatial Euclidean distance calculation formula, we obtain the spatial Euclidean distance between the Guangmingding Scenic Area stop point entity and the Yingkesong Scenic Area stop point entity.
[0075] In its implementation, the improved clustering algorithm calculates the spatial Euclidean distance for each of the 16 matrix positions in the initial distance matrix and fills them in, resulting in a complete filled distance matrix. The element in the first row and first column of the filled distance matrix represents the spatial Euclidean distance between the Huangshan Scenic Area transfer center stop point and the Huangshan Scenic Area transfer center stop point, with a value of zero. The element in the second row and second column of the filled distance matrix represents the spatial Euclidean distance between the Guangmingding Scenic Area stop point and the Yingkesong Scenic Area stop point. The improved clustering algorithm performs dynamic programming to calculate the cumulative distance in the filled distance matrix, starting from the top left corner and calculating the cumulative distance value row by row and column by column. The cumulative distance value is calculated by adding the spatial Euclidean distance of the current position to the cumulative distance values of the positions to the left and above. The minimum of the cumulative distance values of the current position, the top-left position, and the first row of the filling distance matrix is used. If the left position does not exist, the cumulative distance value is equal to the spatial Euclidean distance of the current position plus the cumulative distance value of the previous left position. If the top position does not exist, the cumulative distance value is equal to the spatial Euclidean distance of the current position plus the cumulative distance value of the previous top position. For matrix positions in the filling distance matrix that are neither in the first row nor the first column, the improved clustering algorithm reads the cumulative distance values of the left position, the top position, and the top-left position simultaneously, takes the minimum of the three values, and adds it to the spatial Euclidean distance of the current position to obtain the cumulative distance value of the current position. After completing all calculations row by row and column by column, the cumulative distance matrix is obtained.
[0076] In its implementation, after calculating the cumulative distance matrix, the improved clustering algorithm locates the lower right corner of the cumulative distance matrix, i.e., the matrix position in the 4th row and 4th column, and reads the cumulative distance value stored in this position. This cumulative distance value is the minimum cumulative path distance between the first stop sequence and the second stop sequence after dynamic time bending alignment. The improved clustering algorithm uses this cumulative distance value as the dynamic time bending distance between the first stop sequence and the second stop sequence, and outputs this dynamic time bending distance as a similarity measure between the two tourist behavior sequences of the first and second tourists. The smaller the dynamic time bending distance value, the more similar the sightseeing behavior trajectories of the two tourists are; the larger the dynamic time bending distance value, the more obvious the difference between the sightseeing behavior trajectories of the two tourists are.
[0077] In one embodiment of the present invention, the model building module extracts a set of industry-specific dimension fields from the basic data warehouse of the entire tourism industry. The set of industry-specific dimension fields includes time dimension fields, geographic region dimension fields, tourist profile dimension fields, tourism industry type dimension fields, and consumption level dimension fields. Each behavior trajectory chain in the set of individual tourist behavior trajectory chains is used as a fact record, and the fact record is associated and bound with each dimension field in the set of industry-specific dimension fields to obtain the fact record value corresponding to each dimension field. A time hierarchy structure is established based on the time dimension field, which includes year, quarter, month, day, and hour levels. A spatial hierarchy structure is established based on the geographic region dimension field, which includes country, province, city, scenic area, and attraction levels. The time hierarchy structure, spatial hierarchy structure, tourist profile dimension fields, tourism industry type dimension fields, and consumption level dimension fields are used as the dimension axes of the multidimensional data model, and the fact records are used as the fact measurement values of the multidimensional data model to construct a multidimensional data model of the entire tourism industry.
[0078] In practical implementation, the model building module extracts a set of industry-specific dimension fields from the basic data warehouse of the entire tourism industry. This set includes time dimension fields, geographic region dimension fields, tourist profile dimension fields, tourism industry type dimension fields, and consumption level dimension fields. The time dimension field is parsed from the unified time encoding field of the basic data warehouse of the entire tourism industry, and includes year, quarter, month, date, and hour values. The geographic region dimension field is extracted from the geographic information metadata associated with the standard geographic coordinate encoding field of the basic data warehouse of the entire tourism industry, and includes country and province names. The city name, scenic spot name, and attraction name are among the dimensions. The tourist profile dimension field is extracted from the tourist basic information table associated with the tourist identity identifier field in the basic data warehouse of the entire tourism industry, including age range, gender, and province of origin. The tourism industry type dimension field is mapped from the original data source identifier field of the basic data warehouse of the entire tourism industry, including five industry types: scenic spot visits, hotel accommodation, transportation, catering consumption, and group tours. The consumption level dimension field is segmented and mapped from the standard monetary value field of the basic data warehouse of the entire tourism industry, including three levels: low consumption range, medium consumption range, and high consumption range.
[0079] In practical implementation, the model building module uses each behavioral trajectory chain in the set of individual tourist behavior trajectory chains as a fact record. A single sightseeing route trajectory chain includes four representative stops: the Huangshan Scenic Area transfer center, the Bright Summit scenic area, the Baiyun Hotel, and the Huishang Hometown Restaurant. Each behavioral trajectory chain corresponds to a fact record identifier. The model building module associates and binds these fact records with each dimension field in the business format dimension field set, obtaining the fact record value corresponding to each dimension field. The association operation between the sightseeing route trajectory chain with the fact record identifier "TRACK001" and the time dimension field is to extract the time span of the entire trajectory chain and map the time span to the year level ("2026"), quarter level ("Q2"), month level ("May"), day level ("14th"), and hour level ("09-18"). The association operation between the sightseeing route trajectory chain with the fact record identifier "TRACK001" and the geographic region dimension field is to extract the geographical locations of all representative stops in the trajectory chain and map them to... Corresponding to the spatial hierarchy, the country level is "China", the province level is "Anhui Province", the city level is "Huangshan City", the scenic area level is "Huangshan Scenic Area", and the attraction level is "Guangmingding-Baiyun Hotel-Paiyunting". The association operation between the sightseeing route trajectory chain with the fact record identifier "TRACK001" and the tourist profile dimension field is to extract the statistical values of the tourist group profile generated from the trajectory chain. The age range is "31-40 years old", the gender is "male percentage 68%", and the province of origin is "Anhui Province". The association operation between the sightseeing route trajectory chain with the record identifier "TRACK001" and the tourism industry type dimension field is to statistically analyze the industry types covered in all consumption records associated with the trajectory chain. The tourism industry type dimension field is set to "Scenic Spot Visit - Transportation - Catering Consumption". The association operation between the sightseeing route trajectory chain with the fact record identifier "TRACK001" and the consumption level dimension field is to summarize the total consumption of all tourists associated with the trajectory chain within Huangshan City and calculate the per capita consumption value. The consumption level dimension field is set to "Medium Consumption Range".
[0080] In practical implementation, the model building module establishes a time hierarchy structure based on the time dimension field. This structure comprises five time levels: year, quarter, month, day, and hour. The year level is the root node, the quarter level is a child of the year level, the month level is a child of the quarter level, the day level is a child of the month level, and the hour level is a child of the day level. These five time levels form a top-down, progressively refining time dimension drill-down path through parent-child relationships, and simultaneously a bottom-up, progressively summarizing time dimension roll-up path. The construction module establishes a spatial hierarchy based on the geographic region dimension field. The spatial hierarchy includes five levels: country, province, city, scenic area, and attraction. The country level is the root node of the spatial hierarchy, the province level is a sub-level of the country level, the city level is a sub-level of the province level, the scenic area level is a sub-level of the city level, and the attraction level is a sub-level of the scenic area level. The five spatial levels form a top-down, progressively refined spatial dimension drill-down path through parent-child relationships, and at the same time, a bottom-up, progressively summarized spatial dimension roll-up path.
[0081] In practical implementation, the model building module uses the five levels of the time hierarchy, the five levels of the spatial hierarchy, the age range, gender, and province of origin dimensions included in the tourist profile dimension field, the five tourism industry type dimensions included in the tourism industry type dimension field, and the three consumption level dimensions included in the consumption level dimension field as the dimension axes of the multi-dimensional data model for the entire tourism industry. The time hierarchy defines five drill-down levels on the dimension axis, the spatial hierarchy defines five drill-down levels on the dimension axis, the tourist profile dimension field provides three independent slice dimensions on the dimension axis, and the tourism industry type dimension field provides five enumerations on the dimension axis. The consumption level dimension field provides three enumerated dimension values on the dimension axis. The model building module uses each fact record as a fact metric for the full-process tourism industry multi-dimensional data model. The fact metric includes the number of tourist visits, the number of stops, and the number of consumption records corresponding to the trajectory chain facts. All dimension axes and fact metric values form a star-shaped data cube, thus completing the full-process tourism industry multi-dimensional data model. The full-process tourism industry multi-dimensional data model supports cross-combination queries and summary statistics on the trajectory data of the entire tourism industry at any time level (year-quarter-month-day-hour) and any spatial level (country-province-city-scenic spot-attraction).
[0082] In one embodiment of the present invention, the analysis presentation module receives a user-specified combination of analysis dimensions and a combination of analysis indicators. The combination of analysis dimensions includes a time dimension level selected from a time hierarchy, a spatial dimension level selected from a spatial hierarchy, and a profile slice condition selected from the tourist profile dimension field. Based on the combination of analysis dimensions, the module performs a dimension roll-up or dimension drill-down operation on the multi-dimensional data model of the entire tourism industry, generating an aggregated data cube slice corresponding to the combination of analysis dimensions. Based on the analysis indicator type, the module performs indicator calculation operations on the aggregated data cube slice, including tourist visit frequency statistics, average tourist stay duration calculation, tourist spending summary, and tourism industry conversion rate calculation. The results of the indicator calculation operations are sorted and grouped according to the dimension levels in the combination of analysis dimensions to generate a structured analysis result table. The structured analysis result table is converted into a visual chart object, and all visual chart objects are arranged according to a preset management view template to generate a tourism industry data management view. The visual chart objects include heat maps, trajectory flow maps, stacked bar charts, and trend line charts.
[0083] In practical implementation, the analysis presentation module receives analysis requests submitted by users through the management interface. The analysis request includes a combination of analysis dimensions and types of analysis indicators. The time dimension level in the analysis dimension combination is selected as the quarter level in the time hierarchy structure, specifically the second quarter of 2026. The spatial dimension level in the analysis dimension combination is selected as the scenic area level in the spatial hierarchy structure, specifically Huangshan Scenic Area. The profile slice conditions in the analysis dimension combination select the age range dimension item as 31-40 years old and the gender dimension item as male from the tourist profile dimension field. The types of analysis indicators include four indicators: tourist visit frequency statistics, average tourist stay calculation, tourist spending summary, and tourism industry conversion rate calculation.
[0084] In practical implementation, the analysis and presentation module performs dimensional operations on the multi-dimensional data model of the entire tourism industry based on the combination of analysis dimensions. The time dimension in the analysis dimension combination is specified as the quarterly level. However, the daily and hourly levels of the time hierarchy in the multi-dimensional data model of the entire tourism industry store more granular factual records. The analysis and presentation module performs a dimension roll-up operation, aggregating the factual records from the daily and hourly levels upwards along the parent-child relationship of the time hierarchy to the quarterly level, summarizing all factual records from April 1, 2026 to June 30, 2026. Similarly, the spatial dimension in the analysis dimension combination is specified as the scenic area level. The analysis and presentation module performs a dimension roll-up operation, aggregating the spatial dimension at the scenic area level, and then aggregating the spatial dimension at the scenic area level. The point-level fact records are aggregated upwards along the parent-child relationship from the scenic spot level to the scenic area level in the spatial hierarchy, summarizing the fact records of all scenic spots within the Huangshan Scenic Area. The analysis and presentation module simultaneously applies profile slicing conditions to the tourist profile dimension fields, filtering the fact records corresponding to tourist groups whose age range dimension item is 31-40 years old and whose gender dimension item is male from the multi-dimensional data model of the entire tourism industry. After completing dimension roll-up and slice filtering, an aggregated data cube slice is generated. Each aggregated record in the aggregated data cube slice represents a fact measurement summary value of a tourist behavior trajectory type under the specified time range, specified scenic area range, and specified tourist profile conditions.
[0085] In practical implementation, the analysis and presentation module performs indicator calculation operations on the aggregated data cube slices according to the analysis indicator type. For example, the tourist visit frequency statistics operation involves summing the tourist visit counts in the factual metrics of all factual records in the aggregated data cube slice to obtain the total number of male tourists aged 31-40 who visited Huangshan Scenic Area in the second quarter of 2026. The average tourist stay calculation operation involves extracting the stop point sequence of the trajectory chain associated with each factual record in the aggregated data cube slice, calculating the time span between the time of leaving the last stop point and the time of arriving at the first stop point in the trajectory chain, summing the time spans of all trajectory chains, and dividing by the total tourist visit count to obtain the average tourist stay in hours. Tourist spending... The summarization operation involves iterating through the consumption records associated with each fact record in the aggregated data cube slice, extracting the standard amount value field, and summing them to obtain the total consumption amount of tourists in all tourism sectors within the Huangshan Scenic Area, including sightseeing, hotel accommodation, transportation, and catering. The tourism sector conversion rate calculation involves counting the proportion of tourists switching from one tourism sector type to the next. Taking the conversion from sightseeing to catering as an example, first, the number of trajectory chains in the aggregated data cube slice containing the sightseeing sector type and immediately followed by the catering sector type is counted. Then, the total number of trajectory chains containing the sightseeing sector type in the aggregated data cube slice is counted. Dividing the two yields the sector conversion rate. The formula for calculating the tourism sector conversion rate is:
[0086]
[0087] in: This represents the percentage of tourism business conversion rate. This represents the number of trajectory chains in the aggregated data cube slice that transform from the initial business type to the target business type. This represents the total number of trajectory chains of the initial business type contained in the aggregated data cube slice.
[0088] In its implementation, the analysis and presentation module sorts and groups the results of the indicator calculation operations according to the dimension hierarchy in the analysis dimension combination. The sorting is based on the alphabetical order of the names of the scenic spots under the scenic area level in the spatial hierarchy structure. The grouping is based on the monthly level in the time hierarchy structure, dividing the second quarter of 2026 into three months: April, May, and June. A structured analysis result table is generated. The table lists the monthly groupings of the behaviors in the structured analysis result table, listing the scenic spots under Huangshan Scenic Area and the indicator names, including visitor frequency, average visitor stay duration, total visitor spending, and tourism industry conversion rate. In the structured analysis result table, the visitor frequency of the Bright Summit scenic spot in April is a statistical value, the average visitor stay duration is a calculated value, the total visitor spending is an amount value, and the tourism industry conversion rate is a percentage value.
[0089] In practical implementation, the analysis and presentation module converts the structured analysis results tables into visual charts, transforms the monthly visit frequency data for each attraction into a heatmap, with the horizontal axis representing the attraction name, the vertical axis representing the month, and the color intensity reflecting the visit frequency; the monthly trend data summarizing average visitor stay duration and visitor spending is converted into a trend line chart, with the horizontal axis representing the month and the vertical axis representing both stay duration and spending; the composition data summarizing visitor spending at each attraction is converted into a stacked bar chart, with the horizontal axis representing the attraction name, the vertical axis representing spending, and each stacked bar displaying the breakdown of spending on sightseeing, dining, and accommodation; and the data on visitor spending at each attraction... The mobile traffic data is converted into a trajectory flow map. The trajectory flow map uses the Huangshan Scenic Area map as the base map and uses curves to connect the stops at various scenic spots. The thickness of the curves represents the size of the tourist flow. The analysis and presentation module arranges four types of visualization charts—heat map, trajectory flow map, stacked bar chart, and trend line chart—according to a preset management view template. The management view template is a four-quadrant layout. The upper left quadrant contains the heat map, the upper right quadrant contains the trend line chart, the lower left quadrant contains the stacked bar chart, and the lower right quadrant contains the trajectory flow map. The title bar of each quadrant displays the corresponding analysis indicator name and a summary of the filtering conditions. After the layout is completed, a tourism industry data management view is generated and pushed to the user's terminal screen for display.
[0090] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A big data-based full-process tourism industry data center management system, characterized in that: The system includes: The data aggregation module acquires a collection of original data from multiple sources across the entire tourism industry. This collection includes data from scenic spot ticketing systems, hotel booking systems, transportation ticketing systems, catering consumption systems, and travel agency group systems. The data processing module performs heterogeneous data cleaning and standardization processing on the multi-source tourism industry raw data set to generate a standardized tourism industry basic data warehouse. The trajectory generation module calls an improved clustering algorithm to aggregate tourist behavior trajectories in the tourist identity field of the tourism industry basic data warehouse, generating a set of individual tourist behavior trajectory chains. The improved clustering algorithm calculates the similarity between different tourist behavior sequences based on dynamic time curvature distance. The model building module performs multi-granularity data cube construction processing based on the set of individual tourist behavior trajectory chains and the industry dimension fields of the basic data warehouse of the entire tourism industry, generating a multi-dimensional data model of the entire tourism industry. The analysis and presentation module performs online analysis and processing of all-area tourism indicators based on the multi-dimensional data model of the entire tourism industry, and generates a data management view of the entire tourism industry.
2. The big data-based full-process tourism industry data center management system according to claim 1, characterized in that, The steps of performing heterogeneous data cleaning and standardization processing on the multi-source tourism industry-wide raw data set to generate a standardized tourism industry-wide basic data warehouse specifically include: For each data source in the multi-source tourism industry original data set, perform field semantic parsing processing to identify the tourist identity field, timestamp field, geographical location field, and consumption amount field in each data source; Based on the tourist identity identifier field, cross-data source association matching is performed on the records in the multi-source tourism industry original data set, and records from different data sources belonging to the same tourist are marked as the same tourist identity group. For all records within each tourist identity group, a unified time format conversion process is performed on the timestamp field to convert the time representation format of different data sources into a preset unified time encoding format; For all records within each tourist identity group, perform geocoding standardization processing on the geographic location field to convert geographic location descriptions in different formats into a unified geographic coordinate code; For all records within each tourist identity group, currency unit normalization and exchange rate conversion are performed on the consumption amount field to generate a standard amount value; All records, after undergoing unified time format conversion, geocoding standardization, and currency unit normalization, are grouped and merged according to tourist identity for storage, generating the standardized basic data warehouse for the entire tourism industry.
3. The data center management system for the entire tourism industry based on big data as described in claim 2, characterized in that, The geocoding standardization process calls a third-party geographic information service to convert the geographic location described in the text into latitude and longitude coordinates.
4. The data center management system for the entire tourism industry based on big data as described in claim 1, characterized in that, The step of using an improved clustering algorithm to aggregate tourist behavior trajectories in the tourist identity field of the tourism industry-wide basic data warehouse to generate a set of individual tourist behavior trajectory chains specifically includes: Extract all records corresponding to each tourist identification field from the basic data warehouse of the entire tourism industry, sort and concatenate the geographical location fields of each record according to the order of the timestamp fields in each record, and generate the original behavior location sequence corresponding to each tourist identification field. For each of the original behavioral location sequences, a stop point identification process is performed. Adjacent records whose geographical location changes are less than a preset distance threshold within a continuous time window are merged into a stop point entity to obtain the stop point sequence corresponding to each tourist identity field. Calculate the dynamic time warp distance between the stop point sequences corresponding to every two tourist identification fields, and use the dynamic time warp distance as the behavioral trajectory similarity between the two tourist identification fields. Based on the behavioral trajectory similarity, a tourist similarity matrix is constructed, and spectral clustering segmentation is performed on the tourist similarity matrix to obtain multiple tourist behavioral trajectory clusters; The sequence alignment and fusion of the stop point sequences corresponding to all tourist identity fields within each tourist behavior trajectory cluster are performed to generate a representative behavior trajectory chain for each tourist behavior trajectory cluster. The set of all representative behavior trajectory chains is taken as the set of individual tourist behavior trajectory chains.
5. The big data-based full-process tourism industry data center management system according to claim 4, characterized in that, The preset distance threshold used in the stop point identification process is adaptively adjusted according to the specific geographic spatial scale of the scenic area.
6. The big data-based full-process tourism industry data center management system according to claim 4, characterized in that, The improved clustering algorithm calculates the similarity between different tourist behavior sequences based on dynamic time warp distance. The specific working steps of the improved clustering algorithm include: Obtain the first stop sequence of the first tourist and the second stop sequence of the second tourist, wherein the first stop sequence contains a first sequence length of stop entities and the second stop sequence contains a second sequence length of stop entities; Construct an initial distance matrix between the first stop point sequence and the second stop point sequence, wherein the number of rows in the initial distance matrix is the length of the first sequence and the number of columns in the initial distance matrix is the length of the second sequence; Traverse each matrix position in the initial distance matrix, calculate the spatial Euclidean distance between the stop point entities in the first stop point sequence and the stop point entities in the second stop point sequence corresponding to the matrix position, and fill the spatial Euclidean distance into the corresponding position of the initial distance matrix to obtain the filled distance matrix; Perform dynamic programming cumulative distance calculation on the filled distance matrix. Starting from the top left corner of the filled distance matrix, calculate the cumulative distance value row by row and column by column. The cumulative distance value is calculated by adding the minimum value among the spatial Euclidean distance of the current position, the left position, the top position, and the top left corner position to obtain the cumulative distance matrix. The cumulative distance value at the bottom right corner of the cumulative distance matrix is used as the dynamic time curvature distance between the first stop sequence and the second stop sequence, and this dynamic time curvature distance is output as a similarity metric between the two tourist behavior sequences.
7. The big data-based full-process tourism industry data center management system according to claim 6, characterized in that, The coordinates used to calculate the spatial Euclidean distance are geographic latitude and longitude coordinates based on the WGS-84 coordinate system.
8. The data center management system for the entire tourism industry based on big data as described in claim 1, characterized in that, The steps for generating a multi-dimensional data model of the entire tourism industry based on the set of individual tourist behavior trajectory chains and the industry dimension fields of the basic data warehouse of the entire tourism industry include: Extract a set of industry-specific dimension fields from the basic data warehouse of the entire tourism industry. The set of industry-specific dimension fields includes time dimension fields, geographic region dimension fields, tourist profile dimension fields, tourism industry type dimension fields, and consumption level dimension fields. Each behavior trajectory chain in the set of individual tourist behavior trajectory chains is used as a fact record. The fact record is associated and bound with each dimension field in the set of business format dimension fields to obtain the fact record value corresponding to each dimension field. A time hierarchy structure is established based on the time dimension field, which includes year, quarter, month, day, and hour levels. A spatial hierarchy structure is established based on the geographic region dimension field. The spatial hierarchy structure includes national level, provincial level, city level, scenic area level, and attraction level. The time hierarchy, spatial hierarchy, tourist profile dimension field, tourism industry type dimension field, and consumption level dimension field are used as the dimension axes of the multidimensional data model, and the fact records are used as the fact measurement values of the multidimensional data model to construct the full-process tourism full-industry multidimensional data model.
9. The big data-based full-process tourism industry data center management system according to claim 8, characterized in that, The steps for performing online analysis and processing of all-area tourism indicators based on the aforementioned multi-dimensional data model of the entire tourism industry to generate a data management view of the entire tourism industry specifically include: The system receives a combination of analytical dimensions and a type of analytical indicator specified by the user. The combination of analytical dimensions includes a time dimension level selected from the time hierarchy, a spatial dimension level selected from the spatial hierarchy, and a profile slice condition selected from the visitor profile dimension field. Based on the combination of analytical dimensions, perform a dimension roll-up or dimension drill-down operation on the multi-dimensional data model of the entire tourism industry to generate an aggregated data cube slice corresponding to the combination of analytical dimensions. Based on the analysis indicator type, the aggregated data cube slices are subjected to indicator calculation operations, including tourist visit frequency statistics, average tourist stay duration calculation, tourist spending summary, and tourism industry conversion rate calculation. The results of the indicator calculation operations are sorted and grouped according to the dimension hierarchy in the analysis dimension combination to generate a structured analysis result table. The structured analysis results table is converted into a visual chart object, and all the visual chart objects are arranged according to the preset management view template to generate the tourism industry data management view.
10. The big data-based full-process tourism industry data center management system according to claim 9, characterized in that, The visualization chart objects include heatmaps, trajectory flow charts, stacked bar charts, and trend line charts.
Citation Information
Patent Citations
Motor initialization method and apparatus for electric booster brake system
US20150019096A1