Customized travel AI user operation analysis system

By combining K-Means++ and DBSCAN hybrid clustering with multi-layer graph convolutional networks and Transformer architecture, along with multi-armed slot machine algorithm and reinforcement learning Markov decision closure, the problems of homogeneous resource push and violation of physical laws in customized tourism planning systems are solved, realizing personalized and diversified itinerary planning.

CN122434686APending Publication Date: 2026-07-21SHEXIANGJIA (SICHUAN) TOURISM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHEXIANGJIA (SICHUAN) TOURISM CO LTD
Filing Date
2026-05-06
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing customized travel planning systems struggle to effectively capture complex spatiotemporal dependencies when processing user historical behavior logs and multi-source heterogeneous data, leading to issues such as homogeneous resource recommendations and violations of objective physical laws in itinerary planning.

Method used

We employ a hybrid clustering approach combining K-Means++ and DBSCAN with a multi-layer graph convolutional network. We process temporal features through a Transformer architecture and combine a multi-armed slot machine algorithm with reinforcement learning Markov decision closure to generate a travel planning scheme that conforms to physical and spatiotemporal constraints.

Benefits of technology

It effectively captures the spatiotemporal dependencies of user trajectories, limits the push of homogeneous resources, avoids itinerary planning that violates objective physical laws, and improves the personalization and diversity of customized tourism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434686A_ABST
    Figure CN122434686A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of cross technology of artificial intelligence and smart tourism, and discloses a customized tourism AI user operation analysis system, which comprises a data gathering module, a feature map module, a travel reasoning module and a rendering feedback module. The bottom layer is mixed clustering through K-Means++ and DBSCAN, is combined with a multilayer graph convolution network, accurately completes initial screening of candidate destinations in a heterogeneous graph structure, introduces a Transformer architecture to encode time sequence features, and outputs a prediction result of dynamically distributed recommendation weights through a mathematical detection mechanism of constructing an upper bound confidence interval; then, a Dijkstra algorithm is used in combination with a Markov decision closed loop of reinforcement learning to output a target path according to a multidimensional nonlinear penalty reward function under physical space constraints. The system also has a built-in double machine learning mechanism and an incremental calculation loop, can automatically strip and analyze noise and realize parameter hot updating, and realizes low-delay planning and business closed loop optimization under complex space-time constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and smart tourism, specifically a customized tourism AI user operation analysis system. Background Technology

[0002] With the deep integration of artificial intelligence technology and the tourism industry, customized travel planning systems have gradually evolved from initial manual experience-based route planning to automated algorithmic recommendations. However, existing itinerary generation systems mostly rely on shallow machine learning models or rigid label matching rules when processing user historical behavior logs and multi-source heterogeneous data. These mechanisms struggle to effectively capture the complex spatiotemporal dependencies hidden within user trajectories. In the initial destination screening and recommendation stages, the system is highly susceptible to falling into the trap of local extrema in historical data, frequently pushing highly homogenized popular attractions to users. This limitation of the underlying algorithm severely restricts the system's ability to detect long-tail tourism resources, causing customized travel to lose its inherent personalization and diversity.

[0003] Existing technologies generally rely solely on heuristic algorithms to solve for the shortest path in a mathematical sense, while ignoring the dynamic changes in traffic congestion and tourist capacity of attractions in the real world; or they use pure deep learning sequence models that lack explicit boundary control, resulting in the output itinerary planning frequently violating objective physical laws, such as arranging tourist nodes within the absolute closure time window of attractions, or generating continuous commuting spans that exceed the limits of human physiological tolerance. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a customized tourism AI user operation analysis system, which solves the problems of existing customized tourism planning systems causing high latency and wasted computing power due to global recalculation when dealing with local user interaction modifications, or the inability to remove high-dimensional mixed noise in the model iteration closed loop, leading to distorted effect evaluation.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a customized tourism AI user operation analysis system, comprising: The data aggregation module is used to collect customized tourism data from multiple preset channels and perform cleaning and standardization processing to build multimodal data storage; The feature mapping module is used to perform real-time stream processing and offline feature extraction on the standardized customized tourism data, and to build a unified feature storage. The itinerary reasoning module is used to call features in the feature storage, perform user hierarchical calculations based on K-Means++ and DBSCAN hybrid clustering, and calculate recommendation probabilities using a multi-layer graph convolutional network in a heterogeneous triple graph consisting of users, destinations, and attractions to generate a set of candidate destinations. Subsequently, destination preferences are predicted through a time-series model based on the Transformer architecture, and recommendation weights are dynamically assigned to the set of candidate destinations using a multi-armed slot machine algorithm. Finally, under physical and spatiotemporal constraints, an operations research algorithm combined with a Markov decision loop of reinforcement learning is used to deduce and output an itinerary planning scheme that includes the order of attractions and the duration of stay in the set of candidate destinations, based on a reward function that includes traffic time and resource occupation penalty mechanisms. The rendering feedback module is used to visualize the itinerary planning scheme and receive dynamic adjustment instructions from users to update the itinerary planning scheme.

[0006] Furthermore, the data aggregation module includes a multi-source data acquisition unit and a data cleaning and standardization unit. The multi-source data acquisition unit is used to acquire structured data including itinerary forms, unstructured data including text and voice, behavioral log data including user online interaction, and external data including weather and price information, which together serve as the customized tourism data. The data cleaning and standardization unit is used to use a rule engine combined with machine learning text correction algorithms to deduplicate and detect outliers in the customized tourism data, and to perform word segmentation and entity recognition calculations on the text to extract entity data including destination, attraction type, and travel attributes.

[0007] Furthermore, the data aggregation module also includes a data lake and lake warehouse integrated storage unit; the data lake and lake warehouse integrated storage unit is used to store the original customized tourism data through the data lake, and to build a lake warehouse integrated architecture to perform cross-source data query operations, supporting mixed load computing of real-time and offline analysis.

[0008] Furthermore, the feature mapping module includes a real-time stream processing unit and an offline feature engineering unit. The real-time stream processing unit is used to perform real-time feature calculation on the behavior log data stream in the standardized customized tourism data based on the stream processing engine, and to trace the event source of itinerary creation and modification behaviors. It constructs a user behavior trajectory map in computer memory with users and behavior events as nodes and time sequence as edges. The offline feature engineering unit is used to extract user basic features, demand features, and resource features from the standardized customized tourism data. The demand features include destination type tags, daily average number of attractions and playtime indicators, and preset conditional constraint data extracted based on historical data. The extracted features are written into the feature storage for unified management to provide online real-time query and offline training services.

[0009] Furthermore, the trip reasoning module includes a user segmentation and profiling unit; when performing the hybrid clustering, the user segmentation and profiling unit performs global coarse-grained partitioning based on the principle of maximizing the distance between the initial cluster centers, and performs local fine-grained correction on discrete data points through a density clustering algorithm; when performing the multi-layer graph convolutional network calculation, the user's historical interaction behavior and the physical resource attributes of the destination are mapped into hidden layer feature representation vectors and spatial concatenation operations are performed to calculate the user's recommendation probability value for unvisited destinations, and destinations with recommendation probability values ​​greater than a preset probability threshold are extracted to generate a candidate destination set.

[0010] The formula used by the graph neural network to calculate the probability of a user recommending an unvisited destination is as follows: ; in: For users For unvisited destinations The recommended probability value; The hidden feature representation vector of the user node is obtained by aggregating historical user behavior trajectory data through a multi-layer graph convolutional network. The hidden feature representation vector of the destination node is obtained by aggregating historical destination resource feature data through a multi-layer graph convolutional network. This represents the concatenation operation of feature vectors; The weight matrix is ​​a learnable matrix, obtained by backpropagation training based on historical sample data; This is the bias term, obtained through backpropagation training based on historical sample data; This is the Sigmoid activation function.

[0011] Furthermore, the trip reasoning module also includes a demand prediction and intelligent recommendation unit; the demand prediction and intelligent recommendation unit uses a Transformer-based time series model to predict the user's travel time period and destination preference feature vector as the demand prediction data, and uses a multi-armed slot machine algorithm to calculate and assign system recommendation weights to historically unvisited destinations in the candidate destination set and historically visited destinations with feature matching degree reaching a preset threshold.

[0012] The formula for calculating the recommendation weights of the multi-armed slot machine algorithm allocation system is as follows: ; in: To target users Assigned to destination The system recommendation weight; For users With destination The expected value of feature matching degree is calculated based on the destination preference feature vector and resource features; The total number of system recommendations is calculated based on the cumulative values ​​from the system runtime logs. For this destination The total number of times a system has been recommended in history is obtained based on the cumulative value from the system runtime logs. To explore the control coefficient, a fixed non-negative constant was calibrated based on historical experience test data to balance the weight ratio between exploring new destinations and utilizing known destinations.

[0013] Furthermore, the itinerary reasoning module also includes an intelligent itinerary planning unit, and the Markov decision loop is executed based on a reinforcement learning sub-model. The intelligent itinerary planning unit solves for candidate paths between destinations in the candidate destination set based on the Dijkstra algorithm or a genetic algorithm, and forcibly cuts off path branches that cross external traffic time thresholds or fall within the scenic spot closure time window. Based on the feature matching score between the user profile features and the candidate paths, and combined with the tourist attraction resource occupancy values ​​included in the resource features, the reward function of the reinforcement learning sub-model is constructed. The target path containing the order of scenic spots and the duration of stay is output by the reinforcement learning sub-model as the itinerary planning scheme.

[0014] The reward function of the reinforcement learning sub-model is constructed as follows: ; in: To enhance the reward value of the learning sub-model under the current state action; The feature matching score between the user profile features and the candidate path is obtained by calculating the cosine similarity between the feature vectors. This represents the traffic time for the candidate routes; This is a time-consuming penalty function, applied when the traffic time value... When the traffic time exceeds the preset threshold, the function outputs a penalty value that increases exponentially. The value of tourist attraction resource occupancy is calculated by normalizing based on historical statistical data of tourist flow density or queuing time collected in advance. , , These are the penalty weight hyperparameters for matching degree, time cost, and resource consumption, respectively, which were calibrated based on historical experience data.

[0015] Furthermore, the trip reasoning module also includes an operational effectiveness analysis and closed-loop unit; the operational effectiveness analysis and closed-loop unit is used to construct an A / B test calculation framework containing an experimental group and a control group through an algorithm model, compare the application data indicators of trip planning schemes containing different trip generation parameters, and use a causal inference algorithm based on DoubleML to calculate the influence weight values ​​of different trip generation parameters on the system adoption rate of the trip planning scheme, so as to generate system parameter adjustment instructions.

[0016] Furthermore, the rendering feedback module includes an itinerary preview and adjustment unit; the itinerary preview and adjustment unit is used to display the itinerary planning scheme in a graphical interface that combines tables and maps, respond to input operations on the order of attractions and the duration of stay as the dynamic adjustment instructions, and trigger the itinerary reasoning module to lock the anchor nodes that have not been modified by the operation and perform local incremental calculations according to the dynamic adjustment instructions, and update and display the adjusted itinerary planning scheme and its associated time consumption and cost calculation data in real time.

[0017] This invention provides a customized AI-powered user operation analysis system for tourism. It offers the following advantages: 1. This invention outputs user hierarchical results through a hybrid clustering method combining K-Means++ and DBSCAN, and uses a multi-layer graph convolutional network to perform node feature aggregation calculations on a heterogeneous graph composed of users, destinations, and attractions. This mechanism maps discrete data into topological edges, replacing static label matching rules, thereby solving the problem that existing technologies struggle to capture the spatiotemporal dependencies of user trajectories.

[0018] 2. This invention applies the Transformer architecture to process time-series features and combines a multi-armed slot machine algorithm to calculate confidence compensation values ​​for candidate destinations. This calculation step forcibly assigns calculation weights to destination nodes that have not been exposed in the system records, limits the recommendation probability of nodes that have been exposed multiple times, and solves the problem of the underlying algorithm getting stuck in local extrema and pushing homogeneous resources to the terminal.

[0019] 3. This invention introduces a reinforcement learning Markov decision loop under physical and spatiotemporal constraints, setting road network traffic time indicators and scenic spot passenger flow queuing data as penalty variables in the reward function. During the deduction process, the system blocks calculation branches that cross the time window or commuting threshold, outputs routes with calibrated dwell time, and avoids generating travel data that violates objective spatiotemporal laws.

[0020] 4. This invention receives node order or duration adjustment instructions from the terminal through the rendering feedback module, triggering the system to lock unaffected preceding anchor nodes based on the time offset. This computing power scheduling logic only performs recalculation and numerical overwriting on local network intervals where physical boundary collisions occur, replacing the global graph search process and eliminating computing power consumption caused by local interactive modifications. Attached Figure Description

[0021] Figure 1 This is a diagram of the overall architecture of the present invention; Figure 2 This is a schematic diagram of the multimodal data governance of the present invention; Figure 3 This is a schematic diagram of the initial destination screening based on graph neural networks according to the present invention; Figure 4 This is a schematic diagram illustrating the demand prediction of the multi-armed slot machine algorithm of the present invention; Figure 5 This is a schematic diagram of the intelligent route planning based on the reinforcement learning multimodal reward function of the present invention; Figure 6 This is a schematic diagram illustrating the A / B test effect evaluation of the system of the present invention. Detailed Implementation

[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Please see the appendix Figure 1 To be continued Figure 6 This invention provides a customized tourism AI user operation analysis system, which may include the following components: a data aggregation module, a feature mapping module, an itinerary reasoning module, and a rendering feedback module. These components are stored in a storage medium as computer-executable instructions or program code, and are read and executed by a processor through strict time-series data flow and logical operations.

[0024] Before capturing any user behavior data, the multi-source data acquisition unit obtains explicit authorization and consent from the user through the system's front-end interface. Subsequently, the multi-source data acquisition unit acquires structured data including itinerary forms, unstructured data including text and voice, behavioral log data including user online interactions, and external data including weather and price information through an asynchronous non-blocking mechanism. For the acquired voice data, the system calls acoustic and language models to convert it into a computable text sequence. Before the raw data entering the system is persistently written to the data lake, the processor forces dynamic de-identification processing. A one-way hash algorithm is used to irreversibly de-identify direct identifiers such as user device MAC addresses, real names, and contact information, ensuring that subsequent model training and inference are performed entirely within an anonymized high-dimensional tensor space, physically isolating the risk of privacy leakage.

[0025] The feature mapping module maintains a communication connection with the multimodal data storage space via an internal high-speed bus. This module performs dual-channel processing on the standardized customized tourism data. The real-time channel performs window aggregation and event calculation on the online behavioral data stream; the offline channel performs dimensionality reduction and feature extraction on the full dataset.

[0026] In the process of combining data cleaning and standardization with feature extraction, the underlying data transformation logic follows the following feature mapping formula:

[0027] in: The unified feature representation matrix output is written into the feature storage for downstream artificial intelligence models to call and calculate; The standardized customized tourism data vector is obtained by splicing and extracting the original messages captured from multiple acquisition channels according to preset structured rules. The feature mapping weight matrix is ​​extracted from historical business data through unsupervised training of an autoencoder network. The feature mapping bias term is obtained based on the expected distribution of historical tourism samples in the multidimensional data space. The ReLU or Tanh function is selected based on the numerical gradient convergence requirements of the system to enhance the nonlinear representation capability of the feature matrix in high-dimensional space.

[0028] The trip reasoning module, as the core computing hub of the entire system, is used to access the feature data stored in the feature library. This module takes the extracted feature matrix as input parameters and feeds it into a pre-trained artificial intelligence model containing multiple sub-models. The underlying processor accurately outputs user profile features and demand prediction data through matrix multiplication and tensor operations.

[0029] Based on the demand forecast data, the itinerary reasoning module further performs operations research path solving and reinforcement learning reasoning to generate a specific itinerary planning scheme. This scheme is represented in computer memory as a structured linked list containing precise timestamps, geographic coordinate nodes, and serialized attributes, corresponding to real-world travel routes.

[0030] The rendering feedback module reads the structured linked list data and drives the front-end display device to visualize the itinerary planning scheme. This module also listens for electrical signals generated by the peripheral interface to receive dynamic adjustment commands from the user.

[0031] When the rendering feedback module detects an adjustment command for a scenic spot, time, or route, it immediately parses the command into system parameters and triggers the trip inference module to perform incremental recalculation. The calculated new data structure is then returned to the rendering feedback module, enabling real-time updates and rendering of the display interface.

[0032] In this embodiment, the multi-source data acquisition unit acquires structured data including itinerary forms, unstructured data including text and voice, behavioral log data including user online interactions, and external data including weather and price information through an asynchronous non-blocking mechanism. For the acquired voice data, the system calls the acoustic model and language model to convert it into a computable text sequence.

[0033] Upon data entry into the system, the data cleaning and standardization unit is immediately triggered. This unit first uses a rule engine to remove invalid messages with incomplete formatting or missing key fields. Subsequently, the system combines machine learning text correction algorithms to correct typos and map synonyms to the extracted text content, thus physically deduplicating and detecting outliers. For the corrected clean text sequence, the system performs natural language segmentation and inputs it into a pre-trained named entity recognition sequence labeling model. Using Hidden Markov Models or Conditional Random Fields algorithms, it accurately extracts entity data containing destination, attraction type, and travel attributes.

[0034] The data deluge generated by the aforementioned standardization process is fed into a unified data lake and lake warehouse storage unit. The system first performs distributed block storage of the undisturbed original customized tourism data through the underlying file system of the data lake.

[0035] Building upon this foundation, the system introduces open-source table format engines based on Apache Iceberg or Delta Lake on top of the data lake storage layer to construct an integrated lake warehouse architecture. Regarding metadata layer connectivity, the system abandons the traditional distributed metadata catalog and instead tracks the state changes of each underlying Parquet or ORC data file by constructing unified snapshot logs and tree-structured metadata files.

[0036] Through the shared metadata layer structure described above, the system physically bridges the gap between the massive storage at the bottom of the data lake and the analytical boundaries of the upper-layer data warehouse. This integrated lake-warehouse engine natively supports the ACID transaction consistency protocol using a multi-version concurrency control mechanism. When handling high-concurrency travel order writes and real-time feature queries, compute nodes can directly execute cross-source data queries and micro-batch incremental merging operations concurrently on an open and structured data format by parsing the currently active snapshot pointers in the unified metadata directory. This completely eliminates the transmission latency and storage waste caused by redundant data replication between the lake and the warehouse.

[0037] During the feature engineering phase, the real-time stream processing unit within the big data processing and feature engineering module performs real-time feature computation on the behavioral log data stream from the standardized customized tourism data, based on a distributed event-driven stream computing engine. Addressing the common issues of out-of-order delivery and delayed arrival of mobile behavioral logs during network transmission, this unit employs a watermark mechanism based on the log's native event time, combined with a time-sliding window, for aggregation computation.

[0038] In the specific execution logic, the processor first allocates a preset sliding window size parameter (e.g., set to a 30-minute time span) and a sliding step size parameter (e.g., set to a 55-minute trigger interval) to memory to capture the user's instantaneous interactive interests at high frequency on the timeline. At the same time, the system configures a maximum out-of-order tolerance latency threshold for the data stream (e.g., set to 33 seconds) and continuously generates an increasing dynamic watermark based on the largest event timestamp currently observed in the stream minus this threshold.

[0039] When the underlying engine detects that the value of the dynamic watermark has crossed the absolute end boundary of the current sliding window, it forcibly triggers the closure and state aggregation operations of the accumulated data within that window. For extremely late log data whose arrival timestamp is severely delayed and lags behind the current watermark, the system no longer blocks the calculation triggering of the main window, but initiates a delayed data diversion mechanism to transfer it to a dedicated side output stream storage area for offline compensation aggregation, thereby ensuring low-latency response and high availability of computation for the main feature data stream at the physical computation level.

[0040] Based on the high-quality real-time streaming data generated by the aforementioned sliding window aggregation and out-of-order fault tolerance mechanism, this unit continuously monitors and captures high-frequency user operations such as trip creation and modification, and performs low-level code-level event tracing for each operation.

[0041] Based on the event tracing results, the real-time stream processing unit dynamically constructs a user behavior trajectory graph in the computer's high-speed memory space. This graph structure uses user account identifiers and specific behavioral events as graph nodes, and establishes directed edges according to the absolute chronological order of the behaviors. To accurately quantify the impact weight of different historical interaction events on the current itinerary planning, the system assigns dynamically decaying weights based on time series to the directed edges in the user behavior trajectory graph.

[0042] The edge weight calculation process in the above user behavior trajectory graph follows the following exponential decay model formula:

[0043] in: From the behavior event nodes in the user behavior trajectory graph Pointing to behavior event node The dynamic decay weight value of the directed edge; For behavior event nodes With behavioral event nodes The absolute value of the time interval between them is calculated based on the difference in timestamps of the two events captured by the streaming engine on the server side; The time decay coefficient is obtained by fitting and calibrating the system's historical user immersion and conversion rate log data through a grid search optimization algorithm, and is used to control the forgetting rate of the influence of historical behavioral characteristics. For behavior event nodes The inherent importance score is obtained by directly querying and retrieving the score mapping relationship in the pre-set event type dictionary table; The normalization adjustment parameter is calculated based on the total operation frequency of the current user within a preset historical time period, and is used to map the final weight value to a unified dimension distribution range.

[0044] While running in parallel with the real-time stream processing channel, the offline feature engineering unit initiates periodic batch processing tasks during off-peak periods when system computing resources are redundant. During the feature extraction phase, the system incorporates an algorithm fairness control mechanism. The system calls a pre-defined sensitive field blacklist configuration file. When reading the standardized customized tourism data, it directly skips fields marked in the blacklist. The feature engineering code explicitly removes high-risk biased labels containing user gender, age, geographic residency, and historical spending power, thus blocking biased attributes from entering feature storage at the data source.

[0045] After blocking sensitive attributes, this unit utilizes a distributed computing framework to deeply extract user basic features, demand features, and resource features from the remaining clean data. Demand features are rigorously quantified into destination type tags extracted from historical aggregated data, daily average number of attractions and visit duration metrics, as well as pre-defined physical constraint data. The extracted feature vectors completely strip away the vague, subjective descriptions inherent in human language, transforming them entirely into deterministic machine semantics composed of floating-point numbers.

[0046] Meanwhile, when constructing feature storage and training reinforcement learning sub-models and graph neural networks, the system explicitly introduces a decorrelation penalty term into the loss function of the underlying computation to completely eliminate any implicit association bias that may remain in the data. The processor calculates and minimizes the mutual information between the feature representation vector and protected attributes (such as implicit consumption levels), forcing the computational weights of the multi-armed slot machine algorithm and the recommendation network to rely solely on objective travel preference features and physical spatiotemporal constraints. This mechanism eliminates discriminatory decision-making outputs that violate public order and good morals, such as big data price discrimination and differentiated pricing, from the algorithm's underlying code logic.

[0047] Ultimately, the user dynamic behavior features generated by real-time channel computing and the three-dimensional multimodal solidified features extracted by offline batch processing are uniformly written into an independent feature storage area within the system for centralized scheduling and version isolation management.

[0048] In this embodiment, the system first initiates a hybrid clustering algorithm based on K-Means++ and DBSCAN, utilizing the user basic features and demand features extracted from the feature storage to perform user stratification calculations. During the specific execution process, the processor first calls the K-Means++ algorithm to perform a global coarse-grained partitioning of the user feature vector space. This algorithm calculates the Euclidean distance between feature data points, selects initial centroids based on the principle of maximizing the initial cluster center distance, and iteratively updates them until the sum of squared errors within each cluster converges to a preset lower limit interval, thereby outputting several basic clusters.

[0049] Subsequently, for discrete data points at the edges of basic clusters and high-dimensional spatial regions with uneven density distribution, the system triggers the DBSCAN density clustering algorithm for local fine-grained correction. This algorithm uses a defined neighborhood radius and core point threshold as scanning constraints to forcibly merge high-density connected regions into the final user-hierarchical label system, eliminating the physical interference of noise features on subsequent map analysis.

[0050] After completing user segmentation, the user stratification and profiling units further construct the underlying heterogeneous graph computation space. The system introduces a graph neural network model to perform graph computation modeling on triples consisting of users, destinations, and attractions. In the computer's cache, this heterogeneous graph uses users, destinations, and attractions as entity nodes, and uses users' historical access records, browsing behavior, and the geographical affiliation between destinations and attractions as topological edges to construct a data topology graph that truly maps tourism interaction behavior.

[0051] The graph neural network model comprises an input embedding layer, multiple graph convolutional aggregation layers, and a fully connected output layer at its underlying architecture. The input embedding layer is responsible for reducing the dimensionality of the discrete user node features and destination node features in the feature storage and mapping them to an initial low-dimensional dense vector.

[0052] The multi-layer graph convolutional aggregation layer uses a message-passing mechanism to allow each entity node to aggregate the feature information of its neighboring nodes along the physical topological edges. To prevent the graph neural network from experiencing oversmoothing physical failure due to excessively deep iteration layers causing all node features to become similar, this system limits the setting of this multi-layer graph convolutional aggregation layer to 3 layers in the underlying code.

[0053] In the spatial aggregation operation flow at each layer, the processor introduces a graph attention mechanism. Specifically, the processor first calculates the inner product value between the feature vectors of the target center node and its corresponding neighbor nodes, and inputs this value into the LeakyReLU nonlinear activation function. After Softmax normalization, the relative attention weight coefficients between the two are calculated. Subsequently, the bottom-level operator uses a weighted summation aggregation function to strictly follow the above attention weight allocation ratio to perform a weighted summation update on the hidden feature tensors of all neighbor nodes within the local receptive field. This mechanism enables the system to break the structural limitations of physical connections and dynamically amplify core interactive features that are highly correlated with the current user intent.

[0054] After the above-mentioned feature diffusion, which is strictly limited to 3 rounds, and the summation space aggregation with attention weights, the system generates a high-order hidden layer feature representation vector that highly integrates neighborhood context information.

[0055] Based on this high-order hidden layer feature representation vector, the system enters the fully connected output layer to calculate the probability value of the target user's recommendation for unvisited destinations. This calculation process follows the following nonlinear probability mapping formula:

[0056] in: For users For unvisited destinations The recommended probability value; The hidden feature representation vector of the user node is calculated by aggregating adjacent entity features layer by layer through the multi-layer graph convolutional network based on historical user behavior trajectory data. The hidden feature representation vector of the destination node is calculated by aggregating adjacent entity features layer by layer through the multi-layer graph convolutional network based on historical destination resource feature data. The concatenation operation of feature vectors is represented in the underlying tensor operations as merging two one-dimensional feature tensors end to end into a single-dimensional extended feature tensor according to a specified dimension. The weight matrix is ​​a learnable matrix, obtained through backpropagation iterative training based on a loss function established using historical sample data. The bias term is obtained through backpropagation iterative training based on the loss function established using historical sample data. The Sigmoid activation function is used to strictly map the result of linear matrix multiplication to a continuous real number range from zero to one, in order to conform to the precision definition of probability data at the computer's underlying level.

[0057] To ensure the convergence of the graph neural network model parameters and the generalization ability of the prediction network on unseen data, the system designed a supervised training mechanism based on objective behavior logs. During the offline training phase, the system directly extracts the entity edges of actual travel business orders generated by the target user from the graph database as positive sample data. At the same time, a global random negative sampling mechanism is activated to extract destinations where the user has not generated any interaction records to construct negative sample data.

[0058] Positive and negative sample datasets are input into the network for forward inference. The system then uses a binary cross-entropy loss function to measure the scalar error between the predicted recommendation probability and the actual business label data. After a single forward computation, the underlying processor calls the Adam optimization algorithm to dynamically update the weight matrix along the error backpropagation path based on the calculated loss gradient. With bias term The above fine-tuning process is continuously executed in a parallel architecture with multiple graphics cards until the fluctuation range of the loss value on the validation set remains below the convergence threshold for a long period of time. At this point, the model computation graph and parameter snapshot are saved, completing the preparation work for online inference.

[0059] When online inference computation is triggered, the system batches the target user and all destinations it has not yet visited into a solidified graph neural network for concurrent matrix multiplication calculations, outputting a distribution array of recommendation probability values ​​composed of floating-point numbers. Subsequently, the processor reads a preset probability threshold from the system memory, performs linear traversal and comparison instructions on the above distribution array, accurately extracts destination entities whose recommendation probability values ​​are greater than the preset probability threshold, and assembles them into a candidate destination set in the physical storage area.

[0060] In this embodiment, the demand prediction and intelligent recommendation unit first receives the candidate destination set generated in the preceding steps and the user dynamic behavior sequence in the feature storage. The system uses a time series model based on the Transformer architecture to encode and calculate the above dynamic time series. At the physical level of the underlying network structure, this time series model consists of an input embedding module, a position encoder, and stacked multi-head self-attention mechanism layers and feedforward neural network layers connected in sequence.

[0061] When a user's historical interaction event sequence is transformed into a vector input system, the location encoder forcibly injects an absolute time-dimensional location identifier vector into each sequence node to compensate for the temporal features lost in the parallel processing mechanism. Subsequently, the data flows into the multi-head self-attention mechanism layer, where the system computes in parallel the dot product scaling attention values ​​between the query matrix, key matrix, and value matrix. This mechanism enables the model to capture the hidden coupling relationships between two distant travel behaviors across different time steps within a fixed-size context window. Finally, through the high-dimensional mapping of the feedforward layer, it outputs a clearly directional predicted value for the user's travel time period and a destination preference feature vector.

[0062] After obtaining the destination preference feature vector with high confidence, the multi-armed slot machine algorithm is launched to perform secondary calculation and allocation of recommendation weights. This algorithm performs independent logical branch calculations for two types of targets in the candidate destination set: one is destinations that have not been visited in the history of the system records, and the other is destinations whose feature matching degree reaches a preset threshold and have been visited in the past.

[0063] Specifically, the calculation process of the recommendation weights in the multi-armed slot machine algorithm allocation system follows a mathematical formula that combines exploration and utilization:

[0064] in: To target users Assigned to candidate destination The system ultimately recommends a weight value, which will serve as a direct input parameter for downstream planning algorithms to determine priority ranking. For users With destination The expected value of feature matching between the two is obtained by the processor calling the cosine similarity calculation unit to perform an inner product operation on the destination preference feature vector output by the preceding network and the objective resource features of the corresponding destination. The total number of recommendation interactions accumulated by all users since system initialization is calculated based on the timestamp statistics of the system's background running logs. For specific candidate destinations The total number of times the system has exposed and recommended the data in history is also obtained by accumulating and reading the system's backend logs. To explore the control coefficient, a fixed non-negative constant was calibrated based on the distribution convergence characteristics of historical experience test data. This coefficient directly determines the proportion of hardware computing power tilted towards exploring unknown regions, and is used to amplify or reduce the compensation effect of the confidence interval on the overall weight.

[0065] The above formula clearly defines the system's hardware resource allocation tendency at different stages. When a certain destination... Before it is frequently exposed (i.e., the denominator) If the value is extremely small, the exploration compensation term within the square root will significantly increase, thereby boosting the overall recommendation weight and forcing the system to allocate computing power to long-tail routes that uncover potential demand. Conversely, as the destination is recommended multiple times, the numerical effect of this compensation term gradually diminishes, and the system weight will completely regress and be controlled by the objectively calculated feature matching expectation value. .

[0066] Through the time-series prediction and dynamic weight allocation mechanism constrained by this calculation formula, the system effectively avoids duplicate resource requests for homogeneous, high-popularity destinations at the physical level.

[0067] In this embodiment, after receiving the set of candidate destinations assigned recommendation weight scores by the demand prediction network, the system first imports them into the underlying operations research algorithm engine. The intelligent trip planning unit calls a pre-set Dijkstra's algorithm or genetic algorithm to solve the initial network topology relationship between each destination in a multi-dimensional vector space. During algorithm execution, the system strictly pulls external traffic time thresholds and absolute opening time periods of each attraction as mandatory physical constraints in the graph traversal process. When the cumulative time consumption of any exploration branch exceeds the threshold, or the expected arrival time falls within the physical time window of the attraction's closure, the system hardware will directly trigger a pruning operation, cutting off the computing power allocation of that branch, and finally outputting candidate paths that conform to objective spatial and temporal logic.

[0068] Subsequently, the system enters the intelligent decision-making phase based on a reinforcement learning architecture. Using a pre-built reinforcement learning sub-model, it filters from the numerous legitimate but suboptimal candidate paths and generates the final itinerary plan. This process is rigorously modeled internally as a sequential Markov decision process. The system's state space is defined as a joint feature vector composed of the current user's location, remaining available time, and the list of visited attractions; the action space is defined as the operation instructions for selecting the next target attraction and its precise stay time combination from the candidate node set.

[0069] In each interactive exploration of the reinforcement learning sub-model, the design of the reward function determines the physical direction of network gradient convergence. The system abandons the coarse evaluation mechanism that relies solely on feature matching degree, and performs unified dimensionalization processing on parameters from three heterogeneous dimensions: user personalized matching degree, real-world physical time consumption, and objective occupancy load of tourism resources.

[0070] The reward function construction mechanism in the reinforcement learning sub-model follows the following multimodal penalty fusion formula:

[0071] in: To enhance the overall reward value obtained by the learning sub-model under the current state action pair, so as to guide the weight update of the policy network; The feature matching score between the user profile features and the currently selected candidate path is directly obtained by the underlying computing unit by extracting the feature vectors of both and performing a cosine similarity metric operation. The cumulative traffic time of candidate paths after the current action decision is calculated by summing the absolute time values ​​returned by the road network calculation interface. This is a non-linear time-consuming penalty function used to dampen commuting distances that exceed the limits of human physiological tolerance. Its underlying calculation follows the piecewise exponential function formula:

[0072] in, Preset traffic time thresholds for the system: such as the maximum tolerable time for a single commute. This is a power control constant used to control the severity of the exponential penalty after exceeding a threshold. It is related to the traffic time value. Greater than At that time, the physical output of the function exhibits an exponentially increasing penalty value, forcing the policy network to abandon long-distance jump behavior; This represents the objective resource occupancy value of the currently selected tourist attraction. This value reflects the degree of congestion in the real physical environment. It is calculated by normalizing the data based on the instantaneous visitor flow density or historical queue length statistics of the attraction, which are pre-collected by the system through sensor networks or external interfaces, using the following Min-Max linear mapping formula:

[0073] in, This represents the absolute value of the instantaneous passenger flow currently collected. and These represent the peak and trough values ​​of visitor flow for the attraction within the historical statistical period; , , These are the hyperparameters for matching degree benefit weight, time cost penalty weight, and resource consumption penalty weight, respectively. The parameters of these three subsystems are obtained through offline testing and grid optimization calibration based on historical order acceptance rate data.

[0074] The system is internally equipped with an experience replay storage structure. During the interaction between the reinforcement learning agent and the simulated interactive environment, each generated set of Markov transformation quadruples containing the current state, the executed action, the immediate reward, and the next state is written to the cache. In the backpropagation update phase, the processor extracts historical record slices from the experience replay pool using a random batch sampling strategy, calculates the target Q value based on Bayesian variance, and then constructs the mean squared error loss function. The specific calculation formula is as follows: First, calculate the first... The target Q value for each sample :

[0075] in, The instant reward value recorded in the slice. Let this be the state vector for the next time step. For all possible candidate actions in the next state; The target network set for the system, These are the frozen parameters for the target network; A preset long-term reward discount factor is used to balance the physical weights of current and future returns. Then, the mean squared error loss between the current prediction network output and the target Q-value is calculated. :

[0076] in, The batch size for a single sampling. For the current evaluation network, The parameters of the network being trained are used to evaluate the network. This represents the current state and executed actions recorded in the slice. Based on the gradient signal calculated using the aforementioned mean squared error formula, the system employs a stochastic gradient descent optimizer to continuously fine-tune the deep neural network parameters along the descent direction of the loss function to evaluate the network. And after a fixed number of iterations, value updated to The process continues until the evaluation performance of the end function or value function reaches a stable limit. After online convergence, the reinforcement learning sub-model performs forward inference based on a greedy strategy within the full action space according to the current user's real-time state vector. The final output is a target path containing a clear sequence of attractions and precise dwell time. The system encapsulates this path into a structured itinerary planning scheme and outputs it to the downstream rendering module.

[0077] In this embodiment, while continuously outputting itinerary planning schemes, the system simultaneously activates its built-in statistical algorithm engine to construct an A / B testing computation framework that includes experimental and control groups. The system uses a preset hash-based traffic splitting algorithm to hard-isolate user request traffic on physical computing nodes, ensuring that users in different groups receive itinerary planning schemes containing differentiated itinerary generation parameters. Subsequently, the background log service continuously collects application data metrics for the itinerary planning schemes in each group.

[0078] To address the analytical biases caused by high-dimensional confounding features such as user-defined attributes and external environment on the collected application data metrics, the system introduces a DoubleML-based dual machine learning causal inference algorithm. This algorithm utilizes a cross-fitting mechanism to orthogonalize the target variable and intervention variable separately, thereby removing non-causal data noise and accurately separating the pure business benefits generated by changes in underlying parameters.

[0079] Specifically, the logic for calculating the weights of different trip generation parameters on the system adoption rate of the trip planning scheme follows the following orthogonal residual regression formula:

[0080] in: The calculated unbiased influence weight values ​​of different travel generation parameters on the system adoption rate are used to generate corresponding system parameter adjustment instructions; For the first The residual data of the system adoption rate of a sample is calculated by subtracting the expected value obtained by fitting and predicting the high-dimensional mixed features based on the random forest algorithm from the actual observed value of the system adoption rate of that sample. For the first The residual data of the parameters generated by each sample's journey are calculated by subtracting the expected value obtained by fitting and predicting the high-dimensional mixed features based on the random forest algorithm from the actual set value of the parameters of the test group to which the sample belongs. The total number of test samples participating in the current causal calculation is calculated based on the number of valid log records within the time window captured by the A / B test calculation framework described above.

[0081] When the aforementioned influence weight values ​​exceed the set statistical significance test threshold, the system kernel will automatically generate system parameter adjustment instructions and hot-update the new parameters to the preceding reinforcement learning and recommendation network configuration files, thereby achieving adaptive iteration of the intelligent model without interrupting the service.

[0082] In the front-end execution chain, the system uses a dedicated graphics rendering engine to visualize the issued itinerary planning scheme. The rendering engine extracts the geographic coordinate coefficient values ​​and time sequence arrangement nodes contained in the data packet, and draws a graphical interface that combines tables and maps on the user's terminal device screen. The table area is bound to the absolute time stream and cost details, while the map area loads a continuous geospatial trajectory with directional arrows.

[0083] While the hybrid display interface is running, the system's interactive listener continuously captures input signals from external devices. When it receives input operations from the user, such as dragging specific attractions in order or modifying the duration of stay, the system directly parses the physical operation into structured dynamic adjustment instructions.

[0084] Upon responding to the dynamic adjustment command, the system no longer performs a full calculation and recalculation of the global path, but instead triggers a local incremental calculation mechanism based on the topology propagation of the directed acyclic graph.

[0085] First, the processor parses the user's current line program planning scheme into a form containing... Directed path sequence of time nodes When the system receives a request for the target node ( When an adjustment command is received, the system calculates the absolute time offset caused by the operation using a timestamp comparison algorithm. .

[0086] Subsequently, the system initiates the anchor point locking procedure to strictly lock the sequence in order. All previous preceding node sequences Mark it as an anchor node with an immutable state.

[0087] To identify and locally update the propagation range, the system initiates a forward propagation algorithm based on a sliding window. The system moves along the directed edges from node... Start updating the estimated arrival time of subsequent nodes sequentially, i.e., the new arrival time. With each forward movement, the system will In real time, the traffic congestion enhancement value model for the scenic spot is input into a pre-defined model for mandatory physical constraint verification: if the updated time... It still falls entirely within the legal physical constraints, and the processor only performs lightweight operations. Overwrite the scalar value and continue to the next node. spread; If the updated time Having breached the legal boundaries, the propagation mechanism immediately occurred at that node. The system blocks the previous legitimate node. Set the first valid node that has not timed out after the affected interval as the new temporary starting point, and only within this truncated local small network space, re-awaken the reinforcement learning sub-model to perform local path solving.

[0088] The processor quickly assesses the associated impacts of the changed nodes, recalculates travel time and resource occupancy indicators, and sends the adjusted itinerary plan and its associated time and cost calculations back to the front end. Finally, the rendering engine instantly refreshes the current view through a two-way data binding mechanism, achieving high-concurrency response and low-latency dynamic interaction for customized travel routes under complex constraints.

[0089] Finally, regarding the controllability of social impact, the intelligent planning scheme of this system serves only as an auxiliary decision-making tool, and the rendering feedback module retains the user's highest modification authority. The system has a hard-line blocking mechanism for sudden severe weather warnings and public safety incidents to ensure that the output travel planning scheme is always within the boundaries of safety, controllability, and compliance with public interests.

Claims

1. A customized tourism AI user operation analysis system, characterized in that, include: The data aggregation module is used to collect customized tourism data from multiple preset channels and perform cleaning and standardization processing to build multimodal data storage; The feature mapping module is used to perform real-time stream processing and offline feature extraction on the standardized customized tourism data, and to build a unified feature storage. The trip reasoning module is used to call the features in the feature storage, perform user hierarchical calculation based on K-Means++ and DBSCAN hybrid clustering, and use a multi-layer graph convolutional network to calculate the recommendation probability in the triple heterogeneous graph composed of users, destinations and attractions to generate a set of candidate destinations. Destination preferences are predicted using a time-series model based on the Transformer architecture, and recommendation weights are dynamically assigned to the candidate destination set using a multi-armed slot machine algorithm. Under physical and spatial constraints, a Markov decision loop combining operations research algorithms and reinforcement learning is used. Based on a reward function that includes traffic time and resource occupation penalty mechanisms, a trip planning scheme including the order of attractions and the duration of stay is deduced and output from the candidate destination set. The rendering feedback module is used to visualize the itinerary planning scheme and receive dynamic adjustment instructions from users to update the itinerary planning scheme.

2. The customized tourism AI user operation analysis system according to claim 1, characterized in that, The data aggregation module includes a multi-source data acquisition unit and a data cleaning and standardization unit; The multi-source data acquisition unit is used to acquire structured data including itinerary forms, unstructured data including text and voice, behavioral log data including user online interaction, and external data including weather and price information, which together serve as the customized tourism data. The data cleaning and standardization unit is used to perform deduplication and outlier detection on the customized tourism data by using a rule engine combined with machine learning text correction algorithms, and to perform word segmentation and entity recognition calculations on the text to extract entity data containing destination, attraction type and travel attributes.

3. The customized tourism AI user operation analysis system according to claim 2, characterized in that, The data aggregation module also includes an integrated storage unit for the data lake and lake warehouse; The integrated data lake and lake warehouse storage unit is used to store the original customized tourism data through the data lake and to build an integrated lake warehouse architecture to perform cross-source data query operations, supporting mixed load computing of real-time and offline analysis.

4. The customized tourism AI user operation analysis system according to claim 1, characterized in that, The feature map module includes a real-time stream processing unit; The real-time stream processing unit is used to perform real-time feature calculation on the behavior log data stream in the standardized customized tourism data based on the stream processing engine, and to trace the event source of the trip creation and modification behavior. It constructs a user behavior trajectory map in computer memory with users and behavior events as nodes and time sequence as edges.

5. The customized tourism AI user operation analysis system according to claim 4, characterized in that, The feature mapping module also includes an offline feature engineering unit; The offline feature engineering unit is used to extract user basic features, demand features, and resource features from the standardized customized tourism data. The demand features include destination type tags extracted based on historical data, daily average number of attractions and play duration indicators, and preset conditional constraint data. The extracted features are written into the feature storage for unified management to provide online real-time query and offline training services.

6. The customized tourism AI user operation analysis system according to claim 1, characterized in that, The trip reasoning module includes user segmentation and profiling units; When performing the hybrid clustering, the user segmentation and profiling unit performs global coarse-grained partitioning based on the principle of maximizing the initial cluster center distance, and performs local fine-grained correction on discrete data points through density clustering algorithm. When performing the multi-layer graph convolutional network calculation, it maps the user's historical interaction behavior and the physical resource attributes of the destination into hidden layer feature representation vectors and performs spatial concatenation operation to calculate the recommendation probability value of the user for the unvisited destination, and extracts the destinations whose recommendation probability value is greater than a preset probability threshold to generate a candidate destination set.

7. A customized tourism AI user operation analysis system according to claim 6, characterized in that, The trip reasoning module also includes a demand forecasting and intelligent recommendation unit; The demand prediction and intelligent recommendation unit uses a Transformer-based time series model to predict user travel time periods and destination preference feature vectors as demand prediction data. It also uses a multi-armed slot machine algorithm to calculate and assign system recommendation weights to historically unvisited destinations in the candidate destination set and historically visited destinations with feature matching degrees reaching a preset threshold.

8. The customized tourism AI user operation analysis system according to claim 7, characterized in that, The trip reasoning module also includes a trip intelligent planning unit, and the Markov decision loop is executed based on a reinforcement learning sub-model. The intelligent trip planning unit solves for candidate paths between destinations in the candidate destination set based on Dijkstra's algorithm or genetic algorithm, and forcibly cuts off path branches that cross external traffic time thresholds or fall within the scenic spot closure time window. Based on the feature matching score between the user profile features and the candidate paths, and combined with the tourist attraction resource occupancy values ​​included in the resource features, the reward function of the reinforcement learning sub-model is constructed. The reinforcement learning sub-model outputs a target path containing the order of attractions and the duration of stay, which serves as the trip planning scheme.

9. A customized tourism AI user operation analysis system according to claim 8, characterized in that, The trip reasoning module also includes an operational effectiveness analysis and closed-loop unit; The operational effectiveness analysis and closed-loop unit is used to construct an A / B test calculation framework that includes an experimental group and a control group through an algorithm model. It compares the application data indicators of the itinerary planning scheme with different itinerary generation parameters, and uses a causal inference algorithm based on DoubleML to calculate the influence weight values ​​of different itinerary generation parameters on the system adoption rate of the itinerary planning scheme, so as to generate system parameter adjustment instructions.

10. A customized tourism AI user operation analysis system according to claim 1, characterized in that, The rendering feedback module includes a trip preview and adjustment unit; The itinerary preview and adjustment unit is used to display the itinerary planning scheme in a graphical interface that combines tables and maps. It responds to input operations on the order of attractions and the duration of stay as dynamic adjustment instructions. Based on the dynamic adjustment instructions, it triggers the itinerary reasoning module to lock the anchor nodes that have not been modified by the operation and perform local incremental calculations. It updates and displays the adjusted itinerary planning scheme and its associated time consumption and cost calculation data in real time.