Tourism guidance system based on big data AI

Through the tourism guidance system based on big data AI, combined with large language models and multi-source positioning technology, the problems of insufficient interaction and positioning of existing equipment have been solved, personalized explanations and accurate guidance have been achieved, and the user experience and system intelligence have been improved.

CN120725599APending Publication Date: 2025-09-30HEFEI TINGTING ARTIFICIAL INTELLIGENCE APPLICATION TECHNOLOGY SERVICE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510797719.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing tourist guidance equipment lacks natural language interaction capabilities and cannot provide personalized explanations based on users' real-time needs. Most systems lack contextual memory and semantic continuity, and the positioning system is not accurate enough, especially in indoor areas or areas blocked by high-rise buildings. It cannot accurately trigger the explanation content, cannot make intelligent judgments based on user behavior, and cannot achieve emotional perception and personalized adjustment, resulting in a poor user experience.

Method used

A tourism guidance system based on big data AI is adopted, including an interactive engine design module and a multi-source module information fusion positioning module. The interactive engine module adopts a large language model hybrid architecture, context perception and semantic tracking processing mechanism, multi-round dialogue state management and slot guidance, emotion recognition and emotional interaction mechanism. The positioning module adopts positioning data collection and fusion mechanism, extended Kalman filter algorithm fusion model, Thiessen polygon dynamic fence mechanism and location confidence graded evaluation and adaptive adjustment.

Benefits of technology

It significantly improves the naturalness and accuracy of interaction, realizes personalized content generation, enhances positioning accuracy and user experience, has emotional perception capabilities, supports seamless switching between indoor and outdoor, solves the problems of false positioning triggering and repeated explanations in traditional systems, and improves user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005450244370000051
    Figure BDA0005450244370000051
  • Figure BDA0005450244370000052
    Figure BDA0005450244370000052
  • Figure BDA0005450244370000053
    Figure BDA0005450244370000053
Patent Text Reader

Abstract

The invention discloses a tourism guidance system based on big data AI, and belongs to the technical field of tourism guidance, the tourism guidance system comprises an interaction engine design module and a multi-source module information fusion positioning module, the interaction engine design module comprises: a big language model hybrid architecture; a context sensing and semantic tracking processing mechanism; according to the method, interaction naturalness is improved, context association and language naturalness of dialogues are remarkably improved through a multi-round semantic understanding mechanism based on a large language model, the large language model is used for supporting multi-round semantic interaction, slot position guiding and context maintaining, a system can understand continuous intentions of users, and the interaction naturalness is improved through a dialogue state machine and a semantic memory bank. The system can recognize the emotion of the user through tone, speed and context, adjust the explanation tone and content according to the emotional state, have the emotional common feeling ability and improve the user satisfaction degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tourism guidance, and in particular relates to a tourism guidance system based on big data AI. Background Art

[0002] The current tourist guidance equipment has the following technical problems:

[0003] Lack of natural language interaction capabilities, unable to provide personalized explanations based on users' real-time needs;

[0004] Most systems use pre-recorded audio, which lacks contextual memory and semantic continuity;

[0005] The positioning system is not accurate enough, especially when it is indoors or in areas blocked by tall buildings, and cannot accurately trigger the explanation content.

[0006] Unable to make intelligent judgments based on user behavior (such as dwell time, direction, and movement path);

[0007] Emotional perception and personalized adjustment cannot be achieved, and the user experience is poor.

[0008] Based on this, the present invention designs a tourism guidance system based on big data AI to solve the above problems. Summary of the Invention

[0009] The purpose of the present invention is to propose a tourism guidance system based on big data AI in order to solve the problems in the above-mentioned background technology.

[0010] In order to achieve the above object, the present invention adopts the following technical solutions:

[0011] A tourism guidance system based on big data AI includes an interactive engine design module and a multi-source module information fusion positioning module. The interactive engine design module includes:

[0012] Large language model hybrid architecture;

[0013] Context awareness and semantic tracking processing mechanism;

[0014] Multi-round dialogue state management and slot guidance;

[0015] Emotion recognition and emotional interaction mechanisms;

[0016] The multi-source module information fusion positioning module includes:

[0017] Positioning data collection and fusion mechanism;

[0018] Extended Kalman filter algorithm fusion model;

[0019] Thiessen polygon dynamic fence mechanism;

[0020] Position confidence hierarchical assessment and adaptive adjustment.

[0021] As a further description of the above technical solution:

[0022] The large language model hybrid architecture design includes:

[0023] This system uses edge computing and cloud computing to integrate large language models, achieving reasonable resource allocation and improved response efficiency;

[0024] A lightweight edge model is deployed on the user terminal to receive voice input and perform preliminary feature extraction, such as intonation, sound intensity, and keyword recognition.

[0025] The processed feature vectors are sent to the cloud, where a large language model performs complex semantic reasoning tasks, including dialogue generation, scene understanding, and personalized content generation.

[0026] The model layered deployment uses Neural Architecture Search (NAS) technology to determine the first N layers of the large language model deployed on the edge, and the remaining layers are deployed in the cloud.

[0027] Adopting cloud-edge collaborative communication protocol, it optimizes data transmission efficiency by compressing feature vectors and sharing conversation state structures;

[0028] The system uses unified interface standards to ensure consistency and compatibility in the deployment of models on different terminal platforms.

[0029] As a further description of the above technical solution:

[0030] The context awareness and semantic tracking processing mechanism includes:

[0031] Build a conversation history cache queue to record the user's continuous interaction information, including question content, scene context, and device current location;

[0032] Generate intent vectors using keyword extraction and context analysis, and calculate the current topic focus and reasoning direction through the contextual reasoning network;

[0033] For complex dialogue scenarios, the context annotation mechanism is used to maintain logical continuity and avoid irrelevant answers or semantic deviations.

[0034] As a further description of the above technical solution:

[0035] The multi-round dialogue state management and slot guidance include:

[0036] The system uses a finite state machine (FSM) to control the conversation process, and each round of conversation corresponds to a state node;

[0037] The state transition rules are determined by user input and intent analysis results;

[0038] During the interaction process, a slot filling mechanism is introduced to collect necessary information;

[0039] Unfinished slots will be prompted by guidance to prompt users to complete them, enabling multiple rounds of complete conversations;

[0040] Add an "intent conflict resolution mechanism" to avoid incorrect process jumps caused by user semantic ambiguity.

[0041] As a further description of the above technical solution:

[0042] The emotion recognition and emotion interaction mechanism includes:

[0043] 1. Speech feature extraction: This is used to pre-process and segment the user's input speech signal, extracting the following key features from each frame:

[0044] Time domain characteristics:

[0045] Average energy, Indicates the intensity of speech. Angry / excited emotions tend to have higher energy.

[0046] Where E: the average energy of the speech signal over a period of time;

[0047] N: the total number of sampling points in the signal frame;

[0048] x(n): the amplitude of the speech signal at the nth sampling point;

[0049] Zero crossing rate: Indicates changes in speech frequency, suitable for identifying states such as tension and irritability;

[0050] Among them, ZCR: zero crossing rate, which reflects the rapid change of signal frequency;

[0051] N: the number of sampling points in each frame signal;

[0052] x(n): signal value at the nth sampling point;

[0053] sign(x): represents the sign function, which is 1 if x ≥ 0 and -1 if x < 0;

[0054] Frequency domain characteristics:

[0055] Fundamental frequency (F0): extracted through autocorrelation method, used to measure pitch and reflect emotional fluctuations;

[0056] Mel-frequency cepstral coefficients (MFCC): the most commonly used feature representation for speech recognition, taking the first 13 dimensions:

[0057] MFCC=DCT(log(|FFT(x(n))·Mel Filter Bank))

[0058] Where x(n): sampling point of speech signal;

[0059] FFT: Fast Fourier Transform, which transforms time domain signals into frequency domain signals;

[0060] |·|: module length, indicating signal strength;

[0061] Mel Filter Bank: Mel filter bank, a frequency filter that mimics human auditory perception;

[0062] DCT: Discrete cosine transform, used to remove redundant correlations, and the output is a 13-dimensional vector;

[0063] Second, the emotion recognition model uses the BiLSTM+Attention neural network model to complete speech emotion classification. The specific steps are as follows:

[0064] BiLSTM encoding:

[0065] Input feature sequence X = [x1, x2, ..., x T ]:

[0066]

[0067] h t represents the emotional context encoding vector of the t-th frame;

[0068] Attention Mechanism:

[0069] Aggregate the importance of each frame's emotional features to obtain a global emotional representation:

[0070]

[0071] α t : attention weight;

[0072] C: global weighted sentiment vector;

[0073] T: total number of speech frames;

[0074] C: The final weighted emotion feature vector, which represents the aggregated emotion representation of the entire speech segment;

[0075] Softmax Classifier:

[0076] Output of sentiment classification probability:

[0077]

[0078] P(ek |x): Input speech x belongs to emotion category e k probability;

[0079] K: the total number of emotion categories (such as happy, sad, angry, nervous, calm, etc.);

[0080] W k , b k : Classifier weights and biases;

[0081] C: aggregated emotion feature vector;

[0082] 3. Emotional Response Generation Module: After identifying the user's current emotion, it adaptively adjusts the output results in terms of content and tone;

[0083] 4. System feedback loop and model update:

[0084] The system monitors the user's positive feedback words or negative feedback;

[0085] Construct an "emotion-response-feedback" triple and upload it to the cloud as a training sample;

[0086] Used for fine-tuning large language models and iterative upgrading of emotion recognition models.

[0087] As a further description of the above technical solution:

[0088] The positioning data collection and fusion mechanism includes:

[0089] The system positioning part integrates multiple positioning technologies such as LBS, BLE Bluetooth beacon, and geomagnetic fingerprint to build a hybrid positioning architecture:

[0090] The LBS module provides rough latitude and longitude basic positioning, especially in open outdoor areas;

[0091] Bluetooth beacon modules are deployed in key areas of scenic spots to obtain the relative location of devices through RSSI (received signal strength);

[0092] The geomagnetic matching module relies on a pre-built geomagnetic feature fingerprint database and combines it with the device's magnetic sensor to achieve spatial position comparison;

[0093] All data are time-stamped, aligned to the sampling frequency through linear interpolation, and then fed into the fusion engine for calculation;

[0094] Coordinate system fusion design includes: WGS84→UTM (outdoor) or local coordinate system (indoor), BLE coordinate→global coordinate, geomagnetic heading angle→global heading angle.

[0095] As a further description of the above technical solution:

[0096] The extended Kalman filter algorithm fusion model includes:

[0097] The EKF algorithm is used to fuse multi-source positioning data for accurate position estimation:

[0098] GPS is used as the main observation data, with a weight of about 60%, for basic position calibration;

[0099] The geomagnetic module is used for heading correction, with a weight of about 30%, to avoid sudden changes in heading;

[0100] BLE beacons serve as a supplementary information source, with a weight of approximately 10%, to improve redundancy and fault tolerance in complex environments;

[0101] Introducing a "beacon confidence parameter" to dynamically adjust the observation covariance matrix of the filtering process based on the number and stability of received beacons;

[0102] The result is processed by second-order low-pass filtering before output to eliminate high-frequency jitter.

[0103] As a further description of the above technical solution:

[0104] The Thiessen polygon dynamic fence mechanism includes:

[0105] The scenic area map generates a Thiessen polygon spatial grid according to the coordinate points of each scenic spot, and each polygon represents a guide area;

[0106] Once the user's real-time coordinates enter a certain Tyson area, the guide content of that area will be triggered first;

[0107] Solve the problem of repeated or mis-touch in the boundary area of ​​scenic spots in traditional systems;

[0108] Scenic spots can dynamically adjust polygon boundaries based on tourist behavior heat maps.

[0109] As a further description of the above technical solution:

[0110] The position confidence level evaluation and adaptive adjustment include:

[0111] The system calculates a confidence score (0-100 points) to indicate the trustworthiness of the current location;

[0112] If the score is lower than the set threshold (e.g., 60 points), the sampling frequency is automatically increased and a higher-precision mode is switched (e.g., BLE priority).

[0113] If the user stays in an area with weak signals (such as underground space), the system activates the "position persistence mechanism" to maintain logical continuity based on the previous state.

[0114] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0115] 1. This invention improves the naturalness of interaction: The multi-round semantic understanding mechanism based on a large language model significantly improves the contextual relevance and language naturalness of the conversation. Utilizing the large language model to support multi-round semantic interaction, slot guidance, and context maintenance, the system can understand the user's continuous intent. Through the dialogue state machine and semantic memory, it significantly improves the accuracy of responses and the naturalness of interactions, breaking through the limitations of the traditional "single-round question-and-answer" system.

[0116] Personalized content generation: The system can identify user emotions (such as fatigue, excitement, and confusion) through voice tone, speaking speed, and context, and adjust the tone and content of explanations based on their emotional state. This system has the ability to "empathize with emotions" and improve user satisfaction.

[0117] Improved positioning accuracy: Combined use of GPS, Bluetooth beacons, and geomagnetic positioning signal sources, the fusion algorithm includes the Extended Kalman Filter (EKF), dynamic confidence adjustment, and automatic switching of positioning modes;

[0118] Seamless switching between indoor and outdoor, especially suitable for areas with weak GPS signals such as ancient buildings, forest scenic spots, exhibition halls, etc.

[0119] By using Thiessen polygons to divide the area, we can avoid duplication of explanation content and trigger conflicts. We use Thiessen polygon technology to divide attractions into independent areas, solve the problem of false triggering of adjacent attractions, improve the accuracy of explanations, and combine it with the "dynamic fence" mechanism to support customized changes in scenic area layout.

[0120] Smarter explanation triggering: Using Thiessen polygon space division and confidence judgment, the triggering logic is more reliable;

[0121] User experience optimization: The system has the ability to perceive emotions, making explanations more empathetic and personalized;

[0122] The system has strong scalability: it supports dynamic upgrades, fence adjustments and large model optimization, which facilitates deployment and iteration in scenic spots. DETAILED DESCRIPTION

[0123] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0124] The present invention provides a technical solution: a tourism guidance system based on big data AI, including an interactive engine design module and a multi-source module information fusion positioning module, wherein the interactive engine design module includes:

[0125] Large language model hybrid architecture;

[0126] Context awareness and semantic tracking processing mechanism;

[0127] Multi-round dialogue state management and slot guidance;

[0128] Emotion recognition and emotional interaction mechanisms;

[0129] The multi-source module information fusion positioning module includes:

[0130] Positioning data collection and fusion mechanism;

[0131] Extended Kalman filter algorithm fusion model;

[0132] Thiessen polygon dynamic fence mechanism;

[0133] Position confidence hierarchical assessment and adaptive adjustment.

[0134] The large language model hybrid architecture design includes:

[0135] This system uses edge computing and cloud computing to integrate large language models, achieving reasonable resource allocation and improved response efficiency;

[0136] A lightweight edge model is deployed on the user terminal to receive voice input and perform preliminary feature extraction, such as intonation, sound intensity, and keyword recognition.

[0137] The processed feature vectors are sent to the cloud, where a large language model performs complex semantic reasoning tasks, including dialogue generation, scene understanding, and personalized content generation.

[0138] The model layered deployment uses Neural Architecture Search (NAS) technology to determine the first N layers of the large language model deployed on the edge, and the remaining layers are deployed in the cloud.

[0139] Adopting cloud-edge collaborative communication protocol, it optimizes data transmission efficiency by compressing feature vectors and sharing conversation state structures;

[0140] The system uses unified interface standards to ensure consistency and compatibility in the deployment of models on different terminal platforms.

[0141] The context awareness and semantic tracking processing mechanism includes:

[0142] Build a conversation history cache queue to record the user's continuous interaction information, including question content, scene context, and device current location;

[0143] Generate intent vectors using keyword extraction and context analysis, and calculate the current topic focus and reasoning direction through the contextual reasoning network;

[0144] For complex dialogue scenarios, the context annotation mechanism is used to maintain logical continuity and avoid irrelevant answers or semantic deviations.

[0145] The multi-round dialogue state management and slot guidance include:

[0146] The system uses a finite state machine (FSM) to control the conversation process, and each round of conversation corresponds to a state node;

[0147] The state transition rules are determined by user input and intent analysis results;

[0148] During the interaction process, a slot filling mechanism is introduced to collect necessary information;

[0149] Unfinished slots will be prompted by guidance to prompt users to complete them, enabling multiple rounds of complete conversations;

[0150] Add an "intent conflict resolution mechanism" to avoid incorrect process jumps caused by user semantic ambiguity.

[0151] The emotion recognition and emotion interaction mechanism includes:

[0152] 1. Speech feature extraction: This is used to pre-process and segment the user's input speech signal, extracting the following key features from each frame:

[0153] Time domain characteristics:

[0154] Average energy, Indicates the intensity of speech. Angry / excited emotions tend to have higher energy.

[0155] Where E: the average energy of the speech signal over a period of time;

[0156] N: the total number of sampling points in the signal frame;

[0157] x(n): the amplitude of the speech signal at the nth sampling point;

[0158] Zero crossing rate: Indicates changes in speech frequency, suitable for identifying states such as tension and irritability;

[0159] Among them, ZCR: zero crossing rate, which reflects the rapid change of signal frequency;

[0160] N: the number of sampling points in each frame signal;

[0161] x(n): signal value at the nth sampling point;

[0162] sign(x): represents the sign function, which is 1 if x ≥ 0 and -1 if x < 0;

[0163] Frequency domain characteristics:

[0164] Fundamental frequency (F0): extracted through autocorrelation method, used to measure pitch and reflect emotional fluctuations;

[0165] Mel-frequency cepstral coefficients (MFCC): the most commonly used feature representation for speech recognition, taking the first 13 dimensions:

[0166] MFCC=DCT(log(|FFT(x(n))|·Mel Filter Bank))

[0167] Where x(n): sampling point of speech signal;

[0168] FFT: Fast Fourier Transform, which transforms time domain signals into frequency domain signals;

[0169] |·|: module length, indicating signal strength;

[0170] Mel Filter Bank: Mel filter bank, a frequency filter that mimics human auditory perception;

[0171] DCT: Discrete cosine transform, used to remove redundant correlations, and the output is a 13-dimensional vector;

[0172] Second, the emotion recognition model uses the BiLSTM+Attention neural network model to complete speech emotion classification. The specific steps are as follows:

[0173] BiLSTM encoding:

[0174] Input feature sequence X = [x1, x2, ..., x T ]:

[0175]

[0176] h t represents the emotional context encoding vector of the t-th frame;

[0177] Attention Mechanism:

[0178] Aggregate the importance of each frame's emotional features to obtain a global emotional representation:

[0179]

[0180] α t : attention weight;

[0181] C: global weighted sentiment vector;

[0182] T: total number of speech frames;

[0183] C: The final weighted emotion feature vector, which represents the aggregated emotion representation of the entire speech segment;

[0184] Softmax Classifier:

[0185] Output of sentiment classification probability:

[0186]

[0187] P(e k |x): Input speech x belongs to emotion category e k probability;

[0188] K: the total number of emotion categories (such as happy, sad, angry, nervous, calm, etc.);

[0189] W k , b k : Classifier weights and biases;

[0190] C: aggregated emotion feature vector;

[0191] 3. Emotional Response Generation Module: After identifying the user's current emotion, it adaptively adjusts the output results in terms of content and tone;

[0192] 4. System feedback loop and model update:

[0193] The system monitors the user's positive feedback words or negative feedback;

[0194] Construct an "emotion-response-feedback" triple and upload it to the cloud as a training sample;

[0195] Used for fine-tuning large language models and iterative upgrading of emotion recognition models.

[0196] The positioning data collection and fusion mechanism includes:

[0197] The system positioning part integrates multiple positioning technologies such as LBS, BLE Bluetooth beacon, and geomagnetic fingerprint to build a hybrid positioning architecture:

[0198] The LBS module provides rough latitude and longitude basic positioning, especially in open outdoor areas;

[0199] Bluetooth beacon modules are deployed in key areas of scenic spots to obtain the relative location of devices through RSSI (received signal strength);

[0200] The geomagnetic matching module relies on a pre-built geomagnetic feature fingerprint database and combines it with the device's magnetic sensor to achieve spatial position comparison;

[0201] All data are time-stamped, aligned to the sampling frequency through linear interpolation, and then fed into the fusion engine for calculation;

[0202] Coordinate system fusion design includes: WGS84→UTM (outdoor) or local coordinate system (indoor), BLE coordinate→global coordinate, geomagnetic heading angle→global heading angle.

[0203] The extended Kalman filter algorithm fusion model includes:

[0204] The EKF algorithm is used to fuse multi-source positioning data for accurate position estimation:

[0205] GPS is used as the main observation data, with a weight of about 60%, for basic position calibration;

[0206] The geomagnetic module is used for heading correction, with a weight of about 30%, to avoid sudden changes in heading;

[0207] BLE beacons serve as a supplementary information source, with a weight of approximately 10%, to improve redundancy and fault tolerance in complex environments;

[0208] Introducing a "beacon confidence parameter" to dynamically adjust the observation covariance matrix of the filtering process based on the number and stability of received beacons;

[0209] The result is processed by second-order low-pass filtering before output to eliminate high-frequency jitter.

[0210] The Thiessen polygon dynamic fence mechanism includes:

[0211] The scenic area map generates a Thiessen polygon spatial grid according to the coordinate points of each scenic spot, and each polygon represents a guide area;

[0212] Once the user's real-time coordinates enter a certain Tyson area, the guide content of that area will be triggered first;

[0213] Solve the problem of repeated or mis-touch in the boundary area of ​​scenic spots in traditional systems;

[0214] Scenic spots can dynamically adjust polygon boundaries based on tourist behavior heat maps.

[0215] The position confidence level evaluation and adaptive adjustment include:

[0216] The system calculates a confidence score (0-100 points) to indicate the trustworthiness of the current location;

[0217] If the score is lower than the set threshold (e.g., 60 points), the sampling frequency is automatically increased and a higher-precision mode is switched (e.g., BLE priority).

[0218] If the user stays in an area with weak signals (such as underground space), the system activates the "position persistence mechanism" to maintain logical continuity based on the previous state.

[0219] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A tourism guidance system based on big data AI, characterized by: It includes an interactive engine design module and a multi-source module information fusion positioning module. The interactive engine design module includes: Large language model hybrid architecture; Context awareness and semantic tracking processing mechanism; Multi-round dialogue state management and slot guidance; Emotion recognition and emotional interaction mechanisms; The multi-source module information fusion positioning module includes: Positioning data collection and fusion mechanism; Extended Kalman filter algorithm fusion model; Thiessen polygon dynamic fence mechanism; Position confidence hierarchical assessment and adaptive adjustment.

2. A tourism guidance system based on big data AI according to claim 1, characterized in that: The large language model hybrid architecture design includes: This system uses edge computing and cloud computing to integrate large language models, achieving reasonable resource allocation and improved response efficiency; A lightweight edge model is deployed on the user terminal to receive voice input and perform preliminary feature extraction, such as intonation, sound intensity, and keyword recognition. The processed feature vectors are sent to the cloud, where a large language model performs complex semantic reasoning tasks, including dialogue generation, scene understanding, and personalized content generation. The model layered deployment uses Neural Architecture Search (NAS) technology to determine the first N layers of the large language model deployed on the edge, and the remaining layers are deployed in the cloud. Adopting cloud-edge collaborative communication protocol, it optimizes data transmission efficiency by compressing feature vectors and sharing conversation state structures; The system uses unified interface standards to ensure consistency and compatibility in the deployment of models on different terminal platforms.

3. The tourism guidance system based on big data AI according to claim 1 is characterized in that: The context awareness and semantic tracking processing mechanism includes: Build a conversation history cache queue to record the user's continuous interaction information, including question content, scene context, and device current location; Generate intent vectors using keyword extraction and context analysis, and calculate the current topic focus and reasoning direction through the contextual reasoning network; For complex dialogue scenarios, the context annotation mechanism is used to maintain logical continuity and avoid irrelevant answers or semantic deviations.

4. The tourism guidance system based on big data AI according to claim 1 is characterized in that: The multi-round dialogue state management and slot guidance include: The system uses a finite state machine (FSM) to control the conversation process, and each round of conversation corresponds to a state node; The state transition rules are determined by user input and intent analysis results; During the interaction process, a slot filling mechanism is introduced to collect necessary information; Unfinished slots will be prompted by guidance to prompt users to complete them, enabling multiple rounds of complete conversations; Add an "intent conflict resolution mechanism" to prevent incorrect process jumps caused by user semantic ambiguity.

5. The tourism guidance system based on big data AI according to claim 1 is characterized in that: The emotion recognition and emotion interaction mechanism includes:

1. Speech feature extraction: This is used to pre-process and segment the user's input speech signal, extracting the following key features from each frame: Time domain characteristics: Average energy, Indicates the intensity of speech. Angry / excited emotions tend to have higher energy. Where E: the average energy of the speech signal over a period of time; N: the total number of sampling points in the signal frame; x(n): the amplitude of the speech signal at the nth sampling point; Zero crossing rate: Indicates changes in speech frequency, suitable for identifying states such as tension and irritability; Among them, ZCR: zero crossing rate, which reflects the rapid change of signal frequency; N: the number of sampling points in each frame signal; x(n): signal value at the nth sampling point; sign(x): represents the sign function, which is 1 if x ≥ 0 and -1 if x < 0; Frequency domain characteristics: Fundamental frequency (F0): extracted through autocorrelation method, used to measure pitch and reflect emotional fluctuations; Mel-frequency cepstral coefficients (MFCC): the most commonly used feature representation for speech recognition, taking the first 13 dimensions: MFCC=DCT(log(|FFT(x(n))|·Mel Filter Bank)) Where x(n): sampling point of speech signal; FFT: Fast Fourier Transform, which transforms time domain signals into frequency domain signals; |·|: module length, indicating signal strength; Mel Filter Bank: Mel filter bank, a frequency filter that mimics human auditory perception; DCT: Discrete cosine transform, used to remove redundant correlations, and the output is a 13-dimensional vector; Second, the emotion recognition model uses the BiLSTM+Attention neural network model to complete speech emotion classification. The specific steps are as follows: BiLSTM encoding: Input feature sequence X = [x1, x2, ..., x T ]: h t represents the emotional context encoding vector of the t-th frame; Attention Mechanism: Aggregate the importance of each frame's emotional features to obtain a global emotional representation: α t : attention weight; C: global weighted sentiment vector; T: total number of speech frames; C: The final weighted emotion feature vector, which represents the aggregated emotion representation of the entire speech segment; Softmax Classifier: Output of sentiment classification probability: P(e k |x): Input speech x belongs to emotion category e k probability; K: the total number of emotion categories (such as happy, sad, angry, nervous, calm, etc.); W k , b k : Classifier weights and biases; C: aggregated emotion feature vector; 3. Emotional Response Generation Module: After identifying the user's current emotion, it adaptively adjusts the output results in terms of content and tone; 4. System feedback loop and model update: The system monitors the user's positive feedback words or negative feedback; Construct an "emotion-response-feedback" triple and upload it to the cloud as a training sample; Used for fine-tuning large language models and iterative upgrading of emotion recognition models.

6. The tourism guidance system based on big data AI according to claim 1 is characterized in that: The positioning data collection and fusion mechanism includes: The system positioning part integrates multiple positioning technologies such as LBS, BLE Bluetooth beacon, and geomagnetic fingerprint to build a hybrid positioning architecture: The LBS module provides rough latitude and longitude basic positioning, especially in open outdoor areas; Bluetooth beacon modules are deployed in key areas of scenic spots to obtain the relative location of devices through RSSI (received signal strength); The geomagnetic matching module relies on a pre-built geomagnetic feature fingerprint database and combines it with the device's magnetic sensor to achieve spatial position comparison; All data are time-stamped, aligned to the sampling frequency through linear interpolation, and then fed into the fusion engine for calculation; Coordinate system fusion design includes: WGS84→UTM (outdoor) or local coordinate system (indoor), BLE coordinate→global coordinate, geomagnetic heading angle→global heading angle.

7. The tourism guidance system based on big data AI according to claim 1, characterized in that: The extended Kalman filter algorithm fusion model includes: The EKF algorithm is used to fuse multi-source positioning data for accurate position estimation: GPS is used as the main observation data, with a weight of about 60%, for basic position calibration; The geomagnetic module is used for heading correction, with a weight of about 30%, to avoid sudden changes in heading; BLE beacons serve as a supplementary information source, with a weight of approximately 10%, to improve redundancy and fault tolerance in complex environments; Introducing a "beacon confidence parameter" to dynamically adjust the observation covariance matrix of the filtering process based on the number and stability of received beacons; The result is processed by second-order low-pass filtering before output to eliminate high-frequency jitter.

8. The tourism guidance system based on big data AI according to claim 1 is characterized in that: The Thiessen polygon dynamic fence mechanism includes: The scenic area map generates a Thiessen polygon spatial grid according to the coordinate points of each scenic spot, and each polygon represents a guide area; Once the user's real-time coordinates enter a certain Tyson area, the guide content of that area will be triggered first; Solve the problem of repeated or mis-touch in the boundary area of ​​scenic spots in traditional systems; Scenic spots can dynamically adjust polygon boundaries based on tourist behavior heat maps.

9. The tourism guidance system based on big data AI according to claim 1, characterized in that: The position confidence level evaluation and adaptive adjustment include: The system calculates a confidence score (0-100 points) to indicate the trustworthiness of the current location; If the score is lower than the set threshold (e.g., 60 points), the sampling frequency is automatically increased and a higher-precision mode is switched (e.g., BLE priority). If the user stays in an area with weak signal (such as underground space), the system activates the "position persistence mechanism" to maintain logical continuity based on the previous state.

Citation Information

Cited By

  • Automatic scenic spot explanation content pushing method integrating multi-source perception

    CN121434397A

  • Multi-mode AI interactive global intelligent navigation and culture explanation system and method

    CN122281920A