Multi-mode AI intelligent agent system for personalized service of tourist attraction
The multimodal AI intelligent agent system solves the problems of low knowledge base coverage, insufficient multimodal integration capabilities, and weak personalized services in the cultural and tourism field, realizing intelligent, personalized, and emotional tourism services, and improving the adaptability of services and the accuracy of content matching.
Patent Information
- Application Number
- CN202510962407.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-07
AI Technical Summary
In the cultural and tourism sector, there are shortcomings such as low knowledge base coverage, insufficient multimodal integration capabilities, weak personalized services, lack of emotional interaction, and limited data processing capabilities.
It adopts a multimodal AI intelligent agent system, including a multimodal knowledge base acquisition module, a multimodal knowledge base fusion module, a comprehensive information processing module, a user data management module, and a user interaction module. It achieves the fusion and processing of image, video, text, and geolocation information through CNN, Transformer, and cross-modal attention mechanisms. Combined with user profiles and sentiment analysis, it provides personalized recommendations and emotional interaction.
It enables intelligent, personalized, and emotional tourism services, improves the adaptability of services and the accuracy of content matching, supports multimodal interaction, provides personalized recommendations and emotional interaction, and ensures the real-time nature and security of services.
Smart Images

Figure CN120910344A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence and tourism service technology, and particularly relates to a multi-modal AI agent system for personalized service of a tourist attraction. BACKGROUND
[0002] In recent years, artificial intelligence technology has developed rapidly, and multi-modal data processing has become a research hotspot. In the field of tourism, traditional tourism services mainly rely on paper maps, guide explanations and other single ways to display information. However, tourism information often exists in multiple modalities such as images, videos, texts and geographic locations. The fusion processing of multi-modal data is the key to improving the quality of tourism services and the experience of tourists. For example, the construction of a smart scenic spot requires the combination of surveillance videos, tourist geographic location information and scenic spot text introductions to achieve precise service and management.
[0003] Although there are some multi-modal processing schemes, there are still the following problems in the field of tourism:
[0004] Low knowledge base coverage; insufficient multi-modal fusion capability; weak personalized service; lack of emotional interaction; limited data processing capability. SUMMARY
[0005] The purpose of the present application is to provide a multi-modal AI agent system for personalized service of a tourist attraction to solve the above problems.
[0006] In order to achieve the above application purpose, the technical scheme adopted by the present application is:
[0007] A multi-modal AI agent system for personalized service of a tourist attraction, comprising:
[0008] A multi-modal knowledge base acquisition module for acquiring image, video, text and geographic location multi-modal data of the scenic spot, wherein: image acquisition is through real-time image acquisition by tourist mobile phone, scenic spot monitoring or unmanned aerial vehicle; video acquisition includes scenic spot historical evolution video, performance video and tourist activity video; text acquisition is through the integration of official website, tourism platform and social network text by crawler technology; geographic location acquisition is through real-time acquisition of tourist location by mobile device GPS / Beidou system;
[0009] A multi-modal knowledge base fusion module for semantic modeling and fusion of multi-source heterogeneous data, comprising: a feature extraction unit for extracting image visual features by CNN, extracting text semantic features by Transformer and extracting geographic space features by trajectory encoder; a data alignment unit for realizing semantic alignment between modalities by cross-modal attention mechanism and shared embedding space; a deep fusion unit for generating unified multi-modal semantic representation by multi-modal Transformer or fusion network;
[0010] Comprehensive information processing module: based on the fused semantic representation, multi-class tasks are performed, including: cognitive recognition unit: image recognition, target detection and video behavior analysis are realized; geographic service unit: path planning, surrounding recommendation and positioning correction are performed based on positioning data; content generation unit: automatically generate explanation content and answer inquiries according to user location and historical preferences; personalized recommendation unit: user interests, historical trajectories and behavior patterns are fused to generate customized recommendations; long-term memory unit: records user interaction history, feedback and preferences, and supports continuous service across time periods; emotional interaction unit: identifies user emotions and generates emotional interactive language; error correction and distribution unit: performs semantic correction, completion and task routing to corresponding modules on user input;
[0011] User data management module: used for managing user information, including: information storage unit: structured storage of user basic information, behavior log and preference record; dynamic update unit: update user portrait and recommendation parameters according to real-time behavior and feedback; privacy protection unit: ensure user data security through data encryption and permission management;
[0012] User interaction module: provides multi-modal interaction methods, including: voice interaction unit: supports speech recognition, synthesis and real-time dialogue; text interaction unit: receives text input through APP or web page and responds instantly; image / video interaction unit: identifies user uploaded images / videos and returns related knowledge; geographic information interaction unit: provides location-related services in combination with map navigation.
[0013] As an improvement, the multi-modal knowledge base collection module supports a dynamic update mechanism, which supplements real-time information of scenic spots through images / videos or social network texts taken by tourists in real time.
[0014] As an improvement, the cross-modal deep fusion technology includes a cross-modal attention mechanism and a unified embedding space, realizing semantic alignment and fusion of image, video, text and geographic location information.
[0015] As an improvement, the personalized recommendation unit is realized based on a dynamic user portrait, which is generated by real-time clustering modeling based on user historical behavior data, browsing records, dwell time and preference clicks.
[0016] As an improvement, the emotional interaction unit identifies the emotional tendency in user text / voice through an emotional analysis algorithm, and generates interactive language that conforms to the emotion, thereby strengthening the affinity of human-computer interaction.
[0017] As an improvement, the error correction and distribution unit includes a natural language understanding and error correction mechanism, which can automatically identify spelling errors and ambiguous expressions in user input, and route the corrected intent to corresponding modules such as explanation and recommendation.
[0018] As an improvement, the long-term memory unit records the interaction behavior and access track of the user across time periods and across locations, and supports historical correlation service recommendation in multiple tourism scenarios.
[0019] As an improvement, the image / video interaction unit supports user uploading of live images or short videos to trigger knowledge graph retrieval, and realizes cross-modal question answering from images to text and voice.
[0020] The beneficial effects of the present application are: by deploying a multi-modal intelligent agent based on an AI large model, a 1:1 "virtual tour guide" effect is realized, real-time and flexible response to tourist needs is realized, and the intelligent level and accessibility of the service are significantly improved; the intelligent agent generates text, images, video, voice and other ways through natural language generation, emotion modeling and multi-modal output, making the explanation content more lively and interesting; using multi-modal fusion technology, a graph-text-video-geolocation integrated cognitive system is constructed, supporting tourists to trigger intelligent agent recognition and response through any entry such as taking pictures, voice, location, etc., greatly improving the adaptability and content matching accuracy of the system; the introduction of user portrait and long-term memory mechanism can continuously accumulate tourist preferences, access history, interest types, realize long-term service and dynamic recommendation; through natural language understanding and error correction mechanism, spelling errors, ambiguous expressions and other situations can be automatically identified, and intelligent correction and accurate response can be realized; through data analysis and intelligent scheduling mechanism assisted by large models, tasks can be distributed to voice broadcast, image recognition, location navigation, knowledge explanation and other processing modules according to real-time needs, realizing task-level information shunting and intelligent response, and optimizing service efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 The system architecture diagram of the multi-modal AI intelligent agent system for personalized service of a tourism scenic spot. DETAILED DESCRIPTION
[0022] As shown in Figure 1 A multi-modal AI intelligent agent system for personalized service of a tourism scenic spot includes:
[0023] A multi-modal AI intelligent agent system for personalized service of a tourism scenic spot includes:
[0024] A multi-modal knowledge base acquisition module is used to acquire multi-modal data of images, videos, texts and geographic locations of the scenic spot, wherein: real-time images are acquired through tourist mobile phone photography, scenic area monitoring or unmanned aerial vehicle; video acquisition includes historical evolution images of scenic spots, performance videos and tourist activity videos; text acquisition integrates official website, tourism platform and social network text through crawler technology; geographic location acquisition acquires tourist location in real time through mobile device GPS / Beidou system;
[0025] Multi-modal knowledge base fusion module: for semantic modeling and fusion of multi-source heterogeneous data, including; feature extraction unit: visual features are extracted from images using CNN, text semantic features are extracted using Transformer, and geographic spatial features are extracted using trajectory encoder; data alignment unit: semantic alignment between modalities is achieved through cross-modal attention mechanism and shared embedding space; deep fusion unit: multi-modal Transformer or fusion network is used to generate unified multi-modal semantic representation;
[0026] Comprehensive information processing module: based on the fused semantic representation, multi-class tasks are performed, including: knowledge recognition unit: image recognition, target detection and video behavior analysis are implemented; geographic service unit: path planning, surrounding recommendation and positioning correction are performed based on positioning data; content generation unit: automatically generate explanation content and answer inquiries according to user location and historical preferences; personalized recommendation unit: generate customized recommendations by integrating user interests, historical trajectories and behavior patterns; long-term memory unit: record user interaction history, feedback and preferences to support continuous service across time periods; emotional interaction unit: identify user emotions and generate emotional interactive language; error correction and distribution unit: perform semantic correction, completion and task routing to corresponding modules on user input;
[0027] User data management module: for managing user information, including: information storage unit: structured storage of user basic information, behavior logs and preference records; dynamic update unit: update user portrait and recommendation parameters according to real-time behavior and feedback; privacy protection unit: ensure user data security through data encryption and permission management;
[0028] User interaction module: provides multi-modal interaction methods, including: voice interaction unit: supports speech recognition, synthesis and real-time dialogue; text interaction unit: receives text input through APP or web page and responds instantly; image / video interaction unit: identifies user uploaded images / videos and returns related knowledge; geographic information interaction unit: provides location-related services in combination with map navigation.
[0029] The multi-modal knowledge base collection module supports a dynamic updating mechanism, real-time information of a scenic spot is supplemented through images / videos or social network texts taken by tourists in real time, a cross-modal deep fusion technology includes a cross-modal attention mechanism and a unified embedding space, semantic alignment and fusion of images, videos, texts and geographic location information are realized, a personalized recommendation unit is realized based on a dynamic user portrait, the user portrait is generated through real-time clustering modeling of historical behavior data, browsing records, dwell time and preference clicks of a user, an emotional interaction unit identifies emotional tendencies in texts / voices of a user through an emotional analysis algorithm, and generates interactive languages conforming to the emotions, the affinity of human-computer interaction is strengthened, an error correction and distribution unit includes a natural language understanding and error correction mechanism, can automatically identify spelling errors and ambiguous expressions in input of a user, and routes the corrected intent to corresponding modules such as explanation and recommendation, a long-term memory unit records interactive behaviors and access tracks of a user across time periods and locations, supports historical association service recommendation in multiple tourism scenarios, and an image / video interaction unit supports triggering knowledge graph retrieval from uploaded on-site images or short videos of a user, realizes cross-modal question answering from images to texts and voices.
[0030] In use, a tourist only needs to download a scenic spot exclusive APP on a mobile phone and authorize relevant permissions, and then multi-modal interactive experience can be started, when the tourist arrives at the scenic spot, the system automatically pushes multi-modal introductions such as high-definition images, historical videos and graphic texts of nearby scenic spots through real-time positioning of a mobile device GPS / Beidou, and at the same time, generates personalized tour route suggestions in combination with past browsing records and preferences of the tourist.
[0031] If the tourist is interested in a scenic spot, the tourist can directly ask a question through a voice interaction unit, the system identifies a voice instruction in real time, calls text and video resources fused in a multi-modal knowledge base, and explains in natural language combined with dynamic images; or the tourist can take a photo of the scenic spot and upload the photo to an image / video interaction unit, the system identifies image features through CNN, associates corresponding scenic spot information in a knowledge graph, returns graphic explanation and related legend recommendations, and realizes intuitive interaction of asking a question by taking a photo.
[0032] In the tour process, the system continuously records behavior data such as dwell time and click preferences of the tourist through a long-term memory unit, dynamically updates a user portrait, if the tourist pays attention to red history scenic spots for multiple times, a personalized recommendation unit automatically pushes the same type of scenic spots and related theme activities in the park, when the tourist inputs, an error correction and distribution unit identifies spelling errors and corrects them in real time, and at the same time, routes to an explanation module to provide detailed introduction, and ensures uninterrupted service.
[0033] When the visitor is in a low mood or confused, the emotional interaction unit triggers an empathetic response through emotional keywords in text / speech, adjusts the explanation tone to an encouraging language, and recommends nearby rest areas or simplified explanation content. The system supports cross-period service continuation, and when the visitor visits again, the long-term memory unit automatically recommends unvisited associated scenic spots based on historical trajectories, realizing personalized experience.
[0034] Throughout the service process, the user data management module ensures privacy and security through encrypted storage and permission control, and the multi-modal knowledge base acquisition module dynamically supplements the latest activities, facility changes and other information of the scenic spot by real-time grabbing of live videos and social network evaluations uploaded by visitors, ensuring that the knowledge base always maintains timeliness and richness, and finally forming a closed-loop ecology of data acquisition-fusion processing-intelligent service-feedback optimization, comprehensively improving the intelligent, personalized and emotional level of scenic service.
[0035] The above only describes the preferred embodiments of the present application patent, and is not intended to limit the present application patent. Any modifications, equivalent replacements and improvements made within the spirit and principle of the present application patent should be included in the protection scope of the present application patent.
Claims
1. A multi-modal AI agent system for personalized services of a tourist attraction, characterized in that, The application relates to a multi-modal knowledge base system for scenic spots, which comprises the following modules: a multi-modal knowledge base collection module for collecting image, video, text and geographical position multi-modal data of a scenic spot, wherein image collection is realized by photographing through a tourist's mobile phone, monitoring or a real-time image obtained by a drone; video collection comprises historical evolution images of scenic spots, performance videos and tourist activity videos; text collection is realized by integrating official website, tourism platform and social network texts through a crawler technology; and geographical position collection is realized by real-time acquisition of a tourist's position through a mobile device GPS / Beidou system; a multi-modal knowledge base fusion module for semantic modeling and fusion of multi-source heterogeneous data, which comprises the following units: a feature extraction unit for extracting image visual features by using a CNN, extracting text semantic features by using a Transformer and extracting geographical space features by using a trajectory encoder; a data alignment unit for realizing semantic alignment between modes by using a cross-modal attention mechanism and a shared embedding space; and a deep fusion unit for generating unified multi-modal semantic representation by using a multi-modal Transformer or a fusion network; an integrated information processing module for executing multi-class tasks based on the fused semantic representation, which comprises the following units: a knowledge recognition unit for realizing image recognition, target detection and video behavior analysis; a geographical service unit for realizing path planning, surrounding recommendation and positioning correction based on positioning data; a content generation unit for automatically generating explanation content and answering inquiries according to a user's position and historical preference; a personalized recommendation unit for generating customized recommendation by fusing a user's interest, historical trajectory and behavior mode; a long-term memory unit for recording a user's interactive history, feedback and preference and supporting continuous service across time periods; an emotional interaction unit for identifying a user's emotion and generating emotional interactive language; and an error correction and distribution unit for performing semantic error correction, completion and task routing to corresponding modules on user input; a user data management module for managing user information, which comprises the following units: an information storage unit for structurally storing user basic information, behavior logs and preference records; a dynamic updating unit for updating a user portrait and recommendation parameters according to real-time behavior and feedback; and a privacy protection unit for ensuring user data security by using data encryption and permission management; and a user interaction module for providing multi-modal interaction modes, which comprises the following units: a voice interaction unit for supporting voice recognition, synthesis and real-time dialogue; a text interaction unit for receiving text input through an APP or a webpage and responding instantly; an image / video interaction unit for identifying user-uploaded images / videos and returning related knowledge; and a geographical information interaction unit for providing position-related services in combination with a map navigation. The multi-modal knowledge base collection module supports a dynamic updating mechanism, and real-time information of a scenic spot is supplemented by real-time images / videos or social network texts taken by tourists. The cross-modal deep fusion technology comprises a cross-modal attention mechanism and a unified embedding space, and realizes semantic alignment and fusion of image, video, text and geographical position information. The personalized recommendation unit is realized based on a dynamic user portrait, and the user portrait is generated by real-time clustering modeling based on user historical behavior data, browsing records, dwell time and preference clicks. 2. The multi-modal AI agent system for personalized services of a tourist attraction according to claim 1, wherein, 3.The multi-modal AI agent system for personalized service of a tourist attraction according to claim 1, wherein, 4. The multi-modal AI agent system for personalized services of a tourist attraction according to claim 1, wherein, 5. The multi-modal AI agent system for personalized services of a tourist attraction according to claim 1, wherein, The emotional interaction unit identifies the emotional tendency in the user's text / speech through an emotional analysis algorithm and generates interactive language that conforms to the emotion, thereby enhancing the affinity of human-computer interaction.
6. The multi-modal AI agent system for personalized services of a tourist attraction according to claim 1, wherein, The error correction shunting unit includes a natural language understanding and error correction mechanism, which can automatically identify spelling errors and ambiguous expressions in user input and route the corrected intent to corresponding modules such as explanation and recommendation.
7. The multi-modal AI agent system for personalized services of a tourist attraction according to claim 1, wherein, The long-term memory unit records the user's interaction behavior and access trajectory across time periods and locations, supporting historical association service recommendation in multiple tourism scenarios. 8.The multi-modal AI agent system for personalized service of a tourist attraction according to claim 1, wherein, The image / video interaction unit supports user uploading of live images or short videos to trigger knowledge graph retrieval, enabling cross-modal question answering from images to text and speech.
Citation Information
Cited By
Multi-mode intelligent tour guide method and system
CN121146970A
AI partner system and method based on personal data
CN121787463A
Personalized AI guide system and method for scenic spot
CN122019891A
A personalized AI-guided tour system and method for scenic spots
CN122019891B