Travel note generation method and system based on travel records
By using real-time data collection and scene inference, and dynamic template matching, structured multimodal travelogues are generated, solving the problems of cumbersome and dry content in existing travelogue generation technologies, and realizing personalized, intelligent and efficient travelogue generation.
Patent Information
- Application Number
- CN202511459268.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-01-13
AI Technical Summary
Existing methods for generating travelogues rely on users' post-trip recollections and manual organization, resulting in dull content that lacks emotion and detail. Furthermore, existing applications cannot intelligently understand the travel context, leading to travelogues that lack personalization and narrative.
By acquiring travel data in real time, inferring scene categories based on geographic location and image analysis, dynamically matching travelogue generation templates, guiding users to input their impressions, and automatically integrating multimodal data to generate structured travelogues, including geographic location, images, and user impressions.
The generated travelogues are rich in emotional detail and have strong logic, enhancing their vividness and personalization. The user experience is more natural and efficient, realizing the intelligent and automated generation of travelogues.
Smart Images

Figure CN121327162A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of smart tourism, and in particular to a travel record-based travelogue generation method and system. BACKGROUND
[0002] With the improvement of people's living standards, tourism has become an important way of daily leisure. During the trip, users usually check in, take pictures, and record their experiences according to the pre-planned travel plan. These behaviors generate a large amount of multi-modal travel record data, including geographic coordinate photos, dining consumption records, accommodation environment images, temporary input text notes, and audio clips, etc.
[0003] However, the existing travelogue generation method has obvious deficiencies. First, the traditional travelogue writing completely depends on the user's post-travel memory and manual arrangement. After the trip, the user often has difficulty accurately recalling the specific scene, mood, and story details corresponding to each photo in the face of hundreds of photos and scattered records in the mobile phone, resulting in a boring, lack of emotion and details, i.e. "can't remember", travelogue content. Secondly, some existing auxiliary application functions are single, mostly only providing the function of arranging photos by time or location, and cannot intelligently understand the travel context, nor can they organically integrate multi-dimensional information such as photos, geographic locations, and user real-time thoughts to generate a truly personalized and narrative travelogue, i.e. "can't tell vividly".
[0004] For example, the user checks in and takes pictures at a specialty coffee shop, and may have a strong impression of the unique flavor of the coffee and the interior decoration at that time. However, if not recorded immediately, when looking at the photos a few days later, he may only remember "drinking coffee at a coffee shop", and lose the most moving details when "recommending" to others. Therefore, the existing technology cannot meet the user's demand for efficient and vivid arrangement and sharing of travel experiences after the trip.
[0005] In summary, there is an urgent need in the art for a technical solution that can actively understand the travel context, intelligently integrate multi-source data, and automatically generate a personalized travelogue rich in emotion and details. SUMMARY
[0006] One of the purposes of the present application is to provide a travel record-based travelogue generation method to solve the above problems.
[0007] In order to achieve the above purpose, a travel record-based travelogue generation method is provided, comprising the following steps: Data acquisition step: real-time acquisition of user travel data, the travel data at least including real-time geographic location information of the user device, image data collected through the device camera, and text or audio data collected through the user input interface; A scene inference step: based on the real-time geographic location information, and in combination with a pre-stored geographic information database, a scene category in which the user is currently located is inferred; the scene category includes but is not limited to natural scenery, historical sites, catering places, accommodation sites and transportation hubs; A will analysis and template matching step: according to the inferred scene category, a potential recording will of the user in the current scene is analyzed, and based on the potential recording will, a travelogue generation template matched with the scene category is dynamically called from a preset template library; the travelogue generation template includes guiding questions and a content structure framework related to the scene; A multi-modal data fusion and content generation step: in response to the dynamically called travelogue generation template, the user is guided to supplement input of feeling data related to the current scene; then the geographic location information, image data, text or audio data and the feeling data are fused to generate a structured travelogue segment corresponding to the current scene; A travelogue synthesis and storage step: according to the travel path of the user, a plurality of structured travelogue segments generated in different scenes are automatically classified and integrated in a time line and a spatial sequence to generate a complete travel travelogue document, and the complete travel travelogue document is stored.
[0008] Further, the scene inference step further includes the following steps: The collected image data is analyzed by image recognition technology to identify scene elements in the image; The identified scene elements are cross-verified with the scene category inferred based on the geographic location.
[0009] Further, in the will analysis and template matching step, the guiding questions are presented in the form of a pop-up window through the graphical user interface of the device, or are proposed in the form of voice interaction through a voice assistant.
[0010] Further, in the multi-modal data fusion and content generation step, the feeling data includes user's experience input through voice; and the step of converting the voice data into text data in real time is further included.
[0011] Further, the method further includes: A sharing step: after the travelogue synthesis and storage step, based on the complete travel travelogue document, an abstract promotional script or a graphic sharing card for recommending to others is automatically generated.
[0012] The second purpose of the present application is to provide a travel record-based travelogue generation system, which includes the following modules: Data acquisition module: for acquiring real-time travel data of the user, the travel data at least including real-time geographic location information of the user equipment, image data collected through the equipment camera, and text or audio data collected through the user input interface; Scene inference module: for inferring the scene category in which the user is currently located based on the real-time geographic location information and in combination with a pre-stored geographic information database; the scene category including but not limited to natural landscape, historical site, catering place, accommodation site and transportation hub; Will analysis and template matching module: for analyzing the potential recording will of the user in the current scene according to the inferred scene category, and dynamically calling a travelogue generation template matched with the scene category from a preset template library based on the potential recording will; the travelogue generation template containing guiding questions and content structure framework related to the scene; Multi-modal data fusion and content generation module: for guiding the user to supplement the input of the sentiment data related to the current scene in response to the dynamically called travelogue generation template; and then fusing the geographic location information, image data, text or audio data and the sentiment data to generate a structured travelogue segment corresponding to the current scene; Travelogue synthesis and storage module: for automatically classifying and integrating a plurality of structured travelogue segments generated in different scenes according to the travel path of the user in time sequence and spatial sequence to generate a complete travel travelogue document and store it.
[0013] Further, the scene inference module further includes the following sub-modules: Scene element analysis and identification sub-module: for analyzing the collected image data through image recognition technology to identify scene elements in the image; Cross-validation sub-module: for cross-validating the identified scene elements with the scene category inferred based on the geographic location.
[0014] Further, in the will analysis and template matching module, the guiding questions are presented in the form of pop-up windows through the graphical user interface of the device, or are proposed in the form of voice interaction through the voice assistant.
[0015] Further, in the multi-modal data fusion and content generation step, the sentiment data includes the experience input by the user through voice, and is further used to convert the voice data into text data in real time.
[0016] Further, the method further includes: Sharing module: for automatically generating an abstract promotional script or a graphic sharing card for recommending to others based on the complete travel travelogue document after the travelogue synthesis and storage step.
[0017] Principles and advantages: 1. This solution solves the problem of "forgetting": through real-time inference based on geographic location and scene, and actively guiding users to record their thoughts at the best time (i.e. when the user is in the scene), it effectively captures the user's instantaneous insights and detailed information in a specific context. This avoids the memory blurring and emotional fading caused by post-recalling, making the generated travelogue content not only contain objective facts (such as location, time), but also contain rich subjective feelings and vivid details, greatly improving the emotional value and authenticity of the travelogue.
[0018] It also solves the problem of "not being lively": by dynamically matching scene-based templates, it provides structured guidance for content generation (for example, asking "food taste" for dining scenes, and asking "historical insights" for historical sites), making the generated travelogue no longer a simple pile of photos and scattered text, but a scene-based narrative with logic and readability, making it easier for users to share and "advertise" to others. It realizes the intelligentization and automation of travelogue generation, greatly improving efficiency.
[0019] 2. This invention liberates users from tedious and time-consuming post-editing work. The system automatically completes data collection, scene classification, content preliminary organization and final travelogue synthesis, and users only need to input the minimum (such as voice recording thoughts) under the guidance of the system. This solves the pain point of users "not wanting to edit and not knowing how to edit", and changes travelogue generation from a "task" to a seamless experience that naturally completes during the travel process. It realizes the intelligentization and automation of travelogue generation, greatly improving efficiency.
[0020] 3. Deep fusion and structured storage of multi-modal data are realized: in traditional ways, photos, locations, texts, etc. are isolated from each other. This invention associates and fuses geographic location, image visual information, and user text / audio thoughts under a unified scene label, generating structured travelogue fragments. This data structure not only facilitates user browsing and retrieval (for example, users can easily find "all records in the coffee shop"), but also lays a solid foundation for subsequent data mining and advanced applications (such as generating travel reports and annual reviews).
[0021] 4. Enhances the initiative and naturalness of human-computer interaction: the system is no longer a passive tool waiting for user instructions, but an assistant that can actively understand the context and provide intelligent services. By inferring user intent through scene, and interacting in a natural way such as pop-up windows or voice dialog, it reduces the user's usage threshold and learning cost, making the recording process itself more natural and smooth, and improving user experience.
[0022] 5. Optimized data organization logic and reproducibility: By automatically categorizing and integrating generated travelogue fragments according to travel routes and spatiotemporal sequences, the final travelogue document faithfully reproduces the user's travel trajectory. This organization method aligns well with human memory habits and narrative logic, allowing users to clearly relive the entire travel experience along temporal and spatial threads when reviewing their trip later.
[0023] 6. This invention fundamentally solves the technical problems of cumbersome travelogue generation process, dry content, lack of emotion and contextual relevance in the prior art by synergistic effect of a series of technical means such as scene inference, intention analysis, dynamic template matching and multimodal data fusion. It ultimately realizes personalized, intelligent and efficient travelogue generation, bringing users an unprecedented convenient and high-quality recording experience. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating a method for generating travelogues based on travel records, according to an embodiment of the present invention. Detailed Implementation
[0025] The following detailed description illustrates the specific implementation method: Example A method for generating travelogues based on travel records, basically as follows: Figure 1 As shown, it includes the following steps: Data acquisition steps: Acquire user travel data in real time or near real time. The travel data includes at least the real-time geographic location information of the user device, image data collected through the device's camera, and text or audio data collected through the user input interface. Data sources for real-time geographic location information include: Global Positioning System (GPS), BeiDou Navigation Satellite System (BDS), cellular network positioning, and Wi-Fi positioning.
[0026] Scene inference step: Based on the real-time geographic location information and combined with a pre-stored geographic information database, infer the scene category where the user is currently located; the scene category includes, but is not limited to, natural landscapes, historical sites, restaurants, accommodations, and transportation hubs; the scene inference step also includes the following steps: The collected image data is analyzed using image recognition technology to identify scene elements in the images; The identified scene elements are cross-validated with the scene categories inferred based on geographical location.
[0027] The will analysis and template matching step: according to the inferred scene category, analyze the potential recording will of the user in the current scene, and based on the potential recording will, dynamically call a travel note generation template matched with the scene category from a preset template library; the travel note generation template includes guiding questions and a content structure framework related to the scene; in the will analysis and template matching step, the guiding questions are presented in the form of a pop-up window through the graphical user interface of the device or are proposed in the form of voice interaction through a voice assistant. Through the introduction of "scene inference" and "will analysis", the system can change from a passive data recorder to an intelligent assistant that understands the environment of the user and predicts the user's intentions. Through "scene inference" to determine "where and what", and through "will analysis" to trigger "what to ask and how to organize", which is a qualitative leap, significantly different from the simple logic of existing applications that arrange photos according to a timeline.
[0028] The multi-modal data fusion and content generation step: in response to the dynamically called travel note generation template, guiding the user to supplement input of feeling data related to the current scene; then, the geographic location information, image data, text or audio data and the feeling data are fused to generate a structured travel note segment corresponding to the current scene; in the multi-modal data fusion and content generation step, the feeling data includes user's experience input through voice, and also includes the step of converting voice data into text data in real time. Instead of using a fixed template, the system dynamically calls the most relevant template according to the dynamically inferred scene. This ensures that the generated travel note content is highly scene-based in structure and content guidance. At the same time, it organically "fuses" heterogeneous data such as geographic location, image, and user-initiated feeling data under the framework of the template, generating a "structured travel note segment" rich in contextual information, rather than just a simple list of data.
[0029] The travel note synthesis and storage step: according to the travel path of the user, automatically categorize and integrate multiple structured travel note segments generated in different scenes in chronological and spatial order to generate a complete travel note document and store it.
[0030] The sharing step: after the travel note synthesis and storage step, based on the complete travel note document, automatically generate an abstract promotional script or a graphic sharing card for recommending to others.
[0031] This embodiment takes a user "Xiaowang" traveling in Hangzhou West Lake as an example to illustrate the operation process of the system.
[0032] 1. System startup and data preparation: Xiaowang arrives at the West Lake Broken Bridge scenic spot and opens the travel note generation application in his mobile phone. The system is initialized and the multi-source data acquisition module is called to continuously obtain the following data: Geographical Location: Accurate latitude and longitude (e.g., N 30.2550°, E 120.1495°) obtained through GPS module.
[0033] Image Data: Xiao Wang took panoramic photos of the Broken Bridge and multiple selfies.
[0034] Scene Inference Step 1: Preliminary Inference Based on Location The Scene Inference Module compares the GPS coordinates with a pre-stored geographical information database. The coordinate point exists in the database and is labeled as "West Lake Broken Bridge - Historical Monument Type Attraction." The system preliminarily infers the scene category as "Historical Monument / Landscape."
[0035] Step 2: Cross-Verification Based on Images (Optional but Preferred Step) The system calls the integrated image recognition unit to analyze the Broken Bridge photos taken by Xiao Wang. Key elements in the images are identified, including "classical stone bridge," "lake water," and "weeping willow." These visual features highly coincide with the "Historical Monument / Landscape" category, cross-verifying and strengthening the accuracy of the scene inference. Finally, the system confirms the current scene as "Historical Monument Type - West Lake Broken Bridge."
[0036] 3. Will Analysis and Template Matching Upon receiving the scene information of "Historical Monument Type - West Lake Broken Bridge," the Template Library and Matching Module immediately searches the template library.
[0037] The template library pre-stores templates for "Historical Monument" scenes, which include the following core content: Guiding Questions: "Does the view before your eyes match the legend?" "Have you been moved by any interesting historical stories or legends here?" "At this moment, how do you feel?" Content Structure Framework: Title: [Location Name] Travel Notes Body: Today, I visited the well-known [Location Name], and the view before my eyes made me feel [mood word]. Here, the [historical story / personal association] impressed me.
[0038] The system presents the above guiding questions to Xiao Wang through the human-computer interaction module in the form of a non-intrusive pop-up window on the phone screen.
[0039] 4. Multimodal Data Fusion and Content Generation Xiaowang saw the pop-up window issue and thought it was very timely. He was too lazy to type, so he clicked the voice input button and said into the phone: "The Broken Bridge is wider than I imagined, the story of White Snake and Xu Xian makes me feel very romantic, and a little snow would be even more appropriate, the mood is particularly comfortable!" The voice recognition unit of the human-computer interaction module converts this audio into text in real time: "The Broken Bridge is wider than I imagined, the story of White Snake and Xu Xian makes me feel very romantic, and a little snow would be even more appropriate, the mood is particularly comfortable!" The content generation engine is triggered and begins to perform data fusion: Call template: historical monument template.
[0040] Fill in data: Automatically fill in the geographical location information "West Lake Broken Bridge" as [location name].
[0041] Extract keywords "romantic" and "comfortable" from user voice text and fill in [mood words].
[0042] Use the user's description of "the story of White Snake and Xu Xian" as [historical story / personal association].
[0043] Associate the current timestamp with the user's photos of the Broken Bridge.
[0044] Generate structured travel notes: Automatically generate travel notes in the following format (JSON format example, for system processing): json { "timestamp": "2023-10-01 10:30:00", "location": "West Lake Broken Bridge", "scene_type": "historical monument", "title": "West Lake Broken Bridge Travel Notes", "content": "Today I came to the famous West Lake Broken Bridge, the scene in front of me made me feel romantic and comfortable. Here the story of White Snake and Xu Xian is very impressive to me.", "media": ["photo_001.jpg", "photo_002.jpg"], "user_sentiment": "The Broken Bridge is wider than I imagined... the mood is particularly comfortable." } 5. Travel Notes Synthesis and Storage Xiaowang ends a day of West Lake tour, visits the Broken Bridge, the White Pagoda, Louwai Lou restaurant and other places. At each location, the system generates structured travel notes similar to the above.
[0045] When Xiaowang clicks "Generate Complete Travel Notes", the content generation engine starts the travel note integration function.
[0046] The engine reads all travel note fragments generated that day, sorts them by timestamp field, and generates a travel path map on the map according to the location field.
[0047] Finally, the system smoothly concatenates these fragments, inserts selected photos taken along the way, forms a complete travel note document with text and photos organized in chronological and spatial sequence, and stores it in local or cloud data storage module.
[0048] 6. Derivative function: generate "recommendation" copy The system further analyzes the complete travel notes, extracts key locations and emotional keywords (such as "romance", "openness", "Legend of the White Snake").
[0049] Based on this information, a sharing copy for social media is automatically generated: "Finally, I came to the legendary West Lake Broken Bridge and experienced the romance of the White Snake and Xu Xian! The lake is open and the heart is at ease. Strongly recommend to everyone! # Hangzhou West Lake # Travel Recommendation" Through the above examples, the complete and automated process of the invention from data collection to final travel note generation is clearly demonstrated, fully embodying its intelligent, personalized and efficient technical effects.
[0050] A travel note generation system based on travel records, including a server and a user terminal, the server and the user terminal are remotely connected, the server includes the following modules: Data acquisition module: for real-time acquisition of user travel data, the travel data at least includes real-time geographic location information of user equipment, image data collected through device camera, and text or audio data collected through user input interface; Scene inference module: for inferring the scene category of the user based on the real-time geographic location information and combining the pre-stored geographic information database; the scene category includes but is not limited to natural landscape, historical site, catering place, accommodation site and transportation hub; the scene inference module further includes the following sub-modules: Scene element analysis and identification sub-module: for analyzing the collected image data through image recognition technology to identify scene elements in the image; Cross-validation sub-module: used to cross-validate the identified scene elements with the scene categories inferred based on the geographic location.
[0051] Will analysis and template matching module: used to analyze the potential recording will of the user in the current scene according to the inferred scene categories, and dynamically call a travelogue generation template matching the scene categories from a preset template library based on the potential recording will; the travelogue generation template contains guiding questions and content structure framework related to the scene; in the will analysis and template matching module, the guiding questions are presented in the form of pop-up windows through the graphical user interface of the device, or are proposed in the form of voice interaction through the voice assistant.
[0052] Multi-modal data fusion and content generation module: used to guide the user to supplement the input of feeling data related to the current scene in response to the dynamically called travelogue generation template; then fuse the geographic location information, image data, text or audio data and the feeling data to generate a structured travelogue segment corresponding to the current scene; in the multi-modal data fusion and content generation step, the feeling data includes the experience input by the user through voice, and is also used to convert the voice data into text data in real time. Travelogue synthesis and storage module: used to automatically classify and integrate a plurality of structured travelogue segments generated in different scenes according to the travel path of the user in the time line and spatial sequence, generate a complete travel travelogue document, and store it.
[0053] Sharing module: used to automatically generate an abstract promotional script or a graphic sharing card for recommending to others based on the complete travel travelogue document after the travelogue synthesis and storage step.
[0054] The above is only an embodiment of the present application, and the common knowledge of specific structures and characteristics in the scheme is not described too much. The person skilled in the art knows all the ordinary technical knowledge in the field of the application before the application date or the priority date, can know all the prior art in the field, and has the ability to apply conventional experimental means before that date. The person skilled in the art can perfect and implement the present scheme under the guidance of the present application, and some typical known structures or known methods should not be an obstacle for the person skilled in the art to implement the present application. It should be noted that for those skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which should also be considered as the protection scope of the present application, which will not affect the effect and practicality of the patent. The protection scope of the present application should be subject to the content of its claims, and the specific implementation mode in the specification can be used to explain the content of the claims.
Claims
1. A method for generating travelogues based on travel records, characterized in that, Includes the following steps: Data acquisition steps: Real-time acquisition of user travel data, which includes at least real-time geographic location information of the user device, image data collected through the device's camera, and text or audio data collected through the user input interface; Scene inference step: Based on the real-time geographic location information and combined with the pre-stored geographic information database, infer the scene category where the user is currently located; the scene category includes, but is not limited to, natural landscapes, historical sites, restaurants, accommodations, and transportation hubs; Intention analysis and template matching steps: Based on the inferred scene category, analyze the user's potential recording intention in the current scene, and based on the potential recording intention, dynamically call the travelogue generation template that matches the scene category from the preset template library; The travelogue generation template includes guiding questions and a content structure framework related to the scenario. Multimodal data fusion and content generation steps: In response to the dynamically invoked travelogue generation template, the user is guided to supplement the input of impression data related to the current scene; then the geographic location information, image data, text or audio data and the impression data are fused and processed to generate a structured travelogue fragment corresponding to the current scene; Travelogue synthesis and storage steps: Based on the user's travel route, multiple structured travelogue fragments generated in different scenarios are automatically classified and integrated according to timeline and spatial sequence to generate a complete travelogue document, which is then stored.
2. The method for generating travelogues based on travel records according to claim 1, characterized in that: The scenario inference step also includes the following steps: The collected image data is analyzed using image recognition technology to identify scene elements in the images; The identified scene elements are cross-validated with the scene categories inferred based on geographical location.
3. The method for generating travelogues based on travel records according to claim 2, characterized in that: In the intention analysis and template matching step, the guiding question is presented in the form of a pop-up window through the device's graphical user interface, or raised in the form of voice interaction through a voice assistant.
4. The method for generating travelogues based on travel records according to claim 3, characterized in that: In the multimodal data fusion and content generation steps, the perception data includes the user's experience input through voice; it also includes a step of converting voice data into text data in real time.
5. The method for generating travelogues based on travel records according to claim 4, characterized in that: The method further includes: Sharing steps: After the travelogue synthesis and storage steps, based on the complete travelogue document, a summary recommendation or graphic sharing card for recommending to others is automatically generated.
6. A travelogue generation system based on travel records, characterized in that: Includes the following modules: Data acquisition module: used to acquire the user's travel data in real time. The travel data includes at least the real-time geographical location information of the user's device, image data acquired through the device's camera, and text or audio data acquired through the user input interface. Scene inference module: Based on the real-time geographic location information and combined with a pre-stored geographic information database, it infers the scene category in which the user is currently located; the scene category includes, but is not limited to, natural landscapes, historical sites, restaurants, accommodations, and transportation hubs; Intention Analysis and Template Matching Module: This module is used to analyze the user's potential recording intention in the current scenario based on the inferred scenario category, and dynamically call travelogue generation templates that match the scenario category from a preset template library based on the potential recording intention. The travelogue generation template includes guiding questions and a content structure framework related to the scenario. Multimodal data fusion and content generation module: In response to the dynamically invoked travelogue generation template, it guides the user to supplement the input of impression data related to the current scene; then it fuses the geographic location information, image data, text or audio data and the impression data to generate a structured travelogue fragment corresponding to the current scene; The travelogue synthesis and storage module is used to automatically classify and integrate multiple structured travelogue fragments generated in different scenarios according to timeline and spatial sequence based on the user's travel route, generate a complete travelogue document, and store it.
7. A travelogue generation system based on travel records according to claim 6, characterized in that: The scene inference module also includes the following sub-modules: Scene element analysis and recognition submodule: used to analyze the collected image data using image recognition technology to identify scene elements in the image; Cross-validation submodule: used to cross-validate the identified scene elements with the scene categories inferred based on geographical location.
8. A travelogue generation system based on travel records according to claim 7, characterized in that: In the intention analysis and template matching module, the guiding questions are presented in the form of pop-ups through the device's graphical user interface, or raised through voice interaction via a voice assistant.
9. A travelogue generation system based on travel records according to claim 8, characterized in that: In the multimodal data fusion and content generation step, the perception data includes the user's experience through voice input, and is also used to convert voice data into text data in real time.
10. A travelogue generation system based on travel records according to claim 9, characterized in that: The method further includes: Sharing module: After the travelogue synthesis and storage steps, based on the complete travelogue document, it automatically generates summary recommendation text or graphic sharing cards for recommending to others.