Tourism note generation method and device and storage medium

By analyzing multi-source data from various sources during the trip in multiple dimensions and combining it with style preference parameters, structured text and image content of travelogues is generated. This solves the problems of simplicity and lack of personalization in existing travelogues and achieves the generation of high-quality travel narrative content.

CN121901435APending Publication Date: 2026-04-21CHINA SOUTHERN AIRLINES DIGITAL TECHNOLOGY (GUANGDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA SOUTHERN AIRLINES DIGITAL TECHNOLOGY (GUANGDONG) CO LTD
Filing Date
2025-12-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing methods for generating travelogues produce overly simplistic entries that lack adaptation to available materials, resulting in insufficient personalization and failing to meet users' demands for high-quality travel narratives.

Method used

By acquiring multi-source data during the trip, including image data, video data, and metadata, multi-dimensional analysis and processing are performed to extract feature information in the spatiotemporal dimension, content importance dimension, and emotional dimension. Combined with style preference parameters, structured text and image content is generated to create travelogues.

Benefits of technology

The generated travelogues not only accurately recreate the itinerary, but also output different narrative styles of text and images based on individual preferences, significantly improving the intelligence level and narrative coherence of travelogue generation, and enhancing user experience satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901435A_ABST
    Figure CN121901435A_ABST
Patent Text Reader

Abstract

The invention provides a tourism note generation method and device and a storage medium, relates to the technical field of tourism big data, and is used for improving the suitability of tourism notes and materials and improving the individuation of the tourism notes. The method comprises the following steps: acquiring multi-source data in a travel process, wherein the multi-source data comprises image data, video data and metadata; multi-dimensional analysis processing is conducted on the multi-source data to extract feature information, the feature information comprises at least one of the following items: space-time dimension information, content importance dimension information and emotion dimension information, and the space-time dimension information is used for representing a space path and a time sequence of a journey; the content importance dimension information is used for representing the weight of the landscape element in the image, and the emotion dimension information is used for representing the emotion tendency of the travel experience. And obtaining a style preference parameter, and generating a target travel note based on the feature information and the style preference parameter, the target travel note being a structured image-text content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of tourism big data technology, and in particular to a method, apparatus and storage medium for generating travelogues. Background Technology

[0002] With the continuous development of digital tourism services, people generally record travel details by taking pictures, videos and other multimodal data during their travels. These materials have become important carriers for preserving travel memories and sharing travel experiences. Transforming scattered materials into coherent travelogues has gradually become one of the core needs of users.

[0003] Currently, travelogues can be generated semi-automatically or automatically by collecting user travel-related data and integrating materials using preset templates or basic data processing methods, saving users time and effort in manual writing. However, the travelogues generated by these solutions are too simple and lack adaptability to the materials, resulting in insufficient personalization and failing to meet users' needs for high-quality travel narratives. Summary of the Invention

[0004] This application provides a method, apparatus, and storage medium for generating travelogues, which addresses the problem of generating relatively simple travelogues.

[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, this application provides a method for generating travelogues. This method involves acquiring multi-source data from the travel process, including image data, video data, and metadata, with the metadata including shooting time and location information. The multi-source data is then processed through multi-dimensional analysis to extract feature information, which includes at least one of the following: spatiotemporal dimension information, content importance dimension information, and sentiment dimension information. The spatiotemporal dimension information characterizes the spatial path and time sequence of the journey; the content importance dimension information characterizes the weight of landscape elements in the images; and the sentiment dimension information characterizes the emotional tendency of the travel experience. Style preference parameters are then obtained, and based on the feature information and style preference parameters, a target travelogue is generated. The target travelogue is structured text and image content, and the style preference parameters characterize the user's desired style for the travelogue output.

[0006] This approach achieves comprehensive collection of multi-source heterogeneous data by acquiring images, videos, and metadata including shooting time and location during the trip. Based on this, multi-dimensional analysis is performed to extract feature information covering spatiotemporal dimensions, content importance dimensions, and emotional dimensions. This constructs a structured semantic representation with spatial path continuity, temporal sequence logic, visual focus, and subjective experience authenticity. Furthermore, by incorporating user-inputted style preference parameters, the generated travelogue not only accurately recreates the itinerary but also outputs text and image content with different narrative styles based on individual preferences. This overcomes the problems of fragmented content, templated descriptions, and lack of emotional expression and personalization found in traditional methods, significantly improving the intelligence level and narrative coherence of travelogue generation, thereby enhancing user experience satisfaction.

[0007] In one possible implementation, landscape elements in image and video data are detected using image content recognition algorithms. These landscape elements are then matched with landscape features in a pre-defined scenic spot database to determine a candidate scenic spot set. Based on the shooting location information and the candidate scenic spot set, a revised scenic spot set is generated. Based on the shooting time information and the revised scenic spot set, spatiotemporal information is generated, including a visualized journey path and the duration of stay at each scenic spot.

[0008] In another possible implementation, landscape elements are identified in image and video data, and the proportion of each landscape element in a single image is determined, where images are images from the image and video data. For each landscape element, a weight is determined based on the proportion of the landscape element in each image, the number of images containing the landscape element, and the total number of images, to obtain the weight of each landscape element.

[0009] In another possible implementation, weather feature recognition processing is performed on image and video data to obtain weather status labels. Facial expression feature recognition processing is then performed on the facial regions of people in the image and video data to obtain emotion status labels representing the emotional tendencies of individuals. Based on the weather status labels, weather factor values ​​matching the labels are determined; these weather factor values ​​indicate the degree of influence of weather conditions on the travel experience. Based on the emotion status labels, emotion factor values ​​matching the labels are determined; these emotion factor values ​​indicate the positivity of the individual's emotions. Based on the weather factor values ​​and emotion factor values, emotional dimension information is generated.

[0010] In another possible implementation, initial image data and initial video data are acquired, and their metadata is extracted. Data in the initial image data and initial video data that does not meet the processing conditions are filtered to obtain multi-source data. The processing conditions include at least one of the following: data with metadata and data with identifiable image content.

[0011] In another possible implementation, multiple preset style preference parameters are displayed; these parameters include at least one of the following: literary narrative, practical guide, emotional diary, or minimalist list. The system receives the user's selection and determines the style preference parameter from the multiple preset parameters. Alternatively, it obtains the style preference parameter input by the user.

[0012] In another possible implementation, a travelogue content framework is generated based on style preference parameters. This framework indicates the travelogue's structural form and language style. Based on the travelogue content framework and feature information, a sequence of travel images and a text travelogue are generated, respectively. Finally, a target travelogue is generated based on the travel image sequence and the text travelogue.

[0013] Secondly, this application provides a travelogue generation device, which includes an acquisition module, a processing module, and a display module.

[0014] The acquisition module acquires multi-source data during the trip, including image data, video data, and metadata. Metadata includes shooting time and location information. The processing module performs multi-dimensional analysis on the multi-source data to extract feature information, which includes at least one of the following: spatiotemporal dimension information, content importance dimension information, and sentiment dimension information. The spatiotemporal dimension information represents the spatial path and time sequence of the journey; the content importance dimension information represents the weight of landscape elements in the images; and the sentiment dimension information represents the emotional tendency of the travel experience. The processing module also acquires style preference parameters and, based on the feature information and style preference parameters, generates a target travelogue. The target travelogue is structured text and image content, and the style preference parameters represent the user's desired style for the travelogue output.

[0015] In one possible implementation, the processing module is further configured to detect landscape elements in image and video data using an image content recognition algorithm. The processing module is also configured to match the landscape elements with landscape features in a preset scenic spot database to determine a candidate scenic spot set. The processing module is further configured to generate a revised scenic spot set based on the shooting location information and the candidate scenic spot set. The processing module is further configured to generate spatiotemporal dimension information based on the shooting time information and the revised scenic spot set, including a visualized journey path and the duration of stay at each scenic spot.

[0016] In another possible implementation, the processing module is further configured to identify landscape elements in image and video data, and determine the proportion of each landscape element in a single image, where the image is an image in the image and video data. The processing module is also configured to, for each landscape element, determine its weight based on the proportion of the landscape element in each image, the number of images containing the landscape element, and the total number of images, to obtain the weight of each landscape element.

[0017] In another possible implementation, the processing module is further configured to perform weather feature recognition processing on the image and video data to obtain weather status labels. The processing module is also configured to perform facial expression feature recognition processing on the facial regions of people in the image and video data to obtain emotion status labels representing the emotional tendencies of the people. The processing module is further configured to determine weather factor values ​​matching the weather status labels, which indicate the degree of influence of weather conditions on the travel experience. The processing module is also configured to determine emotion factor values ​​matching the emotion status labels, which indicate the positiveness of the person's emotions. The processing module is further configured to generate emotional dimension information based on the weather factor values ​​and emotion factor values.

[0018] In another possible implementation, the processing module is further configured to acquire initial image data and initial video data, and extract metadata from the initial image data and initial video data. The processing module is also configured to filter data in the initial image data and initial video data that does not meet processing conditions to obtain multi-source data. The processing conditions include at least one of the following: data with metadata, and data with identifiable image content.

[0019] In another possible implementation, a display module is used to display multiple preset style preference parameters; these preset style preference parameters include at least one of the following: literary narrative, practical guide, emotional diary, and minimalist list. The acquisition module is also used to receive the user's selection and determine the style preference parameter from the multiple preset style preference parameters. Alternatively, the acquisition module can also be used to acquire the style preference parameter input by the user.

[0020] In another possible implementation, the processing module is further configured to generate a travelogue content architecture based on style preference parameters, the travelogue content architecture indicating the structural form and language style of the travelogue. The processing module is also configured to generate a travel image sequence and a text travelogue based on the travelogue content architecture and feature information, respectively. Finally, the processing module is configured to generate a target travelogue based on the travel image sequence and the text travelogue.

[0021] Thirdly, this application provides a travelogue generation apparatus, the apparatus comprising: a processor and a memory; the processor and the memory being coupled; the memory being used to store one or more programs, the one or more programs including computer device execution instructions, wherein when the travelogue generation apparatus is running, the processor executes the computer device execution instructions stored in the memory to implement the method as described in the first aspect and any possible implementation thereof.

[0022] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause a computer device to perform the methods described in the first aspect and any possible implementation thereof.

[0023] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run computer programs or instructions to implement the methods described in the first aspect and any possible implementation thereof.

[0024] Sixthly, this application provides a computer program product containing instructions that, when executed by a computer, cause the computer device to perform the methods described in the first aspect and any possible implementation thereof.

[0025] The technical problems that the travelogue generation device, computer equipment, computer storage medium, chip or computer program product can solve and the technical effects that can be achieved in the above solution can be found in the technical problems and effects solved in the first aspect above, and will not be repeated here. Attached Figure Description

[0026] Figure 1 A system architecture diagram of a travelogue generation system provided in this application embodiment; Figure 2 A flowchart illustrating a method for generating travelogues provided in this application embodiment; Figure 3 A flowchart illustrating another method for generating travelogues provided in this application embodiment; Figure 4 A schematic diagram of a travelogue generation device provided in an embodiment of this application; Figure 5 A schematic diagram of the structure of another travelogue generation device provided in an embodiment of this application; Figure 6 A conceptual partial view of a computer program product provided for an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] The terms “first” and “second” in the specification and claims of this application are used to distinguish different objects, rather than to describe a specific order of objects.

[0029] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include other steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or device.

[0030] Furthermore, in the embodiments of this application, the words "exemplary" or "for example" are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design that is described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design options. Specifically, the use of the words "exemplary" or "for example" is intended to present concepts in a concrete manner.

[0031] With the continuous development of digital tourism services, people generally record travel details by taking pictures, videos and other multimodal data during their travels. These materials have become important carriers for preserving travel memories and sharing travel experiences. Transforming scattered materials into coherent travelogues has gradually become one of the core needs of users.

[0032] Currently, travelogues can be generated semi-automatically or automatically by collecting user travel-related data and integrating materials using preset templates or basic data processing methods, saving users time and effort in manual writing. However, the travelogues generated by these solutions are too simple and lack adaptability to the materials, resulting in insufficient personalization and failing to meet users' needs for high-quality travel narratives.

[0033] To address the technical problems identified in the background section of this application, embodiments of this application provide a method for generating travelogues. This method involves acquiring multi-source data during the travel process, including image data, video data, and metadata, with the metadata including shooting time and location information. The multi-source data is then subjected to multi-dimensional analysis to extract feature information, which includes at least one of the following: spatiotemporal dimension information, content importance dimension information, and sentiment dimension information. The spatiotemporal dimension information characterizes the spatial path and time sequence of the journey, the content importance dimension information characterizes the weight of landscape elements in the images, and the sentiment dimension information characterizes the emotional tendency of the travel experience. Style preference parameters are obtained, and based on the feature information and style preference parameters, a target travelogue is generated. The target travelogue is structured text and image content, and the style preference parameters characterize the user's desired style for the travelogue output. This approach achieves comprehensive collection of multi-source heterogeneous data by acquiring images, videos, and metadata including shooting time and location during the trip. Based on this, multi-dimensional analysis is performed to extract feature information covering spatiotemporal dimensions, content importance dimensions, and emotional dimensions. This constructs a structured semantic representation with spatial path continuity, temporal sequence logic, visual focus, and subjective experience authenticity. Furthermore, by incorporating user-inputted style preference parameters, the generated travelogue not only accurately recreates the itinerary but also outputs text and image content with different narrative styles based on individual preferences. This overcomes the problems of fragmented content, templated descriptions, and lack of emotional expression and personalization found in traditional methods, significantly improving the intelligence level and narrative coherence of travelogue generation, thereby enhancing user experience satisfaction.

[0034] The implementation environment of the embodiments of this application is described below.

[0035] like Figure 1 As shown, this application provides a travelogue generation system, which includes: a data acquisition module 101, a multi-dimensional data analysis module 102, and a text generation module 103.

[0036] The data acquisition module 101 is used to collect and preprocess multi-source data to provide structured input for subsequent analysis. The data acquisition module 101 may include: a multimodal data acquisition submodule 1011, a data preprocessing submodule 1012, and a style preference input module 1013.

[0037] The multimodal data acquisition submodule 1011 is used to acquire user-uploaded images (JPEG / PNG) and videos (MP4 / MOV) through a device interface, or to obtain user-authorized data by calling third-party platform application programming interfaces (APIs) (such as social media). The multimodal data acquisition submodule 1011 can also be used to extract metadata from image and video data.

[0038] The data preprocessing submodule 1012 is used to filter out invalid data, such as data that does not meet the processing conditions. For example, it filters out data with missing metadata or data whose image content cannot be identified (such as blurry images). Furthermore, the data preprocessing submodule 1012 can also be used to sort the data according to the shooting time, forming a data sequence with a time dimension. The data preprocessing submodule 1012 can also be used to establish data for the same scene based on geographical proximity algorithms (such as spatial clustering).

[0039] The style preference input module 1013 provides an interactive interface for users to select preset style preference parameters (i.e., copywriting style). The preset style preference parameters include categories such as literary narrative, practical guide, emotional diary, and simple list. Users can customize keywords (such as "emphasizing history and culture" or "highlighting food experience") or upload sample copywriting as style reference.

[0040] The data multi-dimensional analysis module 102 is used to extract spatiotemporal, content, and sentiment three-dimensional features from preprocessed data and generate structured semantic packages. The data multi-dimensional analysis module 102 includes: spatiotemporal dimension analysis submodule 1021, content importance analysis submodule 1022, sentiment dimension analysis submodule 1023, and feature fusion submodule 1024.

[0041] The spatiotemporal dimension analysis submodule 1021 is used to detect landscape elements (such as ancient buildings, natural landscapes, etc.) contained in image or video data. Furthermore, the spatiotemporal dimension analysis submodule 1021 is also used to match landscape elements with a preset scenic spot database, and to correct positioning deviations (such as drift caused by vegetation occlusion) by combining the location information of the image data. The spatiotemporal dimension analysis submodule 1021 is also used to determine the dwell time at each scenic spot, such as dwell time = the difference between the latest and earliest shooting times of the same scenic spot footage.

[0042] The content importance analysis submodule 1022 is used to identify landscape elements (such as landmark buildings) in a single image, calculate the element pixel ratio (landscape element pixel area / total image pixels), and determine the weight of the landscape element among all landscape elements.

[0043] The sentiment dimension analysis submodule 1023 is used to identify weather state labels (such as sunny / rainy / snowy) through image classification and output sentiment state labels (such as happy / calm / depressed) based on facial region analysis. Then, sentiment dimension information is generated based on the weather state labels and sentiment state labels.

[0044] The feature fusion submodule 1024 is used to integrate the output content (i.e. feature information) of the spatiotemporal dimension analysis submodule 1021, the content importance analysis submodule 1022, and the sentiment dimension analysis submodule 1023, and generate a structured feature information set.

[0045] Optionally, the data multi-dimensional analysis module 102 also includes a humanities information parsing module, which is used to link to the local cultural database (such as introductions of historical and cultural sites and schedules of folk activities) based on geographic coordinates, and extract humanities elements by combining image content recognition (such as architectural style, text signs, and clothing) to generate semantic information of cultural tags (historical sites / markets / festivals, etc.) - historical background - folk characteristics.

[0046] The copywriting generation module 103 is used to integrate feature information and user style to generate target travelogues. The copywriting generation module 103 includes a content architecture design submodule 1031, a language generation submodule 1032, and an image and text integration submodule 1033.

[0047] The content architecture design submodule 1031 is used to generate the corresponding content architecture based on style preferences. The language generation submodule 1032 is used to generate text content according to style preferences.

[0048] The image and text fusion submodule 1033 is used to insert media materials (including images, videos, text, etc.) according to the content structure to generate the target travelogue.

[0049] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0050] like Figure 2 The image shows a method for generating travelogues according to an embodiment of this application. The method includes: S201. Obtain multi-source data during the travel process.

[0051] The multi-source data includes image data, video data, and metadata, with the metadata including shooting time information and shooting location information.

[0052] In one possible implementation, initial image data and initial video data are acquired, and metadata of the initial image data and initial video data is extracted. Data in the initial image data and initial video data that does not meet the processing conditions are filtered to obtain multi-source data. The processing conditions include at least one of the following: data with metadata and data with identifiable image content.

[0053] For example, the degree of blur in image or video data can be obtained, and the image content can be determined as to whether it can be recognized based on the degree of blur.

[0054] S202. Perform multi-dimensional analysis and processing on multi-source data to extract feature information.

[0055] The feature information includes at least one of the following: spatiotemporal dimension information, content importance dimension information, and sentiment dimension information. The spatiotemporal dimension information is used to characterize the spatial path and time sequence of the journey, the content importance dimension information is used to characterize the weight of landscape elements in the image, and the sentiment dimension information is used to characterize the emotional tendency of the travel experience.

[0056] In one possible implementation, multi-source data can be processed based on a multi-dimensional analytical model to generate feature information. The multi-dimensional analytical model is used to generate feature information.

[0057] It should be understood that this multi-dimensional analytical model can use a neural network model or the like as an initial model and be trained using a training set.

[0058] In another possible implementation, multiple agents can process multi-source data to generate feature information. These agents include at least one of the following: a spatiotemporal dimension agent, a content importance agent, and a sentiment dimension agent. The spatiotemporal dimension agent generates spatiotemporal dimension information, the content importance agent generates content importance dimension information, and the sentiment dimension agent generates sentiment dimension information.

[0059] The following section introduces another method for obtaining spatiotemporal dimension information, content importance dimension information, and sentiment dimension information.

[0060] In one possible implementation, landscape elements in image and video data are detected using image content recognition algorithms. These landscape elements are then matched with landscape features in a pre-defined scenic spot database to determine a candidate scenic spot set. Based on the shooting location information and the candidate scenic spot set, a revised scenic spot set is generated. Based on the shooting time information and the revised scenic spot set, spatiotemporal information is generated, including a visualized journey path and the duration of stay at each scenic spot.

[0061] For example, image data can be sorted by shooting time and duration to form a complete timeline. Secondly, using the built-in positioning function of the shooting device (such as a smartphone or action camera), the geographic coordinates of the image or video data can be extracted to initially determine the spatial distribution of travelers during their trip. Then, by analyzing the image and video data, landscape elements (such as landmarks, landscape features, or iconic objects) are identified, and the identification results are matched with a scenic spot database to determine the tourist attraction corresponding to the shooting location. If multiple possible matching attractions exist, the characteristics of the landscape elements and the temporal and positional information are combined to determine the tourist attraction corresponding to the shooting location.

[0062] This not only effectively corrects the deviations caused by positioning drift and the problem of inaccurate tourist attraction classification, but also improves the accuracy of attraction identification.

[0063] After obtaining the corrected sequence of attractions, a traveler's journey line can be formed, which is the route of attraction visits arranged in chronological order. This journey line reflects the traveler's spatial movement trajectory. Finally, the timeline and the journey line are fused to generate a travel heatmap (i.e., spatiotemporal dimension information). This heatmap, based on spatial location, maps the duration of stay to color changes, thus visually representing the traveler's itinerary. In the image presentation, darker colored areas correspond to attractions where the stay is longer, thereby intuitively displaying the traveler's activity hotspots and route characteristics.

[0064] In one possible implementation, multi-source data is subjected to multi-dimensional parsing to extract content importance dimension information, including steps 2021a-2022a.

[0065] Step 2021a: Identify landscape elements in image and video data and determine the proportion of each landscape element in a single image.

[0066] The image refers to the image data and the image in the video data.

[0067] For example, landscape elements include landmark buildings, natural landscapes, and special figures.

[0068] In one possible implementation, an element recognition agent can perform content analysis on image and video data one by one to identify the landscape elements contained in the image. Then, the pixel percentage of each landscape element in the image can be obtained; this pixel percentage is the area proportion of the landscape element region within the entire image.

[0069] Step 2022a: For each landscape element, determine the weight of the landscape element based on the proportion of the landscape element in each image, the number of images containing the landscape element, and the total number of images, so as to obtain the weight of each landscape element.

[0070] The weight of a landscape element is used to reflect its importance during the journey.

[0071] In one possible implementation, the number of all images containing the same landscape element is counted, and the frequency of occurrence of that landscape element is generated based on the total number of images. The average percentage of the landscape element's area in all images in which it appears is then averaged to generate the average percentage of the landscape element. Finally, the weight of the landscape element is generated based on its average percentage and frequency of occurrence.

[0072] It should be understood that video data is composed of multiple image frames. Therefore, the images in this application include image frames from the video data.

[0073] In one possible implementation, the weights of landscape elements can satisfy the following formula 1.

[0074] M1= Formula 1.

[0075] Where M1 represents the weight of the landscape element, and n represents the number of images containing that landscape element. The value is used to represent the proportion of a landscape element in the nth image containing that landscape element, and L is used to represent the total number of images.

[0076] Understandably, by introducing weights, we can highlight scenic elements that travelers pay more attention to, and thus achieve a more focused portrayal of the travel experience.

[0077] In one possible implementation, multi-source data is subjected to multi-dimensional parsing to extract sentiment dimension information, including steps 2021b-2025b.

[0078] Step 2021b: Perform weather feature recognition processing on the image data and video data to obtain weather status labels.

[0079] In one possible implementation, weather feature recognition processing is performed on image data and video data to obtain weather status labels corresponding to each shooting time. These weather status labels are used to characterize the type of external meteorological environment in which the user is recording their trip.

[0080] For example, image and video data can be parsed to obtain visual features and shooting parameters. Visual features include at least one of the following: color distribution characteristics, overall brightness level, sky area proportion and its texture pattern. Shooting parameters include at least one of the following: contrast, white balance shift, and shadow intensity parameters to determine the lighting conditions and atmospheric conditions of the current scene. Then, based on the shooting parameters, the lighting conditions and atmospheric conditions of the current scene are determined. Next, the aforementioned visual features, lighting conditions, and atmospheric conditions are input into a preset weather classification model for recognition, outputting weather state labels. For example, weather state labels may include sunny, cloudy, rainy, snowy, foggy, etc.

[0081] Optionally, a third-party meteorological service interface can be invoked based on location information to obtain the actual weather records of the scenic spot as auxiliary verification information. When there is a discrepancy between the visual recognition result and the external meteorological data, a weighted decision is made based on the confidence levels of both to ultimately determine a unified weather status label.

[0082] Step 2022b: Perform facial expression feature recognition processing on the facial regions of people in the image data and video data to obtain emotional state labels used to represent the emotional tendencies of people.

[0083] In one possible implementation, facial expression feature recognition processing can be performed on the facial regions of people in image data and video data to obtain the user's emotional state label corresponding to the image data or video data.

[0084] For example, a face detection algorithm can be used to process the facial region of a person and extract the distribution of facial key points, muscle movement units (such as the curvature of the corners of the eyes and the degree of upward movement of the corners of the mouth), and local texture change features. Then, the distribution of facial key points, muscle movement units, and local texture change features are input into a preset expression classification model to obtain an emotion state label.

[0085] Step 2023b: Based on the weather status label, determine the weather factor value that matches the weather status label.

[0086] Among them, the weather factor value is used to indicate the degree of impact of weather conditions on the travel experience.

[0087] In one possible implementation, the travelogue generation device stores a first correspondence, which is a correspondence between preset weather state labels and preset weather factor values. Based on the weather state labels and the first correspondence, the weather factor value matching the weather state label can be determined.

[0088] For example, when the weather status label is sunny, a higher weather factor value (e.g., 0.9) corresponds to favorable lighting and comfortable environmental conditions. When the weather status label is heavy rain or blizzard, a lower weather factor value (e.g., 0.1) corresponds to adverse travel conditions.

[0089] Step 2024b: Based on the emotional state label, determine the emotional factor value that matches the emotional state label.

[0090] Among them, the emotion factor value is used to indicate the positiveness of a person's emotions.

[0091] In one possible implementation, the travelogue generation device stores a second correspondence, which is a correspondence between emotion state tags and preset emotion factor values. Based on the emotion state tags and the second correspondence, the emotion factor value matching the emotion state tag can be determined.

[0092] For example, when the emotion status label reflects a more positive emotion, it corresponds to a higher emotion factor value (e.g., 0.9). When the emotion status label reflects a more negative emotion, it corresponds to a lower emotion factor value (e.g., 0.1).

[0093] Step 2025b: Generate emotional dimension information based on weather factor values ​​and emotion factor values.

[0094] In one possible implementation, weather factor values ​​and emotion factor values ​​are weighted and fused to generate emotional dimension information.

[0095] In some embodiments, multi-source data can be subjected to multi-dimensional parsing to extract initial feature information. Then, user-inputted adjustment parameters can be received to adjust the initial feature information. Finally, feature information is generated based on the initial feature information and the adjustment parameters.

[0096] For example, the initial spatiotemporal dimension information shows a travel path including attraction a-attraction b-attraction c, while the adjusted travel path includes attraction a-attraction c-attraction b. Alternatively, the initial weight of landscape 1 in the content importance dimension is 0.8, while the adjusted weight of landscape 1 is 0.7. The initial sentiment dimension information is 0.6 (e.g., happy), while the adjusted sentiment dimension information is 0.7 (e.g., active).

[0097] S203, Obtain style preference parameters.

[0098] Among them, the style preference parameter is used to characterize the user's desired travelogue output style.

[0099] In one possible implementation, multiple preset style preference parameters can be displayed; these parameters include at least one of the following: literary narrative, practical guide, emotional diary, or minimalist list. Then, the user's selection can be received, and the style preference parameter can be determined from the multiple preset parameters.

[0100] It should be understood that the number of preset style preference parameters selected in the selection operation can be one or more, and this application embodiment does not limit this.

[0101] In another possible implementation, the style preference parameters input by the user can be obtained.

[0102] For example, a user can input style preference parameters. Alternatively, a user can input historical travelogues, and the travelogue generation device can generate style preference parameters based on those historical travelogues.

[0103] S204. Generate the target travelogue based on feature information and style preference parameters.

[0104] Among them, the target travelogues are structured text and image content.

[0105] In one possible implementation, a text travelogue can be generated based on style preference parameters, sentiment dimension information, and multi-source data. This text travelogue describes the multi-source data according to the style preference parameters. Then, a sequence of travel images can be generated based on spatiotemporal dimension information and content importance dimension information, arranged according to the importance of the travel route and landscape elements.

[0106] Based on the above technical solution, multi-source data during the travel process is acquired. This multi-source data includes image data, video data, and metadata, with metadata including shooting time and location information. Multi-dimensional analysis is performed on the multi-source data to extract feature information, which includes at least one of the following: spatiotemporal dimension information, content importance dimension information, and sentiment dimension information. The spatiotemporal dimension information is used to represent the spatial path and time sequence of the journey; the content importance dimension information is used to represent the weight of landscape elements in the images; and the sentiment dimension information is used to represent the emotional tendency of the travel experience. Style preference parameters are obtained, and based on the feature information and style preference parameters, a target travelogue is generated. The target travelogue is structured text and image content, and the style preference parameters represent the user's desired style of the travelogue output. In this way, by acquiring images, videos, and metadata containing shooting time and location information during the travel process, comprehensive collection of multi-source heterogeneous data is achieved. Based on this, multi-dimensional analysis is performed to extract feature information covering spatiotemporal, content importance, and sentiment dimensions, thereby constructing a structured semantic representation with spatial path continuity, temporal sequence logic, visual focus, and subjective experience authenticity. Furthermore, by incorporating user-inputted style preference parameters, the generated travelogues can not only accurately recreate the itinerary but also output text and image content with different narrative styles based on individual preferences. This overcomes the problems of fragmented content, templated descriptions, and lack of emotional expression and personalization found in traditional methods, significantly improving the intelligence level and narrative coherence of travelogue generation, thereby enhancing user experience satisfaction.

[0107] like Figure 3 As shown, this is another method for generating travelogues provided in an embodiment of this application. In this method, S204 includes: S301. Generate the travelogue content structure based on style preference parameters.

[0108] Among them, the content structure of the travelogue is used to indicate the structural form and language style of the travelogue.

[0109] In this embodiment, the travelogue generation device stores a correspondence between preset style preference parameters and preset content structure. The travelogue content structure can be determined based on the style preference parameters and the correspondence between the preset style preference parameters and the preset content structure.

[0110] For example, when the style preference parameter is literary narrative, a narrative structure is generated based on a timeline, integrating environmental descriptions and subjective feelings, and a rhetorical vocabulary enhancement module is configured. When the style preference parameter is travel guide, an itemized structure is constructed according to themes (such as transportation, accommodation, and attractions), and an information density priority strategy is set. When the style preference parameter is emotional diary, a first-person perspective and a logical progression of emotions are used to arrange the paragraph order. When the style preference parameter is simple list, a key point outline is generated, and sentence length and the use of modifiers are controlled.

[0111] S302. Based on the content structure and feature information of the travelogue, generate a sequence of travel images and a text travelogue, respectively.

[0112] In one possible implementation, multi-source data is filtered and sorted based on pre-defined structural rules within the travelogue's content framework, combined with feature information. Images corresponding to high-weight landscape elements are prioritized and arranged according to shooting time sequence or theme category to generate a travel image sequence with narrative coherence. During text generation, using the travelogue's content framework as a framework, key time nodes, descriptions of important attractions, and corresponding emotional expressions are integrated. A natural language generation model is then invoked to output a text travelogue that is semantically consistent with the image sequence.

[0113] S303. Generate a target travelogue based on travel image sequences and text travelogues.

[0114] In one possible implementation, a sequence of travel images is inserted into the text travelogue according to corresponding time points or thematic paragraphs, forming a layout structure where images and text correspond. A semantically matched caption is generated for each inserted image; the caption is automatically generated based on the context of the paragraph in which the image is located and the image's own identification tags. The layout structure and captions are then integrated to obtain the target travelogue.

[0115] Optionally, users can set the word count for the target travelogue.

[0116] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. It is understood that the travelogue generation apparatus, in order to achieve the above-mentioned functions, includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the travelogue generation method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0117] This application also provides a travelogue generation device. This travelogue generation device can be a mobile terminal, a CPU within the mobile terminal, a generation module within the mobile terminal for generating travelogues, or a client application within the mobile terminal for generating travelogues.

[0118] This application embodiment can divide the travelogue generation device into functional modules or functional units according to the above method example. For example, each function can be divided into a separate functional module or functional unit, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or in software functional modules or functional units. The module or unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0119] This application provides an apparatus for generating travelogues. For example... Figure 4 As shown, the device for generating the travelogue may include: an acquisition module 401, a processing module 402, and a display module 403.

[0120] The acquisition module 401 is used to acquire multi-source data during the trip, including image data, video data, and metadata. The metadata includes shooting time and location information. The processing module 402 is used to perform multi-dimensional analysis on the multi-source data to extract feature information. The feature information includes at least one of the following: spatiotemporal dimension information, content importance dimension information, and sentiment dimension information. The spatiotemporal dimension information is used to represent the spatial path and time sequence of the journey; the content importance dimension information is used to represent the weight of landscape elements in the images; and the sentiment dimension information is used to represent the emotional tendency of the travel experience. The processing module 402 is also used to acquire style preference parameters and, based on the feature information and style preference parameters, generate a target travelogue. The target travelogue is structured text and image content, and the style preference parameters represent the user's desired style for the travelogue output.

[0121] In one possible implementation, processing module 402 is further configured to detect landscape elements in image data and video data using an image content recognition algorithm. Processing module 402 is also configured to match the landscape elements with landscape features in a preset scenic spot database to determine a candidate scenic spot set. Processing module 402 is further configured to generate a revised scenic spot set based on the shooting location information and the candidate scenic spot set. Processing module 402 is further configured to generate spatiotemporal dimension information based on the shooting time information and the revised scenic spot set, the spatiotemporal dimension information including a visualized journey path and the duration of stay at each scenic spot.

[0122] In another possible implementation, processing module 402 is further configured to identify landscape elements in image data and video data, and determine the proportion of each landscape element in a single image, where the image is an image in the image data and video data. Processing module 402 is also configured to, for each landscape element, determine the weight of the landscape element based on the proportion of the landscape element in each image, the number of images containing the landscape element, and the total number of images, to obtain the weight of each landscape element.

[0123] In another possible implementation, processing module 402 is further configured to perform weather feature recognition processing on image data and video data to obtain weather status labels. Processing module 402 is also configured to perform facial expression feature recognition processing on facial regions of people in image data and video data to obtain emotion status labels representing the emotional tendencies of people. Processing module 402 is further configured to determine weather factor values ​​matching the weather status labels, the weather factor values ​​indicating the degree of influence of weather conditions on the travel experience. Processing module 402 is further configured to determine emotion factor values ​​matching the emotion status labels, the emotion factor values ​​indicating the positiveness of the person's emotions. Processing module 402 is further configured to generate emotional dimension information based on the weather factor values ​​and emotion factor values.

[0124] In another possible implementation, processing module 402 is further configured to acquire initial image data and initial video data, and extract metadata from the initial image data and initial video data. Processing module 402 is also configured to filter the initial image data and initial video data that do not meet processing conditions to obtain multi-source data. The processing conditions include at least one of the following: data with metadata, and data with identifiable image content.

[0125] In another possible implementation, the display module 403 is used to display multiple preset style preference parameters; the multiple preset style preference parameters include at least one of the following: literary narrative type, practical strategy type, emotional diary type, and simple list type. The acquisition module 401 is also used to receive the user's selection operation and determine the style preference parameter from the multiple preset style preference parameters. Alternatively, the acquisition module 401 is also used to acquire the style preference parameter input by the user.

[0126] In another possible implementation, processing module 402 is further configured to generate a travelogue content structure based on style preference parameters, the travelogue content structure indicating the structural form and language style of the travelogue. Processing module 402 is also configured to generate a travel image sequence and a text travelogue based on the travelogue content structure and feature information, respectively. Processing module 402 is further configured to generate a target travelogue based on the travel image sequence and the text travelogue.

[0127] Figure 5This is a schematic diagram illustrating the structure of another travelogue generation apparatus according to an exemplary embodiment. The travelogue generation apparatus may include a processor 502, which executes application code to implement the travelogue generation method of this application.

[0128] The processor 502 may be a central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application.

[0129] like Figure 5 As shown, the travelogue generation device may further include a memory 503. The memory 503 stores the application code that executes the scheme of this application, and its execution is controlled by the processor 502.

[0130] Memory 503 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory 503 may exist independently and be connected to processor 502 via bus 504. Memory 503 may also be integrated with processor 502.

[0131] like Figure 5 As shown, the travelogue generation device may further include a communication interface 501, wherein the communication interface 501, processor 502, and memory 503 may be coupled to each other, for example, through a bus 504. The communication interface 501 is used for information exchange with other devices, for example, supporting information exchange between the travelogue generation device and other devices.

[0132] It should be pointed out that, Figure 5The device structure shown does not constitute a limitation on the device for generating this travelogue, except... Figure 5 In addition to the components shown, the travelogue generation device may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0133] In actual implementation, the functions implemented by the processing unit can be derived from... Figure 5 The processor 502 shown calls the program code in memory 503 to implement this.

[0134] This application also provides a computer-readable storage medium storing instructions that, when executed by a processor of a computer device, enable the computer to perform the travelogue generation method provided in the above-described embodiments. For example, the computer-readable storage medium may be a memory 503 including instructions, which may be executed by a processor 502 of a computer device to complete the above method. Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium, such as a ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0135] Figure 6 A conceptual partial view of a computer program product provided in an embodiment of this application is shown schematically. The computer program product includes a computer program for executing computer processes on a computing device.

[0136] In one embodiment, the computer program product is provided using a signal bearer medium 600. The signal bearer medium 600 may include one or more program instructions that, when executed by one or more processors, can provide the above-mentioned... Figure 2 , Figure 3 The described function or part of the function. Therefore, for example, refer to... Figure 2 In the embodiment shown, one or more features of S201-S204 can be fulfilled by one or more instructions associated with the signal carrying medium 600. Furthermore, Figure 6 The program instructions in the document also describe example instructions.

[0137] In some examples, the signal carrying medium 600 may include a computer-readable medium 601, such as, but not limited to, a hard disk drive, a compact disc (CD), a digital video optical disc (DVD), a digital magnetic tape, a memory, a read-only memory (ROM), or a random access memory (RAM), and so on.

[0138] In some implementations, the signal carrying medium 600 may include a computer recordable medium 602, such as, but not limited to, a memory, a read / write (R / W) CD, an R / W DVD, and so on.

[0139] In some implementations, the signal carrying medium 600 may include a communication medium 603, such as, but not limited to, digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.).

[0140] The signal-bearing medium 600 can be transmitted by a wireless communication medium 603. One or more program instructions may be, for example, computer-executable instructions or logical implementation instructions.

[0141] In some examples, the travelogue generation apparatus can be configured to provide various operations, functions, or actions in response to one or more program instructions via a computer-readable medium 601, a computer-recordable medium 602, and / or a communication medium 603.

[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0143] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0144] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0145] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0146] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0147] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating travelogues, characterized in that, The method includes: Acquire multi-source data during the travel process, including image data, video data, and metadata, including shooting time information and shooting location information; The multi-source data is subjected to multi-dimensional analysis to extract feature information, which includes at least one of the following: spatiotemporal dimension information, content importance dimension information, and sentiment dimension information. The spatiotemporal dimension information is used to characterize the spatial path and time sequence of the journey, the content importance dimension information is used to characterize the weight of landscape elements in the image, and the sentiment dimension information is used to characterize the emotional tendency of the travel experience. The style preference parameters are obtained, and a target travelogue is generated based on the feature information and the style preference parameters. The target travelogue is structured text and image content, and the style preference parameters are used to characterize the travelogue output style desired by the user.

2. The method according to claim 1, characterized in that, The feature information includes spatiotemporal dimension information; the multi-dimensional parsing process of the multi-source data to extract feature information includes: Landscape elements in the image data and video data are detected using an image content recognition algorithm; The landscape elements are matched with landscape features in a preset scenic spot database to determine a set of candidate scenic spots; Based on the shooting location information and the candidate scenic spot set, a revised scenic spot set is generated; Based on the shooting time information and the corrected set of attractions, the spatiotemporal dimension information is generated, which includes the visualized journey path and the duration of stay at each attraction.

3. The method according to claim 1, characterized in that, The feature information includes content importance dimension information; the multi-dimensional parsing process of the multi-source data to extract feature information includes: Identify landscape elements in the image data and video data, and determine the proportion of each landscape element in a single image, wherein the image is an image in the image data and video data; For each landscape element, the weight of the landscape element is determined based on the proportion of the landscape element in each image, the number of images containing the landscape element, and the total number of images, so as to obtain the weight of each landscape element.

4. The method according to claim 1, characterized in that, The feature information includes sentiment dimension information; the multi-dimensional parsing process of the multi-source data to extract feature information includes: Weather feature recognition processing is performed on the image data and the video data to obtain weather status labels; Facial feature recognition processing is performed on the facial regions of the people in the image data and video data to obtain emotional state labels that characterize the emotional tendencies of the people; Based on the weather status label, a weather factor value matching the weather status label is determined, and the weather factor value is used to indicate the degree of influence of weather conditions on the travel experience; Based on the emotional state label, an emotional factor value matching the emotional state label is determined, and the emotional factor value is used to indicate the positiveness of the person's emotions. The emotional dimension information is generated based on the weather factor value and the emotion factor value.

5. The method according to any one of claims 1-4, characterized in that, The acquisition of multi-source data during the travel process includes: Acquire initial image data and initial video data, and extract metadata from the initial image data and the initial video data; The initial image data and the initial video data that do not meet the processing conditions are filtered to obtain the multi-source data. The processing conditions include at least one of the following: data with the metadata and data with recognizable image content.

6. The method according to any one of claims 1-4, characterized in that, The process of obtaining style preference parameters includes: Displays multiple preset style preference parameters; the multiple preset style preference parameters include at least one of the following: literary narrative type, practical strategy type, emotional diary type, and minimalist list type; Receive the user's selection operation, and determine the style preference parameter from the plurality of preset style preference parameters; or... Obtain the style preference parameters input by the user.

7. The method according to any one of claims 1-4, characterized in that, The process of generating a target travelogue based on the feature information and the style preference parameters includes: Based on the style preference parameters, a travelogue content structure is generated, which indicates the structural form and language style of the travelogue. Based on the travelogue content structure and the feature information, a travel image sequence and a text travelogue are generated respectively; The target travelogue is generated based on the travel image sequence and the text travelogue.

8. A device for generating travelogues, characterized in that, The device includes: The acquisition module is used to acquire multi-source data during the trip. The multi-source data includes image data, video data, and metadata, including shooting time information and shooting location information. The processing module is used to perform multi-dimensional analysis processing on the multi-source data to extract feature information. The feature information includes at least one of the following: spatiotemporal dimension information, content importance dimension information, and sentiment dimension information. The spatiotemporal dimension information is used to characterize the spatial path and time sequence of the journey. The content importance dimension information is used to characterize the weight of landscape elements in the image. The sentiment dimension information is used to characterize the emotional tendency of the travel experience. The processing module is further configured to obtain style preference parameters and generate a target travelogue based on the feature information and the style preference parameters. The target travelogue is structured text and image content, and the style preference parameters are used to characterize the user's desired travelogue output style.

9. A device for generating travelogues, characterized in that, include: Processor and memory; The processor and the memory are coupled; The memory is used to store one or more programs, the one or more programs including computer device execution instructions, wherein when the travelogue generation device is running, the processor executes the computer device execution instructions stored in the memory to cause the travelogue generation device to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium storing instructions, characterized in that, When the computer device executes the instruction, the computer device performs the method as described in any one of claims 1-7.