Image recognition and large language model-based tourism memory generation method and system

CN122529928APending Publication Date: 2026-08-07LIAONING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAONING UNIVERSITY
Filing Date
2026-06-03
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0010]为了克服现有技术中存在的图像识别与文旅故事生成脱节、大语言模型自由生成容易发散、生成内容与游客真实照片不一致、缺少图像质量门控和内容一致性校验等问题,本发明提供一种基于图像识别和大语言模型的文旅记忆生成方法和系统

Benefits of technology

第一,本发明通过图像识别与结构化标签生成,使游客照片中的人物、动作、表情和景区背景能够作为故事生成的事实依据,解决了现有文旅故事生成与照片内容脱节的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122529928A_ABST
    Figure CN122529928A_ABST
Patent Text Reader

Abstract

The image recognition and large language model-based travel memory generation method and system relate to the technical field of artificial intelligence content generation. The invention faces the scene of visiting scenic spots, receives the information and travel images of tourists through a small program as a tourist end interaction entrance, judges whether the image meets the recognition and generation conditions, extracts the person, action, expression, scene and scenic spot node label in the image through the image recognition and structured label generation module, matches the image label with the scenic spot route, role identity, node task and cultural theme, generates the travel memory text corresponding to the photo and the culture of the scenic spot, checks the generated content, and finally saves and displays the generated results through the achievement storage and feedback optimization module. The invention can improve the matching degree between the travel story generation result and the real photo of the tourists and the cultural resources of the scenic spot, and reduce the content deviation, scenic spot mismatch and fictional errors caused by the free generation of the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical fields of smart cultural tourism, image recognition and artificial intelligence-generated content, and in particular relates to a method and system for generating cultural tourism memories based on image recognition and large language models. Background Technology

[0002] With the changing patterns of cultural and tourism consumption, tourists' demands for scenic spot experiences have gradually shifted from traditional sightseeing, written explanations, and photo ops to immersive experiences, interactive participation, personalized souvenirs, and social sharing. Especially in historical and cultural sites and natural landscapes, tourists not only want to receive introductions to the attractions, but also hope to combine personal photos, travel routes, cultural stories, and travel emotions to create digital travel memories that can be preserved and disseminated for a long time.

[0003] Existing digital service systems for scenic spots typically focus on providing information about attractions, map navigation, ticket reservations, activity displays, order verification, and travel guides. While some scenic spot mini-programs or mobile applications offer services such as image uploading, album saving, travelogue editing, or intelligent copywriting generation, these services mostly remain at the level of general content generation and struggle to fully integrate the actions, expressions, backgrounds, scenic spots, and specific tour routes in tourists' real photos.

[0004] For the creation of cultural and tourism memories, the travel photos uploaded by tourists are not ordinary pictures, but rather contain a variety of visual evidence related to the travel experience. For example, the photos may contain information such as the number of tourists, people's postures, gestures, facial expressions, clothing and props, architectural backgrounds, mountains and trees, waterways and bridges, text signs, shooting angles, and lighting atmosphere. This visual information can reflect the tourists' real experience at a particular scenic spot and is an important basis for generating personalized cultural and tourism stories.

[0005] However, existing cultural and tourism content generation solutions often suffer from a disconnect between image recognition and story generation. Some systems generate text based solely on user-input keywords or scenic spot names, failing to adequately identify the specific content in tourist photos; while others can recognize image content, the recognition results are not structurally matched with scenic routes, cultural themes, character tasks, and story templates, resulting in a lack of correspondence between the generated story content and tourist photos.

[0006] Large language models can generate coherent, rich, and narrative text content. However, if tourist photos or simple text descriptions are directly given to large language models for free generation, problems such as content digression, mismatched attractions, mismatched actions, distorted cultural expression, fabricated non-existent landscapes, incorrect number of characters, and inconsistent story style can easily occur. Especially in specific scenic area scenarios, if there are no constraints on scenic area nodes, cultural themes, and prohibited content, large language models may generate content that is inconsistent with the actual cultural resources of the scenic area.

[0007] Existing image-to-text systems generally lack quality gating mechanisms. Photos taken by tourists in scenic areas often suffer from issues such as blurriness, underexposure, backlighting, closed eyes, subject occlusion, incomplete backgrounds, repeated shots, or irrelevant content to the target scenic area. If the system does not assess image quality and usability, and directly inputs poor-quality or insufficiently supported images into subsequent recognition and generation stages, it will reduce the accuracy of image recognition and further affect the stability and credibility of the generated story.

[0008] Furthermore, existing systems typically lack consistency verification mechanisms for generated content. Even if a large language model can output a fluent story, it cannot guarantee that the number of characters, their actions, the scenic area background, cultural meanings, and route nodes described in the story are consistent with the photographic evidence. For cultural tourism scenarios, generated content must not only be visually appealing but also match the actual experience of tourists and the culture of the scenic area; otherwise, it will affect the formation of tourists' cultural memories of the scenic area.

[0009] Therefore, there is an urgent need for a method and system for generating cultural and tourism memories that can combine image quality judgment, image recognition, structured tag generation, cultural and tourism resource matching, controlled generation of large language models, and content consistency verification, so that the travel photos uploaded by tourists can be transformed into personalized cultural and tourism stories that are consistent with real photos, scenic routes, and cultural themes. Summary of the Invention

[0010] To overcome the problems existing in the prior art, such as the disconnect between image recognition and cultural tourism story generation, the tendency of large language models to diverge easily, inconsistencies between generated content and real tourist photos, and the lack of image quality gating and content consistency verification, this invention provides a method and system for generating cultural tourism memories based on image recognition and large language models.

[0011] This invention is achieved through the following technical solution: A cultural tourism memory generation system based on image recognition and large language model includes a mini-program interaction module, a cultural tourism resource database construction module, an image acquisition and quality gating module, an image recognition and structured tag generation module, a cultural tourism route matching module, a large language model controlled generation module, a content consistency and compliance verification module, and a result storage and feedback optimization module.

[0012] The mini-program's interactive module provides tourists with various entry points, including login, homepage browsing, route selection, role viewing, reservation and order placement, order viewing, offline verification, photo upload, AI story generation, electronic album viewing, result saving, result sharing, and result deletion. After entering the system through the mini-program, tourists can select story routes within the Yiwulü Mountain Scenic Area and upload travel photos after completing their tour or taking pictures. The system then generates corresponding cultural and tourism memory stories based on these photos.

[0013] The cultural tourism resource database construction module is used to establish structured cultural tourism knowledge data related to scenic spots. The cultural tourism resource database includes at least a scenic spot node database, a story route database, a character identity database, an action task database, a cultural theme database, a story template database, and a prohibited content database. The scenic spot node database stores the scenic spot number, name, geographical location, visual features, cultural meaning, and recommended shooting actions; the story route database stores the route number, name, age range, keywords, story background, and route node order; the character identity database stores the character names, descriptions, and behavioral characteristics that tourists can choose or match on the route; the action task database stores the photo actions, interactive tasks, and trigger conditions corresponding to the route nodes; the cultural theme database stores narrative themes related to the Yiwulü Mountain culture; the story template database stores the story structure under different routes, different characters, and different nodes; and the prohibited content database restricts inappropriate content, excessively fictional content, and content inconsistent with the scenic spot's culture.

[0014] The image acquisition and quality gating module receives travel photos uploaded by tourists and performs usability assessments on these photos. The quality gating includes, but is not limited to, sharpness detection, exposure detection, person integrity detection, closed eyes detection, occlusion detection, subject proportion detection, duplicate frame detection, and image format detection. When an image does not meet a preset quality threshold, the system prompts the tourist to retake the photo or re-upload it; when the image meets the quality threshold, the system passes the image to the subsequent recognition module.

[0015] The image recognition and structured label generation module is used to recognize travel photos that have passed quality gating, extract visual information from the images, and generate structured image labels. The structured image labels include labels for the number of people, their locations, actions, expressions, clothing or props, background elements, candidate scenic spots, text labels, lighting and atmosphere, and image confidence scores. The image recognition can be implemented using one or more combinations of object detection models, image classification models, pose recognition models, text recognition models, or multimodal vision models.

[0016] The cultural and tourism route matching module is used to match the structured image tags with the route information, character information, scenic spot node information, action task information, and cultural motif information selected by the tourist, so as to obtain the cultural and tourism story branch that best matches the current image. The matching process can be implemented by one or more combinations of rule matching, weighted scoring, vector similarity calculation, or large language model-assisted determination.

[0017] The cultural and tourism route matching module can make judgments according to the following matching scoring method: S = αA + βB + γC + δD + εE.

[0018] Among them, S represents the comprehensive matching score, A represents the matching degree between the image node features and the scenic spot node library, B represents the matching degree between the character actions and the action task library, C represents the matching degree between the tourist-selected route and the story route library, D represents the matching degree between the expression or atmosphere and the cultural motif library, E represents the matching degree between the character identity and the story template library, and α, β, γ, δ, ε are preset weight coefficients. The system selects the target route, target node, and target story template according to the comprehensive matching score. When the highest matching score is lower than the preset threshold, the system enters the fallback generation process.

[0019] The large language model controlled generation module is used to construct a controlled generation prompt package according to the structured image tags, route matching results, scenic spot culture constraints, story templates, output format requirements, and prohibited content constraints, and call the large language model to generate cultural and tourism memory texts. The controlled generation prompt package includes system role descriptions, image fact packages, route fact packages, node fact packages, cultural motif constraints, story style constraints, prohibited generation content, output format requirements, and generation length requirements.

[0020] The image fact package in the controlled generation prompt package can be represented in a structured format, for example: { "person_count": "2", "action_tags": ["group photo", "wave"], "expression_tags": ["smile"], "background_tags": ["bridge", "trees", "mountain"], "candidate_node": "Holy Water Bridge", "route_id": "R2", "role_name": "Mountain Guard Companion", "cultural_motif": "Guardianship, companionship, prayer", "confidence": "0.86" }

[0021] This structured fact package is used to restrict the large language model to generate content only based on the identified facts and matched cultural and tourism resources, thus preventing the large language model from making unfounded expansions that are detached from the actual situation of the photos and scenic spots.

[0022] The content generated by the controlled generation module of the large language model can be output in a fixed JSON format, for example: { "title": "Story title", “summary”: “story summary” "content": ["First story content", "Second story content", "Third story content"], "route_id": "route number", “node_id”: “Node number”, “tags”: [“Cultural Tag 1”, “Cultural Tag 2”] }

[0023] By limiting the output format, the system can reliably parse the results returned by the large language model and store the title, body text, line number, node number, and cultural tags into the database or display page respectively.

[0024] The content consistency and compliance verification module is used to verify the cultural and tourism memory content generated by the large language model. The verification includes image consistency verification, node consistency verification, character consistency verification, action consistency verification, cultural consistency verification, format consistency verification, and prohibited content verification. If the generated content contains characters, actions, landscapes, or scenic area nodes that are not present in the photos, or contains content that is inconsistent with, inappropriate for, or inaccurate to the cultural theme, the system determines that the verification fails.

[0025] When the generated content fails validation, the system performs different processing methods depending on the reason for the failure. If it is a minor expression problem, the system rewrites the relevant paragraphs; if it is a fact mismatch problem, the system reconstructs the prompt package and requires the large language model to regenerate based on the original structured tags; if it is a problem of insufficient image evidence, the system uses a conservative template to generate a short story; if multiple generation attempts still fail, the system prompts the visitor to re-upload photos or select a manual template.

[0026] The results storage and feedback optimization module is used to save the cultural and tourism memories generated by tourists. These results include at least a story title, story text, travel photo URLs or image data, generation time, route number, node number, character information, cultural tags, generation status, and user feedback information. Tourists can view historical AI albums, read story content, save stories, share results, or delete stories on the mini-program. The system can optimize subsequent route matching weights, story templates, and generation parameters based on tourists' saving, deleting, regenerating, and feedback behaviors.

[0027] This invention also provides a method for generating cultural tourism memories based on image recognition and a large language model, comprising the following steps: Step 1, tourists log in through a mini-program and select a scenic spot story route or role; Step 2, tourists upload travel photos; Step 3, the system performs quality gating on the travel photos; Step 4, the system recognizes the images that pass the quality gating and generates structured tags; Step 5, the system matches scenic spot routes, nodes, roles, and cultural themes according to the structured tags; Step 6, the system constructs a controlled generation prompt package and calls the large language model to generate a cultural tourism story; Step 7, the system performs consistency and compliance verification on the generated content; Step 8, the system returns the verified story content to the mini-program for display and saves it to an electronic photo album.

[0028] Compared with the prior art, the beneficial effects of the present invention are as follows: First, this invention uses image recognition and structured tag generation to enable the people, actions, expressions, and scenic backgrounds in tourist photos to serve as factual evidence for story generation, thus solving the problem of existing cultural tourism story generation being disconnected from photo content.

[0029] Secondly, this invention establishes a correspondence between image recognition results and scenic spot nodes, story routes, character identities, action tasks, and cultural themes through a cultural tourism resource database and route matching mechanism, so that the generated content can revolve around real scenic spot resources, reducing the probability of mismatched attractions, vague stories, and distorted cultural expression.

[0030] Third, this invention limits the free divergence of the large language model by controlling the generation of prompt packages, low divergence generation parameters, fixed output format, and disabling content constraints. This prevents the large language model from creating freely without constraints, and instead generates cultural and tourism memory content within the scope of photographic facts, route facts, and cultural facts.

[0031] Fourth, this invention uses a content consistency and compliance verification mechanism to conduct secondary checks on the number of characters, action descriptions, scene nodes, and cultural themes in the generated content, thereby reducing fictional, erroneous, and inappropriate content and improving the reliability of the generated story.

[0032] Fifth, this invention forms a closed loop through the results storage and feedback optimization module, enabling tourists to save and share generated results. The system can continuously optimize matching rules, generation templates, and quality thresholds based on user feedback, thereby improving the stability and personalization of subsequent generated results. Attached Figure Description

[0033] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0034] Figure 1 This is a schematic diagram of the system structure of the cultural tourism memory generation method and system based on image recognition and large language model of the present invention. Figure 2 This is a flowchart of the cultural tourism memory generation method and system based on image recognition and large language model of the present invention; Figure 3 This is a flowchart of the controlled generation and consistency verification process of the large language model of this invention. Detailed Implementation

[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.

[0036] Example 1: like Figure 1 As shown, a cultural tourism memory generation system based on image recognition and large language models is presented. The system includes a mini-program interaction module, a cultural tourism resource database construction module, an image acquisition and quality gating module, an image recognition and structured tag generation module, a cultural tourism route matching module, a large language model-controlled generation module, a content consistency and compliance verification module, and a results storage and feedback optimization module. These modules are connected via data interfaces, enabling images uploaded by tourists to undergo recognition, matching, generation, verification, and storage, ultimately forming displayable and storable cultural tourism memory results.

[0037] In one implementation, the mini-program interaction module runs within the WeChat mini-program environment as the visitor's entry point. Upon first entering the mini-program, visitors can select an avatar, enter a nickname, and complete the login process. After logging in, visitors are taken to the homepage to browse related story routes, activity highlights, and recommended experiences for the Yiwulü Mountain Scenic Area. Visitors can access the route list page to view the titles, recommended age ranges, keywords, story backgrounds, core characters, and additional services of different story routes, and select the appropriate route based on their own or their group's needs.

[0038] Once a visitor selects a story route, the mini-program's interaction module requests detailed information about that route from the backend. The backend retrieves the route number, route name, story background, age range, keywords, character list, and available services from the story route library and returns this information to the mini-program. Visitors can then select a core experience character based on the character descriptions on the page and initiate a booking. This process allows the visitor's route selection and character information to serve as constraints in subsequent story generation.

[0039] After tourists complete their visit to the scenic area, interactive tasks, or photo shoots, they enter a dedicated documentary or AI photo album page within the mini-program, where they can choose between an e-storybook or AI story generation entry point. The mini-program utilizes the local photo album or camera capabilities to receive the travel photos selected or taken by the tourist and uploads them to the backend service. Upon receiving the image file, the backend first checks if the file type is a valid image format. If the file type is invalid, an error message is returned; if the file type is valid, the image acquisition and quality gate control process begins.

[0040] The image acquisition and quality gating module first reads the basic information of the uploaded image, including image format, image size, file size, color channels, and shooting direction. Then, the system performs sharpness detection, exposure detection, and subject visibility detection. Sharpness detection can be achieved through image edge gradients, Laplacian variance, or depth model scoring; exposure detection can be achieved through brightness histograms and the proportion of overexposed and underexposed pixels; subject visibility detection can be achieved through the proportion of the person detection box, the integrity of human key points, and the degree of occlusion.

[0041] During the quality gating process, if the system detects that the image is severely blurry, the subject has closed eyes, the subject is occluded, the subject is too small, the background is unrecognizable, the exposure is severely abnormal, or the quality of repeated frames is low, the system will send a prompt to the mini-program to re-upload or re-shoot. If the system detects that the image quality meets the preset threshold, the image will be input into the image recognition and structured label generation module. This quality gating process can reduce the impact of low-quality images on subsequent recognition and generation results.

[0042] The image recognition and structured label generation module performs multi-level recognition on travel photos. First, the system identifies whether people are present in the image and determines the number, location, and bounding box of each person. Second, the system recognizes people's actions, such as standing, posing for photos, waving, pointing at scenery, walking, watching, praying, bending down to observe, and parent-child interaction. Third, the system recognizes people's facial expressions and overall emotions, such as smiling, calmness, surprise, focus, and interaction. Finally, the system identifies background elements, such as mountains, trees, bridges, water bodies, buildings, steps, stone tablets, plaques, ancient pine trees, squares, roads, or cultural installations.

[0043] In a preferred embodiment, the image recognition and structured label generation module can also combine text recognition technology to identify scenic area signs, plaque text, or guide sign text in the image. When text related to a scenic area node is identified, the system uses that text as important evidence for matching the scenic area node. When there are no obvious text markers in the image, the system infers candidate scenic area nodes based on building outlines, background elements, shooting angle, and the tourist's current route information.

[0044] After image recognition is completed, the system organizes the recognition results into structured image tags. These structured image tags are not simple natural language descriptions, but rather data structures that can be read and computed by subsequent modules. For example, for a photo of tourists taking a group photo by a bridge, the system can generate the following tags: number of people: 2; action tags: posing for a photo and waving; expression tag: smiling; background elements: bridge, trees, and water; node candidate: Sacred Water Bridge; image quality: acceptable; narrative tendency: traveling together, praying, and protecting. These structured tags can serve as the factual basis for subsequent route matching and story generation.

[0045] The cultural tourism resource database module pre-builds structured data related to the Yiwulü Mountain scenic area. In the scenic area node database, each node is assigned a node number, node name, node type, geographical location, visual characteristics, cultural meaning, suitable actions, recommended shooting angles, and related story segments. In the story route database, each route is assigned a route number, route name, route keywords, appropriate age group, story background, node order, and recommended roles. In the role identity database, each role is assigned a role name, role positioning, behavioral characteristics, narrative perspective, and related story templates.

[0046] For example, the system can define one story route as a "Mountain Recognition Journey," with keywords including mountain recognition, mountain protection, sense of etiquette, and exploration; another story route can be defined as a "Guardian Journey," with keywords including protection, companionship, family and country, and prayer. Different routes correspond to different nodes, action tasks, and cultural themes. When a tourist selects a route and uploads photos, the system will prioritize matching nodes and story templates within that route; if the photo content does not completely match the tourist's selected route, the system will decide whether to maintain the original route, switch to a more matching route, or enter a fallback generation process based on the matching score.

[0047] After receiving structured image tags, the cultural tourism route matching module matches them with node data, route data, character data, action task data, and cultural theme data in the cultural tourism resource database. During matching, the system calculates the matching degree between image background and scenic spot nodes, character actions and node tasks, the matching degree between the tourist's selected route and image semantics, the matching degree between character emotions and cultural themes, and the matching degree between the tourist's role and story template. Based on the comprehensive matching results, the system outputs the target route, target node, target role, and target story template.

[0048] In one implementation, the system can set the following matching logic: if the image contains bridges, water bodies, trees, and a couple taking a photo together, and the theme of the route selected by the tourist includes traveling together, protecting, or praying, then the system increases the matching weight of "bridge-type nodes" and "traveling together story templates"; if the image contains mountains, steps, an upward-looking action, and a single person standing, then the system increases the matching weight of "mountain recognition-type nodes" and "exploration-type story templates"; if the image contains parent-child interaction, pointing at the landscape, and a smiling expression, then the system increases the matching weight of "parent-child study tour story templates".

[0049] Once the route matching results are generated, the controlled generation module of the large language model begins constructing a cue package. This cue package does not simply provide the model with tourist photos for free interpretation; instead, it integrates image recognition results, route matching results, scenic area cultural constraints, and story templates into a structured input that the model can understand. The cue package includes at least the following: system role descriptions, facts about tourist photos, facts about scenic area nodes, route facts, role facts, cultural themes, output format, prohibited items, and generation style.

[0050] In one implementation, system role descriptions are used to limit the identity and writing scope of the large language model. For example, the model, acting as a cultural tourism documentary story generator, is required to generate content only based on input image facts and scenic area cultural facts. Tourist photo facts inform the model of the real content already identified in the photos, such as the number of people, their actions, expressions, and background elements. Scenic area node facts tell the model which node the current story should revolve around. Route facts tell the model which experience route the story belongs to. Role facts tell the model the tourist's identity in the story. Cultural motifs limit the story's thematic theme. Prohibited items explicitly state that the model must not fabricate non-existent attractions, change the number of people, fabricate actions not shown in the photos, or generate content inconsistent with the scenic area's culture.

[0051] To further improve the stability of the generated results, the system can restrict the output format of the large language model to a JSON structure, requiring the model to output a title, summary, array of body text, line numbers, node numbers, and tags. Because the output structure is fixed, the backend can automatically parse the results returned by the large language model. If the model returns a result that is not valid JSON, the system can require the model to re-output or initiate a format correction process. This approach avoids the problem of the mini-program failing to display due to unstable output format of the large language model.

[0052] After the large language model generates an initial draft, the content consistency and compliance verification module checks the generated content. For example... Figure 3 As shown, the system first checks whether the generated content is in a valid structured format; then it checks whether the story title and body text are empty; next, it checks whether the number of characters appearing in the body text matches the image recognition results; then it checks whether the actions described in the body text come from image action tags or action task libraries; next, it checks whether the scenic spot nodes described in the body text match the node matching results; and finally, it checks whether the content contains prohibited words, sensitive content, obviously fictional scenic spots, or expressions that are inconsistent with cultural themes.

[0053] If the system detects that the story describes "tourists ringing a bell in front of an ancient temple," but the image recognition results do not identify a temple, bell tower, or bell-ringing action, and the current matching node does not contain this element, the system determines it as a fact mismatch. If the system detects that the story contains the name of an out-of-town attraction or a fictional scenic area unrelated to the main theme of Yiwulü Mountain culture, the system determines it as a node mismatch. If the system detects that the story contains a third person who is not in the photo, the system determines it as a person count mismatch. In the above cases, the system reconstructs the prompt package and adds stronger prohibition constraints to the prompt package, requiring the large language model to regenerate.

[0054] If the generated content only has minor issues such as repetitive sentences, inconsistent style, excessively long paragraphs, or unnatural expressions, the system can rewrite specific paragraphs instead of regenerating the entire text. If the image evidence is insufficient, such as a photo with an overly blurry background that only identifies people and some mountain elements, the system can enter a conservative generation mode, generating only short documentary text that does not involve specific node names, and indicating that the content is generated based on limited image evidence. By distinguishing between different error types, the system can improve generation efficiency and reduce the cost of repeatedly calling large language models.

[0055] In one implementation, the controlled generation and consistency verification process of the large language model is as follows: The system receives structured image tags; determines whether the confidence level of the image tags reaches a threshold; if it does not reach the threshold, it enters the fallback template; if it reaches the threshold, it constructs a controlled generation prompt package; it calls the large language model to generate a JSON story; it parses the JSON story; it performs format verification; it performs image fact consistency verification; it performs scenic spot node consistency verification; it performs cultural theme consistency verification; it performs prohibited content verification; if all verifications pass, the generated content is returned to the mini-program; if any key verification fails, it performs partial rewriting, regeneration, or template downgrading according to the error type.

[0056] The results storage and feedback optimization module saves verified story content to the backend database. Saved content includes story ID, user ID, story title, cover image, main text paragraphs, route number, node number, character information, cultural tags, generation time, and generation status. After receiving the generated results, the mini-program displays travel photos, story title, story subtitle, and multiple story paragraphs to tourists. Tourists can click the save button to save the current AI story to their electronic photo album; they can also view the historical story list, swipe left and right to browse different stories, share their results, or delete a story.

[0057] In one implementation, the results storage and feedback optimization module also records the user's actions on the generated results. If a user saves a story, the system can assume that the story matches the user's needs highly; if a user regenerates a story multiple times, the system can assume that the current matching result or generation style needs adjustment; if a user deletes a story, the system can reduce the weight of the corresponding template or generation parameters in similar scenarios. By statistically analyzing user feedback, the system can gradually optimize the quality threshold, matching weight, prompt package template, and story template.

[0058] like Figure 2 As shown, the method flow of this invention can be summarized as follows: Tourists enter the mini-program and complete the login; tourists browse the story routes and select routes or roles; tourists complete the experience at scenic spot nodes and upload travel photos; the system performs format detection and quality gating on the photos; the system performs image recognition on qualified photos and generates structured tags; the system matches scenic spot routes, nodes, roles, and cultural themes according to the structured tags; the system constructs a controlled generation prompt package and calls a large language model to generate stories; the system performs consistency and compliance verification on the stories; the system returns the verified stories to the mini-program for display; after tourists save the stories, the system writes the stories to an electronic album and supports subsequent viewing, sharing, and deletion.

[0059] In group visitor scenarios, multiple tourists can take group photos, engage in interactive tasks, or role-play at the same scenic spot location. After identifying the number of people, their formation, interactive actions, and background elements in the photo, the system can generate story content themed around teamwork, shared protection, or collective exploration. For group photos, the system focuses on checking the number of people, interactive actions, and team titles during consistency verification to avoid generating content that does not match the actual number of people.

[0060] In family travel scenarios, the system can match family-oriented study tour story templates based on the relative positions of adults and children, actions such as holding hands, pointing, accompanying, and explaining. The generated content can emphasize joint exploration, nature observation, cultural enlightenment, and parent-child companionship, but it still needs to be constrained by scenic spot nodes and image facts, and cannot generate interactive behaviors that do not appear in the photos out of thin air.

[0061] In study tour application scenarios, the system can match study tour explanation story templates based on the group, the guide's posture, observation actions, recording actions, scenic spot signs, and node backgrounds. The cultural tourism memory content generated by the system can lean towards knowledge-based, task-oriented, and summary-based expression, combining tourists' observation behaviors at nodes with the Yiwulü Mountain cultural theme to form electronic story content suitable for preservation and display during study tour activities.

[0062] In scenarios involving individual tourists, the system can generate personalized travelogue text in the first or third person based on the posture, expression, shooting background, and the tourist's chosen route and role in a single photo. If the tourist photo only shows natural scenery and the person is not prominent, the system can generate a landscape documentary-style story; if the person is prominent in the tourist photo, the system generates a story centered on the person's experience. In this way, the system can adapt to different tourists' shooting habits and travel styles.

[0063] This invention does not require all embodiments to use the same image recognition model or the same large language model. In actual deployment, image recognition can employ open-source object detection models, commercial visual recognition interfaces, pose recognition models, OCR models, or multimodal large models; large language models can employ cloud-based model interfaces, locally deployed models, or hybrid model services. As long as they can achieve image label extraction, line matching, controlled generation, and consistency verification, they can all be considered equivalent implementations of this invention.

[0064] This invention does not limit the specific number and names of fields in the cultural tourism resource database. In different scenic areas, the scenic area node database, route database, character database, action task database, cultural theme database, and story template database can all be adjusted according to the characteristics of the scenic area. For example, in mountainous scenic areas, key node features can be set such as mountains, forests, steps, temples, inscriptions, and viewing platforms; in ancient city scenic areas, key node features can be set such as city gates, streets and alleys, residences, intangible cultural heritage workshops, and historical figures. The above adjustments do not affect the core technical solution of this invention, which generates cultural tourism memory content through image recognition and large language models.

[0065] The core of this invention lies not in simply sending tourist photos to a large language model to generate text freely, but in forming a complete closed loop through image quality gating, image recognition, structured tag generation, cultural and tourism resource matching, controlled prompt package construction, large language model generation, content consistency verification, and result storage and feedback. This closed loop enables the visual facts in tourist photos to be transformed into controllable cultural and tourism narrative evidence, and enables scenic area cultural resources to be transformed into personalized travel memory outcomes.

[0066] Those skilled in the art will understand that the modules and method steps disclosed herein can be implemented in software, hardware, or a combination of both. When implemented in software, the relevant steps can be written as a computer program and stored in a computer-readable storage medium, and executed by a processor to achieve the corresponding function; when implemented in hardware, they can be implemented jointly by servers, mobile terminals, image acquisition devices, storage devices, network communication devices, and artificial intelligence computing devices, etc.

[0067] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in this application, based on the technical solution and inventive concept of this application, should be included within the scope of protection of this application.

Claims

1. A cultural tourism memory generation system based on image recognition and large language model, characterized in that, It includes a mini-program interaction module, a cultural and tourism resource database construction module, an image acquisition and quality gating module, an image recognition and structured tag generation module, a cultural and tourism route matching module, a large language model controlled generation module, a content consistency and compliance verification module, and a results storage and feedback optimization module.

2. The cultural tourism memory generation system based on image recognition and large language model according to claim 1, characterized in that, The aforementioned mini-program interaction module provides tourists with interactive entry points for logging in, browsing the homepage, selecting routes, viewing roles, making reservations, viewing orders, offline verification, uploading photos, generating AI stories, viewing electronic albums, saving results, sharing results, and deleting results.

3. The cultural tourism memory generation system based on image recognition and large language model according to claim 1, characterized in that, The cultural and tourism resource database construction module is used to establish structured cultural and tourism knowledge data related to scenic spots. The cultural and tourism resource database includes a scenic spot node database, a story route database, a character identity database, an action task database, a cultural theme database, a story template database, and a prohibited content database. The aforementioned scenic spot node database is used to store the scenic spot number, scenic spot name, geographical location, visual features, cultural meaning, and recommended shooting actions; The story route library is used to store route number, route name, age range, keywords, story background and route node order; The role identity database is used to store the role names, descriptions, and behavioral characteristics that tourists can select or match in the route. The action task library is used to store photo-taking actions, interactive tasks, and triggering conditions corresponding to line nodes. The aforementioned cultural motif database is used to store narrative themes related to the Yiwulü Mountain culture; The story template library is used to store story structures under different routes, different characters, and different nodes; The aforementioned content ban library is used to restrict inappropriate content, fabricated or excessive content, and content that is inconsistent with the culture of the scenic area.

4. The cultural tourism memory generation system based on image recognition and large language model according to claim 1, characterized in that, The image acquisition and quality gating module is used to receive travel photos uploaded by tourists and to determine the usability of the photos. The quality gating includes one or more of the following: sharpness detection, exposure detection, person integrity detection, closed eyes detection, occlusion detection, subject proportion detection, duplicate frame detection, and image format detection. When an image does not meet the preset quality threshold, the system prompts the tourist to retake the photo or re-upload it. When an image meets the quality threshold, the system passes the image to the subsequent recognition module.

5. A cultural tourism memory generation system based on image recognition and a large language model according to claim 1, characterized in that, The image recognition and structured label generation module is used to recognize travel photos that have passed quality gating, extract visual information from the images, and generate structured image labels. The structured image labels include labels for the number of people, the location of people, actions, expressions, clothing or props, background elements, candidate scenic spots, text labels, lighting atmosphere, and image confidence. The image recognition uses one or more combinations of target detection models, image classification models, pose recognition models, text recognition models, or multimodal visual models.

6. The cultural tourism memory generation system based on image recognition and large language model according to claim 1, characterized in that, The cultural tourism route matching module is used to match structured image tags with the route information, role information, scenic spot node information, action task information and cultural theme information selected by tourists, so as to obtain the cultural tourism story branch that best matches the current image. The matching process is implemented using one or more of the following methods in combination: rule matching, weighted scoring, vector similarity calculation, or large language model-assisted judgment; The cultural tourism route matching module makes its judgment based on the following matching scoring method: S = αA + βB + γC + δD + εE Where S represents the comprehensive matching score, A represents the matching degree between image node features and scenic area node library, B represents the matching degree between character actions and action task library, C represents the matching degree between tourist selected route and story route library, D represents the matching degree between expression or atmosphere and cultural theme library, E represents the matching degree between character identity and story template library, and α, β, γ, δ, and ε are preset weight coefficients. The system selects target route, target node, and target story template based on the comprehensive matching score. When the highest matching score is lower than the preset threshold, the system enters the fallback generation process.

7. A cultural tourism memory generation system based on image recognition and a large language model according to claim 1, characterized in that, The large language model controlled generation module is used to construct a controlled generation prompt package based on structured image labels, route matching results, scenic area cultural constraints, story templates, output format requirements, and prohibited content constraints, and call the large language model to generate cultural tourism memory text; The controlled generation prompt package includes system role descriptions, image fact packages, route fact packages, node fact packages, cultural theme constraints, story style constraints, prohibited generation content, output format requirements, and generation length requirements.

8. A cultural tourism memory generation system based on image recognition and a large language model according to claim 1, characterized in that, The content consistency and compliance verification module is used to verify the cultural and tourism memory content generated by the large language model. The verification includes image consistency verification, node consistency verification, character consistency verification, action consistency verification, cultural consistency verification, format consistency verification, and prohibited content verification. If the generated content contains characters, actions, landscapes, or scenic spot nodes that are not present in the photos, or contains content that is inconsistent with, inappropriate, or inaccurate to the cultural theme, the system determines that the verification fails.

9. A cultural tourism memory generation system based on image recognition and a large language model according to claim 1, characterized in that, The aforementioned results storage and feedback optimization module is used to save the cultural and tourism memory results generated by tourists. The results include story title, story text, travel photo address or image data, generation time, route number, node number, role information, cultural tags, generation status, and user feedback information. Tourists can view historical AI albums, read story content, save stories, share results, or delete stories on the mini-program. The system can optimize subsequent route matching weights, story templates, and generation parameters based on tourists' saving, deleting, regenerating, and feedback behaviors.

10. A method used in the cultural tourism memory generation system based on image recognition and large language model as described in any one of claims 1-9, characterized in that: Step 1: Tourists log in via the mini-program and select a scenic area story route or role. Step 2: Tourists upload travel photos; Step 3: The system performs quality gating on the travel photos; Step four: The system identifies the images that have passed quality gating and generates structured labels; Step 5: The system matches scenic area routes, nodes, roles, and cultural themes based on structured tags; Step six: The system constructs a controlled generation prompt package and calls the large language model to generate cultural and tourism stories; Step 7: The system performs consistency and compliance checks on the generated content; Step 8: The system will return the verified story content to the mini-program for display and save it to the electronic photo album.