A visual generation method and visual system based on a digital tourism scene
By acquiring information from visitors from other regions and cultural tourism animations, and using preset matching rules to generate target video animations, the problem of visitors from other regions not being able to fully understand the highlights of the local area has been solved. This has enabled the intelligent generation of visual cultural tourism scenes and enhanced the experience of cross-regional cultural exchange.
Patent Information
- Application Number
- CN202610182572.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-06-09
AI Technical Summary
In the current technology, visitors from other places do not fully understand the local highlights during their visits, resulting in insufficient cross-regional cultural exchange.
By acquiring basic information of visitors from other locations and original cultural and tourism animations, target video animations are generated using preset matching rules, and mismatched video clips are filtered and adjusted to achieve intelligent visual generation of cultural and tourism scenes.
It enhances the experience of cross-regional cultural exchange, ensures that the displayed videos are smooth and consistent with the cultural characteristics of visitors from other regions, and avoids sensitive content from affecting the display effect.
Smart Images

Figure CN122176138A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart cultural tourism, and in particular to a visual generation method and visual system based on digital cultural tourism scenarios. Background Technology
[0002] In the existing technology, the key link in the development of the digital cultural tourism industry lies in the technology solution providers, who integrate the scattered technologies and content upstream to create a feasible overall solution, such as a smart scenic spot management system and a digital museum system. The relevant technology solution providers are responsible for the integration of hardware equipment, the development and debugging of software, and the on-site deployment of hardware and software, and are the bridge connecting technology and application scenarios.
[0003] In the process of urban development, smart tourism plays a vital role in external and inter-regional exchanges, requiring the presentation of the city's image and culture to visitors from other places. For example, when visiting groups of tourists from other regions come to the city, they may sometimes consider factors such as nationality, cultural background, and political stance, and thus avoid certain cultural and tourism venues or areas in their tour itineraries, resulting in them not fully experiencing the city's highlights. Therefore, the industry needs to design a solution to address the related technical challenges in these situations. Summary of the Invention
[0004] The technical problem solved by this invention is how to design a solution that can realize the intelligent generation of visual results in cultural and tourism scenarios, display reasonable visual generation results according to the specific circumstances of visitors from other places, and thus enhance the experience of cross-regional cultural exchange.
[0005] In a first aspect, this application proposes a visual generation method based on digital cultural tourism scenarios. The method is used in a control unit of a visual system, which controls the display unit of the visual system to display videos to visitors from other locations. The method includes: S1, obtaining the corresponding tags of the basic information of visitors from other locations, and the corresponding tags of the original cultural tourism animation to be displayed; S2, generating a target video animation based on a set of content tags according to a preset matching rule.
[0006] The further technical solution is that step S1 includes: S11, obtaining basic information of visitors from other places, and generating a set of cultural tags corresponding to visitors from other places based on the basic information; S12, obtaining the original cultural and tourism animation, dividing the original cultural and tourism animation into N video segments, and parsing the content of the N video segments according to the preset tag framework to generate a set of content tags corresponding to the N video segments.
[0007] The further technical solution is that S2 includes: S21, matching the set of cultural tags corresponding to visitors from other places with the set of content tags of N video clips according to the preset matching rules, filtering out the mismatched video clips that do not match the set of cultural tags, optimizing the mismatched video clips from the original cultural tourism animation, and obtaining the target video animation based on the remaining NM video clips.
[0008] A further technical solution is that S11 includes: S311, obtaining basic information of the group corresponding to the visitors from other places, the basic information corresponding to the common attributes of the group, and generating a set of cultural tags corresponding to the group based on the basic information.
[0009] A further technical solution is that S12 includes: S321, segmenting the original cultural and tourism animation into segments that conform to a preset number range, and then obtaining N video segments; S322, performing multimodal analysis on the N video segments to extract image features and / or audio features and / or text features from the video segments; S323, converting the image features and / or audio features and / or text features into attribute semantic features, and matching them with a preset content tag library, and then obtaining a set of content tags for the N video segments.
[0010] Further, S21 includes: matching the cultural tag set corresponding to visitors from other places with the content tag set of N video clips according to preset matching rules; using the cultural tag set as a benchmark, filtering out a portion of the content tag set that does not match and marking the corresponding portion of the video clips as mismatched video clips; determining whether the fluency of the preceding and following adjacent segments of each mismatched video clip in the original cultural tourism animation after merging them in chronological order is lower than a preset requirement; if the fluency after merging in chronological order is not lower than the preset requirement, deleting the corresponding mismatched video clip, so that the preceding and following adjacent segments are directly merged into the first target video animation; if the fluency after merging in chronological order is lower than the preset requirement, adjusting the audio-to-text content of the mismatched video clips based on a preset instruction set until the audio-to-text content of the mismatched video clips matches the cultural tag set, obtaining the first selected text segment, and generating the second target video animation that sequentially includes the preceding adjacent segment, the first selected text segment, and the following adjacent segment.
[0011] Further, the step of adjusting the audio-to-text content of mismatched video clips based on a preset instruction set until the audio-to-text content of the mismatched video clips matches the cultural tag set includes: adjusting the audio-to-text content of the mismatched video clips based on the preset instruction set, causing the semantics of the audio-to-text content to change, and randomly generating a first semantic fragment whose semantics point to the cultural tag set; determining whether the content of the first semantic fragment matches the cultural tag set; if the content of the first semantic fragment matches the cultural tag set, mapping the content of the first semantic fragment to the first selected text fragment; if the content of the first semantic fragment does not match the cultural tag set, continuing to adjust the first semantic fragment based on the preset instruction set, and randomly generating a second semantic fragment whose semantics are closer to the cultural tag set; determining whether the content of the second semantic fragment matches the cultural tag set; if the content of the second semantic fragment matches the cultural tag set, mapping the content of the second semantic fragment to the first selected text fragment; if the content of the first semantic fragment does not match the cultural tag set, the visual system generates a warning signal.
[0012] Secondly, this invention discloses a visual system comprising a control unit and a display unit interconnected with each other. The display unit is located in the visitor area and is used to display videos to visitors from other regions entering the area. The control unit in the visual system is used to implement the visual generation method based on digital cultural tourism scenarios as described in the first aspect. The technical effect of the visual system is that it can realize the intelligent generation of visual results for cultural tourism scenarios, display reasonable visual generation results according to the specific circumstances of visitors from other regions, thereby enhancing the experience of cross-regional cultural exchange.
[0013] In conclusion, in the process of urban development, smart cultural tourism plays an important role in external or inter-regional exchanges, and needs to showcase the city's image or culture to visitors from other places. Based on this, the solution described in this application can realize the intelligent generation of visual results for cultural tourism scenes, and display reasonable visual generation results according to the specific circumstances of visitors from other places, thereby enhancing the experience of cross-regional cultural exchanges. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart of a visual generation method based on digital cultural tourism scenarios provided by the present invention.
[0017] Figure 2 Another flowchart of the visual generation method based on digital cultural tourism scenarios provided by the present invention.
[0018] Figure 3 This is another flowchart of the visual generation method based on digital cultural tourism scenarios provided by the present invention.
[0019] Figure 4 A simplified diagram of the electronic device provided by the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0022] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0023] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to one or any combination of the associated listed items and all possible combinations, and includes such combinations.
[0024] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0025] In this specification and the appended claims, there may be multiple ways of expressing the same technical feature or technical term, such as using a superordinate generalization, a subordinate limitation, or a synonym substitution. Those skilled in the art can clearly understand the substantially the same technical meaning referred to by different expressions based on their professional knowledge and in conjunction with the overall content of the specification and the drawings. The differences in different expressions are only reflected in the diversity of words and do not constitute a substantial modification or limitation to the technical solution, nor will they affect the certainty of the scope of protection of this patent claim or the full disclosure of the technical content of the specification.
[0026] Example 1
[0027] Please see Figures 1 to 3 The illustration shows a visual generation method based on digital cultural tourism scenarios proposed in this invention. In a first aspect, the method is used in a control unit within a visual system. The control unit controls the display unit of the visual system to display videos to visitors from other regions. The method includes: S1, obtaining corresponding tags for the basic information of visitors from other regions, and corresponding tags for the original cultural tourism animation to be displayed; S2, generating a target video animation based on a set of content tags according to preset matching rules. The scope of visitors from other regions can include guests from overseas, guests with specific cultural backgrounds, guests with specific religious beliefs, etc., who are considered potential targets for friendly exchange. This definition is understood in the art. The technical effect of the above solution is that it can achieve intelligent generation of visual results for cultural tourism scenarios, displaying reasonable visual generation results according to the specific circumstances of visitors from other regions, thereby enhancing the experience of cross-regional cultural exchange.
[0028] In one embodiment, step S1 includes: S11, obtaining basic information of visitors from other locations, and generating a set of cultural tags corresponding to the visitors from other locations based on the basic information; S12, obtaining the original cultural tourism animation, dividing the original cultural tourism animation into N video segments, and performing content parsing on the N video segments according to a preset tag framework to generate a set of content tags corresponding to the N video segments. The preset tag framework for parsing the N video segments according to the preset tag framework is a pre-standardized classification framework, which can be constructed by staff or reviewed and confirmed by staff after referring to clustering results. Those skilled in the art can understand the definition of the preset tag framework.
[0029] In one embodiment, S2 includes: S21, matching the set of cultural tags corresponding to visitors from other places with the set of content tags of N video clips according to a preset matching rule, filtering out video clips that do not match the set of cultural tags, optimizing the mismatched video clips from the original cultural tourism animation, and obtaining the target video animation based on the remaining NM video clips. Corresponding to filtering out the video clips that do not match the set of cultural tags, for example, visitors from other places in cross-strait exchanges (or religious exchanges) correspond to the set of cultural tags, and a certain video clip corresponds to the set of content tags; then, in a certain video clip, most of the content can be displayed, and only a small part of the content is not suitable for display. Filtering out the video clips that do not match the set of cultural tags, i.e., the small number of mismatched video clips, means that even if these mismatched video clips are not present in the original cultural tourism animation, it does not affect the main theme. The target video animation is obtained based on the remaining video clips. Therefore, the generated target video animation has hidden potentially sensitive segments and can be displayed and communicated normally.
[0030] In one embodiment, S11 includes: S311, obtaining basic information about the group corresponding to the visitors from other locations, wherein the basic information corresponds to the common attributes of the group, and generating a set of cultural tags corresponding to the group based on the basic information. The basic information includes at least one of nationality information, cultural background information, religious information, and stance information. For example, if a visiting group comes to a certain city, then this visiting group corresponds to the same cultural background information.
[0031] In one embodiment, S12 includes: S321, segmenting the original cultural and tourism animation into segments conforming to a preset number range, resulting in N video segments; S322, performing multimodal analysis on the N video segments to extract image features and / or audio features and / or text features from the video segments; S323, converting the image features and / or audio features and / or text features into attribute semantic features, and matching them with a preset content tag library, resulting in a set of content tags for the N video segments. Specifically, the segmentation of the original cultural and tourism animation into segments conforming to a preset number range limits the upper and lower limits of the preset number to avoid overly broad or detailed segmentation, which could lead to excessive resource consumption later. Furthermore, the conversion of image features and / or audio features and / or text features into attribute semantic features can correspond to historical evaluation, cultural evaluation, or religious evaluation, and is matched with a preset content tag library, i.e., using the content tag library as the benchmark for the framework.
[0032] In one embodiment, S21 includes: matching the cultural tag set corresponding to visitors from other places with the content tag set of N video segments according to a preset matching rule; using the cultural tag set as a benchmark, filtering out a portion of the content tag set that does not match and marking the corresponding portion of the video segments as mismatched video segments; determining whether the fluency of the preceding and following adjacent segments of each mismatched video segment in the original cultural tourism animation after merging them in chronological order is lower than a preset requirement; if the fluency after merging in chronological order is not lower than the preset requirement, deleting the corresponding mismatched video segment, so that the preceding and following adjacent segments are directly merged into the first target video animation; if the fluency after merging in chronological order is lower than the preset requirement, adjusting the audio-to-text content of the mismatched video segment based on a preset instruction set until the audio-to-text content of the mismatched video segment matches the cultural tag set, obtaining the first selected text segment, and generating the second target video animation that sequentially includes the preceding adjacent segment, the first selected text segment, and the following adjacent segment.
[0033] Furthermore, whether the fluency is below a preset requirement is a judgment that can be made by an engineer in the art using a large model; the deletion of the corresponding mismatched video segments refers to the fact that each mismatched video segment exists independently. The aforementioned intelligent review process is not limited to fully intelligent or semi-human / semi-intelligent methods. Furthermore, the adjustment of the audio-to-text content of the mismatched video segments based on a preset instruction set can utilize a large model or an intelligent agent; furthermore, the audio-to-text content is a term understood in the art, such as video subtitles.
[0034] In one embodiment, adjusting the audio-to-text content of mismatched video clips based on a preset instruction set until the audio-to-text content of the mismatched video clips matches the cultural tag set includes: adjusting the audio-to-text content of the mismatched video clips based on the preset instruction set, causing the semantics of the audio-to-text content to change, and randomly generating a first semantic segment whose semantics point to the cultural tag set; determining whether the content of the first semantic segment matches the cultural tag set; if the content of the first semantic segment matches the cultural tag set, mapping the content of the first semantic segment to the first selected text segment; if the content of the first semantic segment does not match the cultural tag set, continuing to adjust the first semantic segment based on the preset instruction set, and randomly generating a second semantic segment whose semantics are closer to the cultural tag set; determining whether the content of the second semantic segment matches the cultural tag set; if the content of the second semantic segment matches the cultural tag set, mapping the content of the second semantic segment to the first selected text segment; if the content of the first semantic segment does not match the cultural tag set, the visual system generates a warning signal. In the above scheme, the random generation of the first semantic segment whose semantics point to the cultural tag set can be achieved through AI generation, as understood by those skilled in the art.
[0035] In the process of urban development, smart cultural tourism plays an important role in external or inter-regional exchanges, requiring the presentation of the city's image or culture to visitors from other places. For example, when visiting groups of visitors from other places come to the city, they may sometimes avoid certain cultural and tourism venues or areas in the design of their tour routes due to factors such as nationality, cultural background, and political stance, resulting in visitors not being able to fully understand the city's highlights. The solution described in this application first selects mismatched video clips that are not suitable for direct presentation to visitors from other places, and then deletes or modifies these mismatched video clips to ensure the smoothness of the entire video. In this way, the presentation of the entire video will not be affected by a single mismatched video clip, and the efficiency of video adjustment and optimization is significantly improved (no need to repeatedly modify the video for different visitors from other places). Therefore, it is conducive to further enhancing the experience of cross-regional cultural exchanges.
[0036] Secondly, this invention discloses a visual system comprising a control unit and a display unit interconnected with each other. The display unit is located in the visitor area and is used to display videos to visitors from other regions entering the area. The control unit in the visual system is used to implement the visual generation method based on digital cultural tourism scenarios as described in the first aspect. The technical effect of the visual system is that it can realize the intelligent generation of visual results for cultural tourism scenarios, display reasonable visual generation results according to the specific circumstances of visitors from other regions, thereby enhancing the experience of cross-regional cultural exchange.
[0037] In conclusion, in the process of urban development, smart cultural tourism plays an important role in external or inter-regional exchanges, and needs to showcase the city's image or culture to visitors from other places. Based on this, the solution described in this application can realize the intelligent generation of visual results for cultural tourism scenes, and display reasonable visual generation results according to the specific circumstances of visitors from other places, thereby enhancing the experience of cross-regional cultural exchanges.
[0038] Example 2
[0039] Please see Figure 4 , Figure 4 This invention provides a block diagram of an electronic device. The electronic device can be a terminal or a server. The terminal can be a smartphone, tablet computer, laptop computer, desktop computer, personal digital assistant, wearable device, or other electronic device with communication capabilities. It includes a processor 111, a communication interface 112, a memory 113, and a communication bus 114. The processor 111, communication interface 112, and memory 113 communicate with each other via the communication bus 114.
[0040] Memory 113 is used to store computer programs.
[0041] In one embodiment of the present invention, the processor 111, when executing the program stored in the memory 113, implements the method provided in any of the foregoing method embodiments.
[0042] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program may be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0043] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions, but such implementations should not be considered beyond the scope of this invention.
[0044] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is only a logical functional division, and there may be other division methods in actual implementation. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0045] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0046] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0047] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0048] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Since these modifications and variations fall within the scope of the claims and their equivalents, this invention also intends to include these modifications and variations.
[0049] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A visual generation method based on digital cultural tourism scenarios, characterized in that, The method is used in a control unit of a vision system, the control unit being used to control the display unit of the vision system to display video to remote visitors, the method comprising: S1, obtain the corresponding tags for the basic information of visitors from other places, and the corresponding tags for the original cultural and tourism animation to be displayed; S2 generates the target video animation based on the content tag set according to the preset matching rules.
2. The visual generation method based on digital cultural tourism scenes according to claim 1, characterized in that, Step S1 includes: S11: Obtain basic information of visitors from other locations and generate a set of cultural tags corresponding to the visitors from other locations based on the basic information. S12, obtain the original cultural and tourism animation, divide the original cultural and tourism animation into N video segments, and perform content parsing on the N video segments according to the preset tag framework to generate a set of content tags corresponding to the N video segments.
3. The method according to claim 2, characterized in that, S2 includes: S21. According to the preset matching rules, the set of cultural tags corresponding to visitors from other places is matched with the set of content tags of N video clips. The video clips that do not match the set of cultural tags are filtered out. The video clips that do not match the set of cultural tags are optimized from the original cultural tourism animation. The target video animation is obtained based on the remaining NM video clips.
4. The method according to claim 3, characterized in that, S11 includes: S311, Obtain basic information about the group corresponding to visitors from other places, the basic information corresponds to the common attributes of the group, and generate a set of cultural tags corresponding to the group based on the basic information.
5. The method according to claim 4, characterized in that, S12 includes: S321, the original cultural tourism animation is segmented into segments that meet the preset number range, resulting in N video segments; S322, Perform multimodal analysis on N video segments to extract image features and / or audio features and / or text features from the video segments; S323, transform image features and / or audio features and / or text features into attribute semantic features, and match them with a preset content tag library to obtain a set of content tags for N video segments.
6. The method according to claim 5, characterized in that, S21 includes: According to the preset matching rules, the cultural tag set corresponding to visitors from other places is matched with the content tag set of N video clips. Based on the cultural tag set, a portion of the content tag set that does not match is filtered out and the corresponding portion of the video clips is marked as mismatched video clips. Determine whether the fluency of the preceding and following adjacent segments of each mismatched video segment in the original cultural tourism animation, after being merged in chronological order, is lower than the preset requirements. If the smoothness of the merged video clips in chronological order is not lower than the preset requirements, delete the corresponding mismatched video clips so that the preceding adjacent clips and the following adjacent clips can be directly merged into the first target video animation; If the fluency of the merged video clips in chronological order is lower than the preset requirements, the audio-to-text content of the mismatched video clips is adjusted based on the preset instruction set until the audio-to-text content of the mismatched video clips matches the cultural tag set, thus obtaining the first selected text clip and generating a second target video animation that sequentially includes the preceding adjacent clip, the first selected text clip, and the following adjacent clip.
7. The method according to claim 6, characterized in that, The step of adjusting the audio-to-text content of mismatched video clips based on a preset instruction set until the audio-to-text content of the mismatched video clips matches the cultural tag set includes: Based on a preset set of instructions, the audio-to-text content of mismatched video segments is adjusted, so that the semantics of the audio-to-text content changes, and a first semantic segment is randomly generated that points to the set of cultural tags. Determine whether the content of the first semantic fragment matches the cultural tag set. If the content of the first semantic fragment matches the cultural tag set, map the content of the first semantic fragment to the first selected text fragment. If the content of the first semantic fragment does not match the cultural tag set, the first semantic fragment is adjusted based on the preset instruction set, and a second semantic fragment with semantics closer to the cultural tag set is randomly generated; Determine whether the content of the second semantic fragment matches the cultural tag set. If the content of the second semantic fragment matches the cultural tag set, map the content of the second semantic fragment to the first selected text fragment. If the content of the first semantic fragment does not match the cultural tag set, the visual system generates a warning signal.
8. A vision system, characterized in that, The visual system includes a control unit and a display unit that are interconnected. The display unit is located in the visitor area and is used to display videos to visitors from other places who enter the visitor area. The control unit in the visual system is used to implement the visual generation method based on digital cultural tourism scenes as described in any one of claims 1 to 7.