IP content and real scene sv interactive experience travel value-added data driven system
By deeply integrating cultural IP with real-world 3D scenes, user-participatory content co-creation, and behavioral data analysis, the problems of low user participation and unutilized data in cultural tourism experience systems have been solved, thereby enhancing immersion and commercial value.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JIUYUNGUANG DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-29
AI Technical Summary
In existing cultural tourism experience systems, the connection between cultural IPs and 3D scenes is superficial, user participation is low, and user behavior data is not deeply mined, resulting in a rigid experience, lack of vitality, and a broken data value chain, which fails to drive commercial value-added.
Through modules for scene 3D modeling, map construction, IP content processing, and content optimization, the system achieves deep and dynamic integration of cultural IP with real-world 3D scenes, enabling user-participatory content co-creation, real-time collection and analysis of user behavior data, and generation of personalized value-added services.
It enhances the immersiveness and interactivity of the user experience, strengthens the sustainable appeal of the experience, forms a complete value chain of data collection, analysis, optimization, and value-added services, provides quantitative operational decision-making basis for cultural tourism projects, and significantly enhances commercial value.
Smart Images

Figure CN122115802A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart cultural tourism technology, and in particular to a data-driven system for interactive cultural tourism experiences based on IP content and real-world SV. Background Technology
[0002] With the widespread application of digital technologies such as augmented reality (AR) and 3D modeling in the cultural and tourism industry, the digital upgrade of cultural and tourism experiences has become an important trend in industry development. Currently, common digital cultural and tourism experience solutions in the industry mainly include overlaying fixed graphic and textual information on real-world 3D models, setting up virtual characters for one-way explanations, and setting up limited scanning trigger points to trigger preset content. However, these existing solutions suffer from several technical shortcomings that urgently need to be addressed: First, the connection between cultural IP content and 3D scenes is mostly static and superficial, lacking in-depth integration. For example, a certain ancient city tourism app can only display historical text descriptions of specific buildings in a 3D ancient city model, with no dynamic interaction logic between IP elements and scenes, resulting in a stiff user experience and weak immersion. Second, the interactive content in the system is mostly pre-set during the development phase, with users acting only as passive recipients, unable to effectively participate in content creation and enrichment. For example, in a museum's AR guide system, users can only view fixed exhibit descriptions along a preset path, unable to share their own insights or create related interactive content, making the experience lack vitality and sustainable appeal. Third, the large amount of behavioral data generated by users during interaction has not been deeply mined and utilized. This data is stored in a scattered manner and has not formed an effective analysis system, making it difficult to transform the process data of interactive experiences into quantifiable assets that can assess operational effectiveness, guide service optimization, and drive consumption conversion. This leads to a break in the data value chain and fails to provide strong support for the continuous optimization and commercial value-added of cultural tourism projects.
[0003] To address the problems existing in the prior art, this invention proposes a data-driven system for interactive cultural tourism experiences based on IP content and real-world SV. Summary of the Invention
[0004] This invention provides an interactive experience cultural tourism value-added data-driven system based on IP content and real-world SV, in order to solve the aforementioned technical problems.
[0005] This invention provides an interactive experience-based cultural tourism value-added data-driven system based on IP content and real-world SV, comprising:
[0006] The scene 3D modeling module is used to perform 3D digital reconstruction of the target cultural tourism scene, track the historical virtual interaction screens of each scene object in the reconstructed 3D space, and parse each historical virtual interaction screen according to the range of coarse elements of the screen specified by the interaction command to construct the event interaction type-interaction feedback sequence.
[0007] The graph construction module is used to extract all event interaction types and interaction feedback sequences under the same scene object according to the same event interaction type, to obtain the interaction feedback set of the same event interaction type, and to construct facial expression feedback graph, voice feedback graph, touch dwell feedback graph and physiological fluctuation change graph. In addition, by combining the cultural and tourism theme of the target cultural and tourism scene and the object attributes of the scene object, a comprehensive defect graph of the scene object is obtained.
[0008] The IP content processing module is used to parse the original cultural IP materials of the target cultural tourism scene and construct an IP knowledge architecture. It dynamically binds each element point in the IP knowledge architecture with the comprehensive defect map of each scene object involved in the element point to generate initial interactive content.
[0009] The content optimization module is used to obtain original interactive content submitted by users, extract detailed content from the original interactive content that matches the scene objects involved in each element point, and optimize the initial interactive content according to the detailed content to obtain optimized interactive content.
[0010] The value-added module is used to collect and analyze user behavior data in real time during the interaction process, and combine it with the optimized interaction content to generate experience optimization strategies and user behavior profiles, and output personalized value-added service guidance signals.
[0011] Preferably, the scene 3D modeling module includes:
[0012] The single-dimensional intent determination submodule is used to obtain a decomposition model that matches the interaction dimension and dimension quantity of the interaction instruction from the dimension-quantity-decomposition lookup table, input the interaction instruction into the decomposition model, and output the single-dimensional temporal-intent sequence of each interaction dimension, wherein the interaction dimension includes: voice dimension, touch dimension, facial expression dimension or physiological dimension.
[0013] The intent fusion submodule is used to fuse the single-dimensional temporal-intent sequences under all interaction dimensions to obtain a fused temporal-intent sequence, and to perform temporal alignment processing with the scene interaction behavior chain of the corresponding historical virtual interaction scene to obtain a sub-interaction surface that is in the same temporal sequence as the fused temporal-intent.
[0014] The mapping analysis submodule is used to perform screen mapping analysis on the fused time-intention of the same time sequence, and apply it to the corresponding sub-interaction surface to perform coarse selection of screen boundaries, and extract the existing elements of the coarsely selected screen.
[0015] The position locking submodule is used to lock the screen positions of all existing elements on all sub-interaction surfaces under the same interaction command, and to perform salience processing on the corresponding locked positions according to the locking frequency of each locked position and the screen attributes of the locked position to obtain the coarse element range of the screen.
[0016] The event acquisition submodule is used to align the coarse element range of the screen with the historical virtual interactive screen, perform fine event acquisition on the aligned part according to the saliency processing result, and perform coarse event acquisition on the remaining part of the historical virtual interactive screen.
[0017] The feedback and event association submodule is used to obtain the first capture feedback based on each fine-grained event and the second capture feedback based on each coarse-grained event, and associate them with the corresponding fine-grained events and coarse-grained events to obtain the event interaction type-interaction feedback sequence.
[0018] Preferably, the intent fusion submodule includes:
[0019] The function determination unit is used to align and sort all single-dimensional time-intent sequences according to the time sequence to obtain the overall time sequence matrix, and input each column vector into the intent divergence model to obtain the first intent divergence function, wherein each divergence term of the first intent divergence function corresponds one-to-one with the corresponding interaction dimension.
[0020] The clustering analysis unit is used to perform clustering analysis on all first intentional divergence functions, determine several first clusters, and lock the time nodes corresponding to all column vectors in the first clusters to obtain the first position distribution;
[0021] The proportion determination unit is used to determine if the proportion of the first position distribution is greater than or equal to a preset proportion, and then obtain the column vector corresponding to the first position in the first position distribution, which is regarded as the first column; at the same time, obtain the column vector corresponding to the second position in the first position distribution, which is regarded as the second column.
[0022] The moving unit is configured to determine the total sequence interval between the first column and the second column if the divergence terms of the first intention divergence function of the first column and the second column are the same, and to move the row vector corresponding to the same term to the right by the total sequence interval number starting from the first element, while keeping the elements in the first column before the position corresponding to the same term unchanged.
[0023] The first fusion unit is used to obtain a new matrix after all identical items have been moved, and to perform fusion processing on each column vector to obtain a fused temporal-intent sequence;
[0024] The discrete determination unit is used to obtain a divergence matrix by each column vector corresponding to the first position distribution if the divergence terms of the first intention divergence function of the first column and the second column are different terms, and to obtain the divergence discrete points of each row vector in the divergence matrix, and then count the total number of divergence discrete points of each column vector in the divergence matrix.
[0025] The numerical partitioning unit is used to partition all total quantities according to the same numerical value, calculate the ratio of the number of partition column vectors with the same numerical value to the total number of column vectors in the bifurcation matrix, and filter the maximum ratio.
[0026] A single elimination unit is used to obtain the similarity between each column vector corresponding to the maximum ratio and any adjacent column vector if the value corresponding to the maximum ratio is the maximum value among all the total quantities. The column vector corresponding to the maximum ratio with a similarity less than a preset value is eliminated by single elimination. The column vectors in the overall time series matrix that have not been eliminated by single elimination are fused to obtain a fused time series-intent sequence.
[0027] The complete elimination unit is used to eliminate all column vectors corresponding to the maximum ratio if the value corresponding to the maximum ratio is not the maximum value among all the total quantities, and to perform fusion processing on each column of the retained vectors of the overall time series matrix to obtain the fused time series-intent sequence.
[0028] Preferably, the intent fusion submodule further includes:
[0029] A window function construction unit is used to lock the column vector corresponding to each first cluster in the overall time series matrix as the fourth column if the distribution ratio of the first position distribution is less than a preset ratio, and construct a window function according to the number of clusters of the first cluster.
[0030] The extraction unit is used to divide and extract each fourth column from the overall time series matrix according to the window function to obtain the window matrix corresponding to each fourth column. The window function corresponding to each fourth column contains the fourth column, and its column position is random and not unique.
[0031] An adjustment unit is used to obtain the basic rate of change of each row vector in the window matrix, and adjust the row elements corresponding to the fourth column based on the basic rate of change to obtain the adjusted fourth column;
[0032] The second fusion unit is used to construct a fusion matrix by combining the adjusted fourth column and the unadjusted columns in the overall time series matrix, and to perform fusion processing on each column vector in the fusion matrix to obtain a fused time series-intent sequence.
[0033] Preferably, the map construction module includes:
[0034] The vector and weight determination submodule is used to extract the feature vector of each feedback map, and simultaneously determine the first weight matrix based on the cultural tourism theme of the target cultural tourism scene and the object attributes of the scene object. Second weight matrix The first weight matrix It is a 4×4 diagonal matrix, with diagonal elements representing facial expression weight, voice weight, touch persistence weight, and physiological weight, respectively. This is the second weight matrix. It is a 4×4 matrix, and each element in the 4×4 matrix This represents the strength of the attribute association between the i-th feature dimension and the j-th feature dimension;
[0035] The comprehensive vector acquisition submodule is used to obtain the feature vector and the first weight matrix. Second weight matrix Determine the comprehensive defect feature vector Z of the scene object;
[0036] The dimension reduction processing submodule is used to perform dimension reduction processing on the comprehensive defect feature vector Z to obtain the comprehensive defect value Sc of the scene object;
[0037] The comprehensive defect map generation submodule is used to generate a comprehensive defect map of the scene object based on the comprehensive curve value Sc and the relative weights of each dimension in the comprehensive defect feature vector Z. The comprehensive defect map includes at least a defect intensity dimension and a defect type distribution dimension.
[0038] Preferably, the IP content processing module includes:
[0039] The dimension extraction submodule is used to extract the core attribute dimensions of each element point. The core attribute dimensions include cultural connotation dimension, expression form dimension, and interaction adaptation dimension.
[0040] The priority determination submodule is used to map and match the core attribute dimensions with each defect dimension in the comprehensive defect map of each scene object involved, calculate the matching degree with each scene object involved, and obtain the priority coefficient of each element point.
[0041] The dynamic binding submodule is used to dynamically bind element points to the comprehensive defect map of each scene object in order of priority coefficient. For the defect dimension - IP element point adaptation gap that exists after binding, it calls the related element points in the IP knowledge architecture to supplement the binding and obtain the initial interactive content.
[0042] Preferably, the content optimization module includes:
[0043] The mapping and association submodule is used to extract the initial feature set of the initial interactive content, and according to the feature attributes of each initial feature in the initial feature set, match the derivative conditions that are consistent with the feature attributes from the attribute-derivative lookup table, and perform mapping and association analysis on each derivative condition with the item attributes of each sub-item of sub-content to obtain the item subset of each sub-content involved in the corresponding initial feature.
[0044] The array construction submodule is used to determine the number of entries in each subset of entries, the total weight of entries, and the main attributes pointing to entries, and to construct a three-dimensional array;
[0045] The position analysis submodule is used to analyze the three-dimensional array of each derived condition under the same initial feature, and determine the placement position of the corresponding entry subset of each three-dimensional array based on the corresponding initial feature. The placement position includes: placing it on the adjacent left, adjacent right, or any randomly divided middle position of the corresponding initial feature.
[0046] The new feature submodule is used to extract new features corresponding to each initial feature after all item subsets have been placed in their respective positions, and to obtain optimized interactive content.
[0047] Preferably, the driving value-added module includes:
[0048] The State Field Construction Submodule is used to collect and analyze user behavior data in real time during the interaction process to construct a dynamic user immersion state field.
[0049] The strategy generation submodule is used to generate real-time experience optimization strategies based on the immersion field and the optimized interaction content.
[0050] The guidance submodule is used to output personalized value-added service guidance signals based on the experience optimization strategy and user behavior profile.
[0051] Compared with the prior art, the beneficial effects of this application are as follows:
[0052] By constructing a complete technical system encompassing scene 3D modeling, graph analysis, IP content processing, content optimization, and data-driven value-added services, we can achieve deep and dynamic integration of cultural IP with real-world 3D scenes, user-participatory content co-creation, and the value transformation of user behavior data, thereby significantly enhancing the immersiveness, interactivity, and commercial value of cultural tourism experiences.
[0053] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0054] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0055] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0056] Figure 1 This is a structural diagram of an interactive experience cultural tourism value-added data-driven system based on IP content and real-world SV, as described in an embodiment of the present invention. Detailed Implementation
[0057] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0058] This invention provides a data-driven system for interactive cultural tourism experiences based on IP content and real-world SV (Site-Based Virtual Reality), such as... Figure 1 As shown, it includes:
[0059] The scene 3D modeling module is used to perform 3D digital reconstruction of the target cultural tourism scene, track the historical virtual interaction screens of each scene object in the reconstructed 3D space, and parse each historical virtual interaction screen according to the range of coarse elements of the screen specified by the interaction command to construct the event interaction type-interaction feedback sequence.
[0060] The graph construction module is used to extract all event interaction types and interaction feedback sequences under the same scene object according to the same event interaction type, to obtain the interaction feedback set of the same event interaction type, and to construct facial expression feedback graph, voice feedback graph, touch dwell feedback graph and physiological fluctuation change graph. In addition, by combining the cultural and tourism theme of the target cultural and tourism scene and the object attributes of the scene object, a comprehensive defect graph of the scene object is obtained.
[0061] The IP content processing module is used to parse the original cultural IP materials of the target cultural tourism scene and construct an IP knowledge architecture. It dynamically binds each element point in the IP knowledge architecture with the comprehensive defect map of each scene object involved in the element point to generate initial interactive content.
[0062] The content optimization module is used to obtain original interactive content submitted by users, extract detailed content from the original interactive content that matches the scene objects involved in each element point, and optimize the initial interactive content according to the detailed content to obtain optimized interactive content.
[0063] The value-added module is used to collect and analyze user behavior data in real time during the interaction process, and combine it with the optimized interaction content to generate experience optimization strategies and user behavior profiles, and output personalized value-added service guidance signals.
[0064] In this embodiment, Real Scene SV refers to a SceneView constructed based on Real Scene 3D scanning technology. Spatial and texture data of the target cultural and tourism scene are collected through technologies such as laser scanning and photogrammetry, and then processed by modeling software to form a 3D visualization scene that is 1:1 restored to the real scene.
[0065] In this embodiment, the target cultural tourism scene refers to a specific physical space or themed area that has cultural tourism development value and is available for users to visit and experience.
[0066] Historical virtual interactive screens refer to the visual records generated during users' past interactions with the reconstructed 3D virtual scene when using this system. Each frame is associated with corresponding metadata such as the interaction timestamp and interaction command type. For example, in the 3D Forbidden City Hall of Supreme Harmony scene, the continuous frames of images where users previously zoomed in to view the details of the roof ridge beasts of the Hall of Supreme Harmony by touching the screen, the scene images corresponding to when users asked about the construction date of the Hall of Supreme Harmony by voice command, and the superimposed images of the virtual guide's response are all considered historical virtual interactive screens.
[0067] Interactive commands refer to operation requests or information query commands issued by users to the system through the interactive methods supported by the system. These commands contain core information such as interaction dimensions and interaction intent. For example, users can use voice to introduce the purpose of a bronze artifact, slide their fingers on a touch screen to rotate an ancient building model in a 3D scene, use a smiling facial expression to trigger virtual interactions in the scene, or use a smart bracelet to collect heart rate data and trigger scene difficulty adjustments.
[0068] In this embodiment, the coarse element range of the screen refers to a rectangular or irregular screen area centered on the core object corresponding to the interaction command, which is determined by the interaction intent. After the core object is identified by the YOLOv8 object detection algorithm, the range is expanded outward by 50-100 pixels (adaptively adjusted according to the size of the scene object) to focus on the core interaction area.
[0069] The event interaction type-interaction feedback sequence refers to the orderly data set formed by classifying the interactive events extracted from historical virtual interactive screens according to their types and the corresponding system feedback information in chronological order. For example, if the event interaction type is "touch to view exhibit details" and the corresponding interaction feedback is "display high-definition images of exhibits + voice introduction", then the sequence of "touch to view exhibit details - display high-definition images of exhibits - voice introduction" is formed.
[0070] The comprehensive defect map refers to a visual map that integrates four types of feedback maps: facial expressions, voice, touch persistence, and physiological fluctuations. It combines cultural and tourism themes and scene object attributes to generate a map that includes defect type (cultural connotation defects, interactive response defects, performance form defects, etc.), defect intensity (quantitative score of 0-10 points), and defect distribution location (specific area of scene object).
[0071] Original cultural IP materials refer to original cultural resources with unique cultural connotations and commercial value that are associated with the target cultural tourism scene, including historical documents, folk legends, cultural symbols, works of art, intangible cultural heritage skills, etc.
[0072] The IP knowledge architecture includes the core elements of the IP and the relationships between them. In the IP knowledge architecture built with the Dunhuang Flying Apsaras IP as the core, the core elements include the image of the flying apsaras, the clothing of the flying apsaras, the musical instruments of the flying apsaras, the posture of the flying apsaras, and the mythological background. There are relationships between the elements, such as the image of the flying apsaras being associated with the clothing of the flying apsaras, and the musical instruments of the flying apsaras being associated with the mythological background. At the same time, each element contains specific sub-elements, such as the clothing of the flying apsaras containing sub-elements such as shawls, necklaces, and long skirts.
[0073] Original interactive content refers to interactive content created and submitted by users during their use of the system that is related to the target cultural tourism scene and IP content. This content can be used to supplement or optimize the existing interactive content of the system. For example, interactive Q&A on the theme of flying apsaras (containing 5 questions and answers about the myth of flying apsaras) created by users after visiting the 3D Mogao Grottoes scene, virtual decorative elements drawn in the style of Dunhuang murals, and audio content recorded in dialect explaining the history of a cave in Mogao Grottoes are all considered original interactive content.
[0074] User behavior data refers to various quantifiable data generated by users during their interaction with the system, including interaction operation data, dwell time data, feedback data, physiological state data, etc. For example, the coordinates of a user's click position in a 3D scene, the frequency and direction of swipe operations, the duration of dwell time on a certain scene object, likes / favorites / shares of interactive content, speech rate and emotional tendency of voice interaction, heart rate variability data, etc., all belong to user behavior data.
[0075] User behavior profiles refer to structured user models formed by analyzing and refining user behavior data, which can represent user interests, preferences, interaction habits, consumption tendencies, and other characteristics. For example, a user's behavior profile may show the following interests: ancient architecture and culture, intangible cultural heritage skills; interaction habits: preference for voice interaction and liking to view high-definition details; consumption tendencies: willingness to buy cultural and creative products and accept paid guided tours; and experience needs: emphasis on in-depth cultural explanations and preference for personalized interaction.
[0076] Value-added service guidance signals refer to the value-added service recommendations related to culture and tourism that the system pushes to users based on user behavior profiles and experience optimization strategies. For example, based on the user's behavior profile showing a liking for intangible cultural heritage skills, the system may push guidance signals for booking intangible cultural heritage handicraft experience classes; based on the user's multiple views of a certain type of cultural and creative products in a 3D scene, the system may push guidance signals for purchasing cultural and creative products at limited-time discounts; and based on the user's travel route, the system may push guidance signals for booking nearby characteristic restaurants.
[0077] The beneficial effects of the above technical solution are as follows: Through the coordinated operation of five modules—scene 3D modeling, map construction, IP content processing, content optimization, and value-added driving—a deep and dynamic integration of cultural IP and real-world 3D scenes is achieved. This changes the static and superficial association mode in traditional cultural tourism interaction, significantly enhancing the user's sense of immersion and engagement. By introducing user-generated interactive content and granting users the right to co-create content, the interactive content becomes more vibrant and diverse, enhancing the sustainable appeal of the experience. Simultaneously, by deeply mining the value of user behavior data, a complete value chain of data collection, analysis, optimization, and value-added is formed. This not only provides users with personalized interactive experiences and value-added services but also provides cultural tourism operators with quantitative operational decision-making basis, effectively driving the commercial value-added of cultural tourism projects.
[0078] This invention provides an interactive experience cultural tourism value-added data-driven system based on IP content and real-world scene SV, wherein the scene 3D modeling module includes:
[0079] The single-dimensional intent determination submodule is used to obtain a decomposition model that matches the interaction dimension and dimension quantity of the interaction instruction from the dimension-quantity-decomposition lookup table, input the interaction instruction into the decomposition model, and output the single-dimensional temporal-intent sequence of each interaction dimension, wherein the interaction dimension includes: voice dimension, touch dimension, facial expression dimension or physiological dimension.
[0080] The intent fusion submodule is used to fuse the single-dimensional temporal-intent sequences under all interaction dimensions to obtain a fused temporal-intent sequence, and to perform temporal alignment processing with the scene interaction behavior chain of the corresponding historical virtual interaction scene to obtain a sub-interaction surface that is in the same temporal sequence as the fused temporal-intent.
[0081] The mapping analysis submodule is used to perform screen mapping analysis on the fused time-intention of the same time sequence, and apply it to the corresponding sub-interaction surface to perform coarse selection of screen boundaries, and extract the existing elements of the coarsely selected screen.
[0082] The position locking submodule is used to lock the screen positions of all existing elements on all sub-interaction surfaces under the same interaction command, and to perform salience processing on the corresponding locked positions according to the locking frequency of each locked position and the screen attributes of the locked position to obtain the coarse element range of the screen.
[0083] The event acquisition submodule is used to align the coarse element range of the screen with the historical virtual interactive screen, perform fine event acquisition on the aligned part according to the saliency processing result, and perform coarse event acquisition on the remaining part of the historical virtual interactive screen.
[0084] The feedback and event association submodule is used to obtain the first capture feedback based on each fine-grained event and the second capture feedback based on each coarse-grained event, and associate them with the corresponding fine-grained events and coarse-grained events to obtain the event interaction type-interaction feedback sequence.
[0085] In this embodiment, the dimension-quantity-decomposition lookup table refers to a pre-established mapping table for matching interaction commands, the number of dimensions, and the corresponding decomposition models, as shown in Table 1:
[0086] Table 1. Dimension-Quantity-Decomposition Comparison Table
[0087]
[0088] In this embodiment, the single-dimensional temporal-intent sequence refers to an ordered data sequence formed by arranging the intent information after parsing the interaction command in chronological order for a single interaction dimension. Each element in the sequence represents the core intent of the interaction dimension at the corresponding time node. For example, the single-dimensional temporal-intent sequence for the voice dimension is: [0s: inquire about the age of the cultural relic, 2s: confirm the purpose of the cultural relic, 5s: query the cultural relic protection measures, 8s: end the interaction]; the single-dimensional temporal-intent sequence for the touch dimension is: [0s: click on the exhibit icon, 1s: slide to zoom in on details, 3s: drag to rotate the view, 6s: deselect].
[0089] In this embodiment, the overall time-series matrix refers to a two-dimensional matrix constructed according to the rule that rows correspond to a single interaction dimension and columns correspond to the same time node. Each row of the matrix corresponds to a complete time-series-intent sequence of an interaction dimension (row vector), and each column corresponds to the intent set of all interaction dimensions at the same time node (column vector). The matrix is represented as shown in Table 2:
[0090] Table 2 Overall Time Series Matrix
[0091]
[0092] In this embodiment, the scene interaction behavior chain refers to the chain-like data structure formed by a series of interactive behaviors executed by the user in a historical virtual interactive scene in chronological order. For example, the scene interaction behavior chain of the user in the 3D museum scene is: [t1: Enter the bronze ware exhibition area (behavior type: scene switching, behavior object: bronze ware exhibition area), t2: Click on bronze ware A (behavior type: touch operation, behavior object: bronze ware A), t3: Query the history of bronze ware A by voice (behavior type: voice interaction, behavior object: bronze ware A), t4: View high-definition images of bronze ware A (behavior type: visual interaction, behavior object: high-definition resources of bronze ware A), t5: Exit the exhibition area (behavior type: scene switching, behavior object: museum main interface)].
[0093] A sub-interaction surface refers to a local interactive scene area corresponding to a single fusion time sequence-intent after aligning the scene interaction behavior chain of a historical virtual interactive scene with the fusion time sequence-intent sequence. For example, if the intent of a certain time period in the fusion time sequence-intent sequence is to view the decorative details of bronze artifact A, and the behavior corresponding to the time-aligned scene interaction behavior chain is that the user touches the decorative part of bronze artifact A and slides to zoom in, then the corresponding sub-interaction surface is the scene area containing the decorative part of bronze artifact A and the user's touch operation mark within that time period.
[0094] In this embodiment, a lock frequency weight and a screen attribute weight are set. The lock frequency weight is determined based on the ratio of the number of locks to the total number of sub-interaction surfaces. The screen attribute weight is preset based on the scene object attributes (core area, related area, background area) corresponding to the location (e.g., the core area weight is 0.7, the related area weight is 0.3, and the background area weight is 0.1). The comprehensive salience score of each locked location is calculated as: lock frequency weight × 0.6 + screen attribute weight × 0.4. Based on the score, the location is divided into three salience levels: high, medium, and low, to achieve salience processing. For example, under the same interaction command, if a location is locked 8 times in 10 sub-interaction surfaces (high lock frequency) and the location corresponds to the core functional area of the scene object (important screen attributes), then the location is marked as high salience; if another location is locked only 2 times and corresponds to a non-core background area, then it is marked as low salience.
[0095] In this embodiment, a fine-grained event refers to a detailed interactive event extracted from a region in a historical virtual interactive screen that is aligned with the position of the coarse elements of the screen and is highly salient after salientization. This event has a clear interactive intent and a specific interactive object. For example, if the coarse elements of the screen are the display area of a cultural relic and the highly salient position is the inscription part of the cultural relic, the event in which the user touches the inscription part and triggers the interpretation of the inscription is a fine-grained event.
[0096] Coarse events refer to generalized interactive events extracted from areas other than those aligned with the coarse elements in a historical virtual interactive screen. These events are characterized by relatively vague interactive intent or unclear interactive objects. For example, in a historical virtual interactive screen, events such as users moving their perspective to browse the scene or pausing to observe the scene in the background area outside the artifact display area are coarse events. These events only describe general interactive behaviors and do not have a specific, clear object.
[0097] The first capture feedback is highly matched with the interaction intent of the fine event and can meet the specific needs of the user. For example, when the user touches the part of the inscription and triggers the inscription interpretation, the feedback generated by the system is to display a high-definition rubbing of the inscription, a word-by-word voice interpretation, and a text introduction of the historical background of the inscription. This is the first capture feedback.
[0098] Secondary capture feedback is mainly used to respond to the user's general interactive behavior and maintain the continuity of the interaction. For example, in response to coarse events such as the user moving their viewpoint to browse the scene, the system generates feedback that smoothly follows the scene viewpoint and slightly enhances the ambient sound effects. This is secondary capture feedback.
[0099] The beneficial effects of the above technical solution are as follows: The dimension-quantity-decomposition comparison table enables rapid matching of decomposition models, improving the efficiency of intent parsing; the construction of the overall temporal matrix makes the temporal alignment of multi-dimensional intents clearer, providing a well-organized data foundation for subsequent fusion processing; the saliency processing accurately delineates the range of coarse elements in the image, ensuring the targeted nature of event extraction; and the classification and extraction of fine and coarse events, along with differentiated feedback association, makes the sequence data more hierarchical and practical, providing high-quality, refined basic data for the subsequent construction of a comprehensive defect map, further enhancing the system's ability to understand and analyze user interaction behavior.
[0100] This invention provides an interactive experience-based cultural tourism value-added data-driven system based on IP content and real-world SV, wherein the intent fusion submodule includes:
[0101] The function determination unit is used to align and sort all single-dimensional time-intent sequences according to the time sequence to obtain the overall time sequence matrix, and input each column vector into the intent divergence model to obtain the first intent divergence function, wherein each divergence term of the first intent divergence function corresponds one-to-one with the corresponding interaction dimension.
[0102] The clustering analysis unit is used to perform clustering analysis on all first intentional divergence functions, determine several first clusters, and lock the time nodes corresponding to all column vectors in the first clusters to obtain the first position distribution;
[0103] The proportion determination unit is used to determine if the proportion of the first position distribution is greater than or equal to a preset proportion, and then obtain the column vector corresponding to the first position in the first position distribution, which is regarded as the first column; at the same time, obtain the column vector corresponding to the second position in the first position distribution, which is regarded as the second column.
[0104] The moving unit is configured to determine the total sequence interval between the first column and the second column if the divergence terms of the first intention divergence function of the first column and the second column are the same, and to move the row vector corresponding to the same term to the right by the total sequence interval number starting from the first element, while keeping the elements in the first column before the position corresponding to the same term unchanged.
[0105] The first fusion unit is used to obtain a new matrix after all identical items have been moved, and to perform fusion processing on each column vector to obtain a fused temporal-intent sequence;
[0106] The discrete determination unit is used to obtain a divergence matrix by each column vector corresponding to the first position distribution if the divergence terms of the first intention divergence function of the first column and the second column are different terms, and to obtain the divergence discrete points of each row vector in the divergence matrix, and then count the total number of divergence discrete points of each column vector in the divergence matrix.
[0107] The numerical partitioning unit is used to partition all total quantities according to the same numerical value, calculate the ratio of the number of partition column vectors with the same numerical value to the total number of column vectors in the bifurcation matrix, and filter the maximum ratio.
[0108] A single elimination unit is used to obtain the similarity between each column vector corresponding to the maximum ratio and any adjacent column vector if the value corresponding to the maximum ratio is the maximum value among all the total quantities. The column vector corresponding to the maximum ratio with a similarity less than a preset value is eliminated by single elimination. The column vectors in the overall time series matrix that have not been eliminated by single elimination are fused to obtain a fused time series-intent sequence.
[0109] The complete elimination unit is used to eliminate all column vectors corresponding to the maximum ratio if the value corresponding to the maximum ratio is not the maximum value among all the total quantities, and to perform fusion processing on each column of the retained vectors of the overall time series matrix to obtain the fused time series-intent sequence.
[0110] In this embodiment, the intent divergence model adopts a neural network model. The input is a column vector (intent of each interaction dimension), and the output is the first intent divergence function. The model structure is: input layer (dimension = number of interaction dimensions) → hidden layer (2 layers, with 32 and 16 neurons respectively) → output layer (dimension = number of divergence terms). It is trained with 500 sets of labeled multi-dimensional intent data and has an accuracy of 87%.
[0111] In this embodiment, the interaction situations corresponding to the four dimensions under the same time series should theoretically be consistent. However, there may be differences in the interaction feedback before and after. There may be actual inconsistencies during the time series alignment process. Therefore, it is necessary to perform divergence term analysis to adjust the overall time series matrix to ensure the reliability and accuracy of the sequence.
[0112] In this embodiment, the first intent divergence function consists of multiple divergence terms, each corresponding to the difference in intent across different dimensions in the column vector. For example, if column 1 (time node t1) of the overall time-series matrix is: [I2 (voice), I6 (touch)], and there are differences in intent, the corresponding divergence term of the first intent divergence function is... Column 2 (time node t2) vector is [I3 (voice), I3 (touch)], meaning Figure 1 There are no divergent terms, and the divergence function value is 0.
[0113] The first cluster refers to the set of column vectors with similar divergence characteristics obtained by grouping the first intention divergence functions corresponding to all column vectors (multi-dimensional intentions at the same time point) of the overall time series matrix through cluster analysis algorithms. For example, if the overall time series matrix contains 5 column vectors t0-t4, the number of divergence terms in the corresponding first intention divergence functions of column 1 (t1) and column 3 (t3) are both 2, and the degree of divergence is similar, so they are classified into the same first cluster; the number of divergence terms in column 0 (t0), column 2 (t2), and column 4 (t4) are 1 or 0, so they are classified into another first cluster.
[0114] The first position distribution refers to locking the time node positions corresponding to all column vectors (multi-dimensional intentions at the same time point) in each first cluster. For example, if a first cluster contains column 1 (t1) and column 3 (t3), and the corresponding time points are t1 and t3, then the first position distribution is: {t1: cluster 1, t3: cluster 1}; another cluster contains column 0 (t0), column 2 (t2), and column 4 (t4), and the first position distribution is {t0: cluster 2, t2: cluster 2, t4: cluster 2}.
[0115] In this embodiment, the preset percentage ranges from 50% to 70%, with a default value of 60%, determined based on experimental data from 100 sets of cultural and tourism interactive scenarios. Among them, 68 sets of scenarios achieved an intent fusion accuracy of 89% under this default value. Adjustment is supported according to the scenario type, with 65% for historical and cultural scenarios and 55% for natural scenery scenarios.
[0116] The divergent terms (identical / different terms) refer to the divergent content in the first and second columns' respective initial divergent functions. Identical terms mean the divergent terms in both columns are completely identical in terms of the combination of divergent dimensions and the degree of divergence; different terms mean the divergent terms in both columns differ in terms of the combination of divergent dimensions or the degree of divergence. For example, the divergent terms in the first column (t1) are... The divergence term in the second column (t2) is also... Here, only the dimension combination is consistent; if the degree of divergence is also consistent, then it is considered a common item; if the divergent items in the second column are... If , then they are different terms.
[0117] The total sequence interval refers to the number of intervals between the time nodes corresponding to the first and second columns, i.e., the index difference between the two time nodes. It is used to determine the movement distance of the row vector in the overall time series matrix. For example, if the first column corresponds to time node t1 (column index 1) and the second column corresponds to time node t3 (column index 3), and the index difference between the time nodes is 2, then the total sequence interval = 2. At this time, the corresponding row vector [I1,I2,I3,I1,I4] is shifted 2 positions to the right from the first element, becoming [I1,I2,I1,I2,I3].
[0118] A bifurcation matrix is a submatrix formed by all column vectors corresponding to the first position distribution when the bifurcation terms in the first and second columns are different terms.
[0119] A divergence point refers to a position in a divergence matrix where the intent identifiers of different column vectors differ within a given row vector. This represents a situation where the intent of the same interaction dimension is inconsistent at different time points, and is used to measure the degree of dispersion of the intent in that dimension. For example, the intent sequence for the voice dimension (row 1) is [I2, I3, I1], and all three intents are different, therefore there are two divergence points in this row (the intent differences between t1-t2 and t2-t3); the intent sequence for the touch dimension (row 2) is [I6, I3, I7], and there are two divergence points; the intent sequence for the facial expression dimension (row 3) is [I9, I10, I11], and there are two divergence points.
[0120] In this embodiment, the similarity distribution under different interaction scenarios is statistically analyzed through experiments, and the preset value range is determined to be 0.6-0.8, with a default value set to 0.7. Manual adjustment is supported according to scenario requirements. The similarity calculation adopts the cosine similarity algorithm, which converts the intent identifiers of each dimension of the column vector into vector form before calculating the similarity value.
[0121] The beneficial effects of the above technical solution are as follows: Through steps such as column vector-level first intent divergence function analysis, clustering, and statistical analysis of divergence discrete points, the divergence characteristics of multi-dimensional intents at the same time node are accurately identified, enabling targeted processing of scattered or conflicting intents. The row vector movement strategy for identical divergence items and the column vector filtering and elimination mechanism for different divergence items effectively improve the consistency and accuracy of column vector fusion, solving the problems of chaotic temporal alignment and difficulty in reconciling intent conflicts in traditional fusion algorithms. The final output fused temporal-intent sequence more closely matches the user's actual interaction intent, providing high-quality core data for subsequent sub-interaction surface generation and event extraction.
[0122] This invention provides an interactive experience-driven cultural tourism value-added data system based on IP content and real-world SV, wherein the intent fusion submodule further includes:
[0123] A window function construction unit is used to lock the column vector corresponding to each first cluster in the overall time series matrix as the fourth column if the distribution ratio of the first position distribution is less than a preset ratio, and construct a window function according to the number of clusters of the first cluster.
[0124] The extraction unit is used to divide and extract each fourth column from the overall time series matrix according to the window function to obtain the window matrix corresponding to each fourth column. The window function corresponding to each fourth column contains the fourth column, and its column position is random and not unique.
[0125] An adjustment unit is used to obtain the basic rate of change of each row vector in the window matrix, and adjust the row elements corresponding to the fourth column based on the basic rate of change to obtain the adjusted fourth column;
[0126] The second fusion unit is used to construct a fusion matrix by combining the adjusted fourth column and the unadjusted columns in the overall time series matrix, and to perform fusion processing on each column vector in the fusion matrix to obtain a fused time series-intent sequence.
[0127] In this embodiment, a window function refers to a function used to extract local submatrices (window matrices) from the overall time series matrix. For example, the fourth column is column 1 (t1) and column 3 (t3). The constructed window function is a rectangular window function with column range covering t0-t4 (including the fourth column and adjacent time nodes) and row range covering all interaction dimensions (voice, touch, and expression). It is used to extract the local matrix containing the fourth column.
[0128] In this embodiment, the window matrix refers to a local submatrix extracted from the overall time series matrix through a window function.
[0129] The basic rate of change refers to the degree of change of the intention of each dimension within a column vector of the window matrix, or the degree of change of the intention at different time points within a row vector. The basic rate of change of a column vector = the logarithm of the inconsistency of intention between dimensions / the total number of logarithms of dimensions; the basic rate of change of a row vector = the logarithm of the inconsistency of intention between time points / the total number of logarithms of time points, with a value range of 0-1.
[0130] The matrix to be fused refers to the matrix formed by the adjusted fourth column and the unadjusted column vectors in the overall time series matrix.
[0131] In this embodiment, the column vector of the fourth column is adjusted according to the basic rate of change of the window matrix. The specific execution process is as follows:
[0132] Iterate through each window matrix and calculate two types of basic rates of change: one is the basic rate of change of column vectors (the degree of change of intent across multiple dimensions at the same time point), and the other is the basic rate of change of row vectors (the degree of change of intent across multiple time points within the same interaction dimension). For example, the basic rate of change of column vector in the fourth column 1 of the window matrix is 1 (intents are different in each dimension), and the basic rate of change of row vectors in the speech dimension is 0.75 (intent changes drastically).
[0133] Adjustment rules are formulated based on the fundamental rate of change.
[0134] If the basic rate of change of the column vector is >0.7 (significant conflict in multi-dimensional intents), refer to the dominant intent (the intent with the highest frequency of occurrence) of adjacent column vectors in the reference window matrix, and correct the divergent dimension intents in the fourth column to improve the correlation between intents across dimensions. For example, if the column vector of the fourth column 1 is [I2,I6,I10], and the dominant intent of the adjacent column 0 is inquiry / click, then I2 will be corrected to I1 (inquiry details), which is related to inquiry.
[0135] If the basic rate of change of the row vector is >0.7 (the intention fluctuates drastically within the same dimension), interpolation is used to smooth the intention in the fourth column of that dimension, making the temporal changes more coherent. For example, if the row vector of the speech dimension is [I1,I2,I3], and the basic rate of change is 1, after interpolation, I2 is adjusted to the transitional intention I1.5 between I1 and I3 (inquiry + curiosity).
[0136] It should be noted that the thresholds for the basic rate of change of column vectors and the basic rate of change of row vectors were determined by analyzing the intent fluctuation patterns of 50 sets of user interaction data, which can cover more than 90% of intent conflict scenarios.
[0137] In this embodiment, the adjusted fourth column vector replaces the original column vector in the window matrix and is synchronously updated to the overall time series matrix to obtain the adjusted fourth column.
[0138] The beneficial effects of the above technical solution are: by focusing on the core fourth column through a window function, the distortion of intent caused by global adjustments is avoided; the bidirectional calculation of the basic rate of change accurately captures spatial conflicts and temporal fluctuations of intent; and the targeted adjustment rules effectively improve the stability and relevance of local intents. Finally, through the overall fusion of the matrices to be fused, the output fused temporal-intent sequence not only retains the core information of the original interactive intent, but also solves the problem of intent fragmentation caused by dispersed distribution.
[0139] This invention provides a data-driven system for interactive cultural tourism experiences based on IP content and real-world SV (Site-Based Virtual Reality), wherein the map construction module includes:
[0140] The vector and weight determination submodule is used to extract the feature vector of each feedback map, and simultaneously determine the first weight matrix based on the cultural tourism theme of the target cultural tourism scene and the object attributes of the scene object. Second weight matrix The first weight matrix It is a 4×4 diagonal matrix, with diagonal elements representing facial expression weight, voice weight, touch persistence weight, and physiological weight, respectively. This is the second weight matrix. It is a 4×4 matrix, and each element in the 4×4 matrix This represents the strength of the attribute association between the i-th feature dimension and the j-th feature dimension;
[0141] The comprehensive vector acquisition submodule is used to obtain the feature vector and the first weight matrix. Second weight matrix Determine the comprehensive defect feature vector Z of the scene object;
[0142] ,in, Each element in the facial expression feedback feature vector, voice feedback feature vector, touch dwell feature vector, and physiological feedback feature vector is normalized to the [0,1] interval by Min-Max and then the arithmetic mean of the corresponding vector is taken. For Hadamah accumulation; This is the correlation coefficient, used to control the amplification effect of attribute association;
[0143] The dimension reduction processing submodule is used to perform dimension reduction processing on the comprehensive defect feature vector Z to obtain the comprehensive defect value Sc of the scene object;
[0144] ,in, Denotes the L2 norm of Z; This represents the variance of all elements in Z; , This is the balance coefficient; It is the maximum value among all variances;
[0145] The comprehensive defect map generation submodule is used to generate a comprehensive defect map of the scene object based on the comprehensive curve value Sc and the relative weights of each dimension in the comprehensive defect feature vector Z. The comprehensive defect map includes at least a defect intensity dimension and a defect type distribution dimension.
[0146] In this embodiment, the correlation coefficient ranges from 0.5 to 1.0, with a default value of 0.8, and is used to control the amplification effect of attribute correlation. This value is obtained through fitting multiple sets of defect assessment experiments. The overall defect vector Z has the highest discriminative power when the value is 0.8.
[0147] In this embodiment, , Based on the defect assessment results of 100 sets of scene objects, the comprehensive defect value Sc is determined to ensure that it reflects both the overall strength of the vector and the influence of dimensional variance.
[0148] In this embodiment, the first weight matrix Default value:
[0149] Facial expression weighting: 0.25 (historical and cultural scenarios), 0.3 (parent-child interaction scenarios);
[0150] Voice weighting: 0.35 (historical and cultural scenarios), 0.3 (parent-child interaction scenarios);
[0151] Touch dwell weight: 0.25 (historical and cultural scenarios), 0.3 (parent-child interaction scenarios);
[0152] Physiological weight: 0.15 (historical and cultural scenarios), 0.1 (parent-child interaction scenarios).
[0153] In this embodiment, the facial expression feedback feature vector includes at least three feature dimensions: mean pleasure level (mean displacement amplitude of facial key points), proportion of negative expressions (frame proportion of expressions such as anger and irritability), and frequency of expression changes (number of expression switching times per unit time). The voice feedback feature vector includes at least three feature dimensions: voice emotion tendency value = pleasure level - anger level, speech rate change rate (speech rate standard deviation / average speech rate), and frequency of specific keywords (number of occurrences of keywords related to the cultural tourism theme / total speech duration). The touch dwell feature vector includes at least three feature dimensions: average dwell time (mean dwell time of all interaction points), interaction trigger frequency (number of interactions per unit time), and interaction action complexity (mean curvature of touch trajectory). The physiological feedback feature vector includes at least three feature dimensions: heart rate variability index (RR interval standard deviation), cardiac cycle abnormality ratio (number of abnormal cardiac cycles / total number of cardiac cycles), and physiological arousal level (proportion of the difference between heart rate and baseline heart rate).
[0154] In this embodiment, the attribute association strength refers to the quantitative index of the elements in the second weight matrix, with a value range of 0-1. It is calculated by the scene object attribute similarity algorithm. The larger the value, the closer the attribute association between the i-th feature dimension and the j-th feature dimension. For example, the association strength between the clarity of voice feedback and the completeness of cultural explanation is 0.8.
[0155] The beneficial effects of the above technical solution are as follows: the extraction of four types of feedback feature vectors comprehensively covers the core dimensions of user interaction; the dynamic determination of the first and second weight matrices ensures a high degree of adaptation between defect assessment and scene theme and object attributes; the fusion operation of the comprehensive defect feature vector Z takes into account both the importance of a single dimension and the correlation effect between dimensions; the comprehensive defect value Sc achieves accurate dimensionality reduction of high-dimensional data; the finally generated comprehensive defect map intuitively presents the distribution of defect intensity and type, solves the problems of ambiguity and lack of quantitative standards in traditional defect assessment, and significantly improves the practicality and guiding value of the map construction module.
[0156] This invention provides a data-driven system for interactive cultural tourism experiences based on IP content and real-world SV (Site-Based Virtual Reality) elements. The IP content processing module includes:
[0157] The dimension extraction submodule is used to extract the core attribute dimensions of each element point. The core attribute dimensions include cultural connotation dimension, expression form dimension, and interaction adaptation dimension.
[0158] The priority determination submodule is used to map and match the core attribute dimensions with each defect dimension in the comprehensive defect map of each scene object involved, calculate the matching degree with each scene object involved, and obtain the priority coefficient of each element point.
[0159] The dynamic binding submodule is used to dynamically bind element points to the comprehensive defect map of each scene object in order of priority coefficient. For the defect dimension - IP element point adaptation gap that exists after binding, it calls the related element points in the IP knowledge architecture to supplement the binding and obtain the initial interactive content.
[0160] In this embodiment, the core attribute dimensions include cultural connotation dimensions (the historical significance and cultural value of the element), expression dimensions (the way the element is displayed), and interaction adaptation dimensions (the interaction methods supported by the element), which are used to match the defect dimensions of the scene object. For example, the core attribute dimensions of the IP element point embroidery craft are: cultural connotation dimension: intangible cultural heritage skills and historical inheritance; expression dimension: 3D animation, graphics and text, and handmade simulation; and interaction adaptation dimension: touch operation and voice query.
[0161] In this embodiment, the overall matching degree = Σ (matching degree of each core dimension × corresponding defect dimension weight), with a value range of 0 to 1. At this time, the overall matching degree is the priority coefficient.
[0162] Dynamic binding refers to the process of associating the core attribute dimensions of IP elements with the defect dimensions of scene objects one by one according to the priority coefficient from high to low. That is, based on the mapping relationship library between the core attribute dimensions and the defect dimensions: cultural connotation → insufficient cultural explanation, expression form → monotonous expression form, interaction adaptation → single interaction method. The binding relationship is automatically associated according to the priority order and recorded, and manual adjustment is supported.
[0163] In this embodiment, the dynamic binding process with IP element points is as follows: the cultural connotation, expression form, and interaction adaptation dimension of IP element points are converted into semantic vectors; the defect dimension of the comprehensive defect map is converted into semantic vectors; the cosine similarity of the two types of semantic vectors is calculated, and a binding relationship is established when the similarity is ≥0.6; for the adaptation gaps that still exist after binding (the defect dimension is not covered or the similarity is <0.6), related element points (semantic similarity ≥0.5) are extracted from the IP knowledge architecture for supplementary binding.
[0164] The initial interactive content refers to the initial interactive scheme generated after dynamic binding and supplementary binding of related element points to make up for the defects of scene objects. It includes core information such as interactive form, content theme, and execution logic. For example, for the embroidery display scene, the initial interactive content is touch simulation of embroidery operation (making up for the single interaction) + 3D animation demonstration of embroidery history (making up for the lack of culture) + voice explanation of the characteristics of the craft (making up for the unclear explanation).
[0165] The beneficial effects of the above technical solution are as follows: By extracting the core attribute dimensions of IP elements and establishing a precise matching mechanism with the defect dimensions of scene objects, targeted binding of IP content and scene defects is achieved. The priority coefficient setting ensures that highly adaptable elements play a leading role, and the identification of adaptation gaps and the supplementary binding of related elements achieve comprehensive defect coverage and avoid binding omissions. The final generated initial interactive content can accurately compensate for the interactive experience defects of scene objects, deeply integrating cultural IP with real-world 3D scenes, providing users with an initial experience that combines cultural connotation and interactivity.
[0166] This invention provides a data-driven system for interactive cultural tourism experiences based on IP content and real-world SV (Site-Based Virtual Reality), wherein the content optimization module includes:
[0167] The mapping and association submodule is used to extract the initial feature set of the initial interactive content, and according to the feature attributes of each initial feature in the initial feature set, match the derivative conditions that are consistent with the feature attributes from the attribute-derivative lookup table, and perform mapping and association analysis on each derivative condition with the item attributes of each sub-item of sub-content to obtain the item subset of each sub-content involved in the corresponding initial feature.
[0168] The array construction submodule is used to determine the number of entries in each subset of entries, the total weight of entries, and the main attributes pointing to entries, and to construct a three-dimensional array;
[0169] The position analysis submodule is used to analyze the three-dimensional array of each derived condition under the same initial feature, and determine the placement position of the corresponding entry subset of each three-dimensional array based on the corresponding initial feature. The placement position includes: placing it on the adjacent left, adjacent right, or any randomly divided middle position of the corresponding initial feature.
[0170] The new feature submodule is used to extract new features corresponding to each initial feature after all item subsets have been placed in their respective positions, and to obtain optimized interactive content.
[0171] In this embodiment, the initial feature set refers to the set of features extracted from the initial interactive content that can characterize the core logic and form of the interaction. For example, the initial feature set of the initial interactive content of touch simulation embroidery + 3D animation demonstration + voice explanation is {touch interaction - embroidery operation (attribute: interactive operation class), 3D animation - embroidery history (attribute: visual demonstration class), voice explanation - craft characteristics (attribute: audio explanation class)}.
[0172] The attribute-derivative mapping table refers to the preset feature attribute-derivative condition mapping table. The derivative condition is the type of detailed content that can be added to the initial feature corresponding to the attribute, as shown in Table 3:
[0173] Table 3 Attribute-Derivative Reference Table
[0174]
[0175] In this embodiment, detailed content refers to specific details extracted from user-generated interactive content that can supplement the initial features. It includes multiple detailed entries, each of which is labeled with corresponding entry attributes and derived conditions. Specifically, based on the cultural tourism IP keyword library, which contains keywords such as the history, culture, folk customs, and core elements of the target scene, semantic matching is performed on the user-generated content to filter out detailed content with a relevance of ≥0.7 to the IP content, thus avoiding interference from irrelevant content.
[0176] In this embodiment, the item subset refers to the set of detailed items whose attributes are consistent with the initial feature derivation conditions, selected from the detailed content. For example, if the derivation conditions of the initial feature touch interaction - embroidery operation are operation skills and step refinement, then the item subset is: {embroidery tool selection skills embroidery step breakdown}.
[0177] A three-dimensional array refers to a three-dimensional data structure describing the characteristics of a subset of entries. The dimensions are the number of entries (the number of sub-entries), the total entry weight (the sum of the quality weights of the sub-entries), and the primary attribute pointing to the entry (the entry attribute with the highest frequency of occurrence). For example, the three-dimensional array for the entry subset {Tool Selection Techniques Step Breakdown} is: [2, 1.5, Operation Techniques], where 2 is the number of entries, 1.5 is the total weight (0.8 + 0.7), and Operation Techniques is the primary attribute. The placement logic is based on a decision tree model using the three-dimensional array (number of entries, total entry weight, and primary attribute pointing to the entry) to determine the placement of the sub-content.
[0178] If the total item weight is ≥1.2 and the main item attribute is consistent with the initial feature attribute, place it on the adjacent right side;
[0179] If the total item weight is ≥1.0 and the number of items is ≥3, place it in the middle position;
[0180] In other cases, place it on the adjacent left.
[0181] Optimizing interactive content refers to integrating a subset of items into the initial interactive content according to a predetermined placement, resulting in an optimized interactive solution that supplements user-generated detailed content and better meets user needs. For example, optimized interactive content could include tool selection guidance → touch simulation embroidery with detailed step-by-step prompts + 3D animation with close-up details + dialect explanation of the craft's characteristics with story sharing.
[0182] The beneficial effects of the above technical solution are as follows: By establishing a precise correlation mechanism between initial features and user-generated detailed content, personalized optimization of initial interactive content is achieved. The attribute-derived lookup table ensures the targeting of detailed content selection, and the three-dimensional array analysis scientifically determines the placement of detailed content, avoiding confusion in content integration. The final optimized interactive content integrates the core framework of the system design with high-quality user-generated details, enriching the depth and practicality of the interaction and enhancing user engagement and experience satisfaction.
[0183] This invention provides a data-driven system for interactive cultural tourism experiences based on IP content and real-world SV (Site-Based Virtual Reality) elements. The data-driven module includes:
[0184] The State Field Construction Submodule is used to collect and analyze user behavior data in real time during the interaction process to construct a dynamic user immersion state field.
[0185] The strategy generation submodule is used to generate real-time experience optimization strategies based on the immersion field and the optimized interaction content.
[0186] The guidance submodule is used to output personalized value-added service guidance signals based on the experience optimization strategy and user behavior profile.
[0187] In this embodiment, the state field construction submodule includes:
[0188] Behavioral data sensing unit, used to synchronously collect the user's visual gaze point coordinate sequence. Interactive operation event sequence Voice emotion feature vector and time-series signals of heart rate variability They are all considered as behavioral data;
[0189] Immersion field calculation unit, used to calculate the user's location in the narrative space at time t. Immersion field scalar value at the location :
[0190] ;
[0191] in, For visual fixation point coordinate sequence In position The kernel density estimate within the neighborhood characterizes the degree of visual attention concentration; The frequency of valid interactive events per unit of time; The magnitude of the speech emotion feature vector represents the intensity of emotional arousal. For heart rate variability signals The sympathetic nervous system activity inhibition index obtained through frequency domain analysis characterizes the depth of psychological engagement. , , , These are adaptive weight coefficients that are semantically relevant to the current scene.
[0192] In this embodiment, , , , The values are 0.3, 0.25, 0.25, and 0.2.
[0193] Field gradient analysis unit, used for calculation Gradient in narrative space and time and Identify the spatial evolution path and temporal decay points of user immersion;
[0194] The strategy generation submodule includes:
[0195] The strategy decision function library stores multiple preset optimization strategy base functions. Where C is the feature vector of the current interactive content;
[0196] The policy synthesizer is used to select a matching optimized policy basis function from the policy decision function library based on the output of the field gradient analysis unit, and to select the policy strength coefficients as follows. Perform weighted synthesis to generate real-time experience optimization strategies:
[0197] ,in, , These are the scene tuning parameters; The function is a hyperbolic tangent, which causes the intensity of the optimized intervention to decrease rapidly with increasing immersion (high). And basic immersion level ( (It increases when it is not high.)
[0198] In this embodiment, , The algorithm was trained using 200 sets of user immersion data to ensure that the intensity of the optimized intervention is precisely enhanced when immersion decreases rapidly.
[0199] The bootstrap submodule includes:
[0200] The guiding value potential function calculation unit is used to simultaneously calculate the guiding value potential function for recommending value-added service item j to the user when generating real-time experience optimization strategies. ;
[0201] ,in, It is the sigmoid function; For the preference vector in the user profile; The feature vector of service item j; The target experience distribution of the current optimization strategy Distribution of experiences available with service item j The KL divergence between them is used to measure the consistency of the experience; Current narrative spatial location Location of service item associated scenarios semantic distance; , To adjust the parameters;
[0202] In this embodiment, , This is used to balance the impact of experience consistency and semantic distance on guidance value.
[0203] The triggering and synthesis unit is used when a service item exists. Make The dynamic threshold is triggered to generate the guidance. ,in, Based on the threshold, The adjustment coefficient is used; the unit then generates augmented reality guidance content that is logically coherent in the narrative and corresponds to the instructions of the real-time experience optimization strategy.
[0204] In this embodiment, , Based on experiments on the conversion rate of value-added service recommendations, it was determined that the recommendation conversion rate under this parameter combination reached 28%, which is 15% higher than the default parameters.
[0205] In this embodiment, an immersion field is introduced. This method quantifies and simulates the distribution and flow of user immersion states in the narrative space, fusing discrete, multi-dimensional behavioral data (visual, operational, auditory, and physiological) into a scalar field with clear gradients and evolutionary patterns through a mathematical model. Establish the spatiotemporal variation rate of strategic intervention intensity and immersion field and field strength itself The nonlinear, adaptive relationship between them ensures that the system only applies significant intervention when immersion is rapidly lost and the base level is low.
[0206] Through a potential function containing multiple depth factors Conduct an assessment and determine the strength of the strategy. As a multiplicative factor, it means that the guiding value is directly positively correlated with the necessity and intensity of this optimization intervention. At the same time, the divergence is introduced to measure the consistency between the optimization goal and the value-added service in the experience distribution, ensuring that the guidance and experience optimization are aligned at the philosophical level.
[0207] User behavior data, such as visual gaze point coordinate sequences =[(x1,y1,t1),(x2,y2,t2)], which records the gaze position at different time points; interactive operation event sequence =[Click to start (t1), slide to adjust (t2), complete operation (t3)]; Voice emotion feature vector =[0.8,0.2,0.6] (pleasure, anger, surprise); heart rate variability signal =[70,72,68] (times / minute).
[0208] The user immersion state field refers to a dynamic field constructed based on user behavior data, expressed through immersion field scalar values. Quantifying the user's location in time t and narrative space The level of immersion includes the spatial evolution path (the spatial trajectory of immersion changes) and the time decay point (the moment when immersion rapidly decreases). The user's immersion level in t2 and r2 (embroidery operation area) is... =0.85 (high immersion), t5, r5 (text area) =0.3 (low immersion), t5 is the time decay point; the spatial evolution path is r1→r2→r3→r5 (immersion first increases and then decreases).
[0209] Experience optimization strategies refer to real-time optimization solutions generated based on the immersion state field and optimized interactive content, used to improve user immersion. These include content adjustments, interaction optimization, and pacing control. For example, if a user's immersion decreases in the text area (…), then… =0.3), generating a text-to-interactive question-and-answer strategy; high immersion in the operation area ( =0.85), generating a strategy to extend operation time.
[0210] The beneficial effects of the above technical solution are as follows: the field perception-strategy generation-guided decision chain constructed by mathematical modeling realizes the deep coupling and collaborative adaptation of the three subsystems under a unified mathematical model. Experience optimization and value-added guidance are no longer two independent subsequent links, but a symbiotic process that shares the same set of immersion field state inputs and is driven by the coupling formula. This solves the problem of the separation or even conflict between optimization and guidance in the existing technology, and generates significant synergistic gains in ensuring the smoothness of the experience and improving the efficiency of business conversion.
[0211] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A data-driven system for interactive cultural tourism experiences based on IP content and real-world SV, characterized in that, include: The scene 3D modeling module is used to perform 3D digital reconstruction of the target cultural tourism scene, track the historical virtual interaction screens of each scene object in the reconstructed 3D space, and parse each historical virtual interaction screen according to the range of coarse elements of the screen specified by the interaction command to construct the event interaction type-interaction feedback sequence. The graph construction module is used to extract all event interaction types and interaction feedback sequences under the same scene object according to the same event interaction type, to obtain the interaction feedback set of the same event interaction type, and to construct facial expression feedback graph, voice feedback graph, touch dwell feedback graph and physiological fluctuation change graph. In addition, by combining the cultural and tourism theme of the target cultural and tourism scene and the object attributes of the scene object, a comprehensive defect graph of the scene object is obtained. The IP content processing module is used to parse the original cultural IP materials of the target cultural tourism scene and construct an IP knowledge architecture. It dynamically binds each element point in the IP knowledge architecture with the comprehensive defect map of each scene object involved in the element point to generate initial interactive content. The content optimization module is used to obtain original interactive content submitted by users, extract detailed content from the original interactive content that matches the scene objects involved in each element point, and optimize the initial interactive content according to the detailed content to obtain optimized interactive content. The value-added module is used to collect and analyze user behavior data in real time during the interaction process, and combine it with the optimized interaction content to generate experience optimization strategies and user behavior profiles, and output personalized value-added service guidance signals.
2. The interactive experience cultural tourism value-added data-driven system according to claim 1, characterized in that, The scene 3D modeling module includes: The single-dimensional intent determination submodule is used to obtain a decomposition model that matches the interaction dimension and dimension quantity of the interaction instruction from the dimension-quantity-decomposition lookup table, input the interaction instruction into the decomposition model, and output the single-dimensional temporal-intent sequence of each interaction dimension, wherein the interaction dimension includes: voice dimension, touch dimension, facial expression dimension or physiological dimension. The intent fusion submodule is used to fuse the single-dimensional temporal-intent sequences under all interaction dimensions to obtain a fused temporal-intent sequence, and to perform temporal alignment processing with the scene interaction behavior chain of the corresponding historical virtual interaction scene to obtain a sub-interaction surface that is in the same temporal sequence as the fused temporal-intent. The mapping analysis submodule is used to perform screen mapping analysis on the fused time-intention of the same time sequence, and apply it to the corresponding sub-interaction surface to perform coarse selection of screen boundaries, and extract the existing elements of the coarsely selected screen. The position locking submodule is used to lock the screen positions of all existing elements on all sub-interaction surfaces under the same interaction command, and to perform salience processing on the corresponding locked positions according to the locking frequency of each locked position and the screen attributes of the locked position to obtain the coarse element range of the screen. The event acquisition submodule is used to align the coarse element range of the screen with the historical virtual interactive screen, perform fine event acquisition on the aligned part according to the saliency processing result, and perform coarse event acquisition on the remaining part of the historical virtual interactive screen. The feedback and event association submodule is used to obtain the first capture feedback based on each fine-grained event and the second capture feedback based on each coarse-grained event, and associate them with the corresponding fine-grained events and coarse-grained events to obtain the event interaction type-interaction feedback sequence.
3. The interactive experience cultural tourism value-added data-driven system according to claim 2, characterized in that, The intent fusion submodule includes: The function determination unit is used to align and sort all single-dimensional time-intent sequences according to the time sequence to obtain the overall time sequence matrix, and input each column vector into the intent divergence model to obtain the first intent divergence function, wherein each divergence term of the first intent divergence function corresponds one-to-one with the corresponding interaction dimension. The clustering analysis unit is used to perform clustering analysis on all first intentional divergence functions, determine several first clusters, and lock the time nodes corresponding to all column vectors in the first clusters to obtain the first position distribution; The proportion determination unit is used to determine if the proportion of the first position distribution is greater than or equal to a preset proportion, and then obtain the column vector corresponding to the first position in the first position distribution, which is regarded as the first column; at the same time, obtain the column vector corresponding to the second position in the first position distribution, which is regarded as the second column. The moving unit is configured to determine the total sequence interval between the first column and the second column if the divergence terms of the first intention divergence function of the first column and the second column are the same, and to move the row vector corresponding to the same term to the right by the total sequence interval number starting from the first element, while keeping the elements in the first column before the position corresponding to the same term unchanged. The first fusion unit is used to obtain a new matrix after all identical items have been moved, and to perform fusion processing on each column vector to obtain a fused temporal-intent sequence; The discrete determination unit is used to obtain a divergence matrix by each column vector corresponding to the first position distribution if the divergence terms of the first intention divergence function of the first column and the second column are different terms, and to obtain the divergence discrete points of each row vector in the divergence matrix, and then count the total number of divergence discrete points of each column vector in the divergence matrix. The numerical partitioning unit is used to partition all total quantities according to the same numerical value, calculate the ratio of the number of partition column vectors with the same numerical value to the total number of column vectors in the bifurcation matrix, and filter the maximum ratio. A single elimination unit is used to obtain the similarity between each column vector corresponding to the maximum ratio and any adjacent column vector if the value corresponding to the maximum ratio is the maximum value among all the total quantities. The column vector corresponding to the maximum ratio with a similarity less than a preset value is eliminated by single elimination. The column vectors in the overall time series matrix that have not been eliminated by single elimination are fused to obtain a fused time series-intent sequence. The complete elimination unit is used to eliminate all column vectors corresponding to the maximum ratio if the value corresponding to the maximum ratio is not the maximum value among all the total quantities, and to perform fusion processing on each column of the retained vectors of the overall time series matrix to obtain the fused time series-intent sequence.
4. The interactive experience cultural tourism value-added data-driven system according to claim 3, characterized in that, The intent fusion submodule also includes: A window function construction unit is used to lock the column vector corresponding to each first cluster in the overall time series matrix as the fourth column if the distribution ratio of the first position distribution is less than a preset ratio, and construct a window function according to the number of clusters of the first cluster. The extraction unit is used to divide and extract each fourth column from the overall time series matrix according to the window function to obtain the window matrix corresponding to each fourth column. The window function corresponding to each fourth column contains the fourth column, and its column position is random and not unique. An adjustment unit is used to obtain the basic rate of change of each row vector in the window matrix, and adjust the row elements corresponding to the fourth column based on the basic rate of change to obtain the adjusted fourth column; The second fusion unit is used to construct a fusion matrix by combining the adjusted fourth column and the unadjusted columns in the overall time series matrix, and to perform fusion processing on each column vector in the fusion matrix to obtain a fused time series-intent sequence.
5. The interactive experience cultural tourism value-added data-driven system according to claim 1, characterized in that, The map construction module includes: The vector and weight determination submodule is used to extract the feature vector of each feedback map, and simultaneously determine the first weight matrix based on the cultural tourism theme of the target cultural tourism scene and the object attributes of the scene object. Second weight matrix The first weight matrix It is a 4×4 diagonal matrix, with diagonal elements representing facial expression weight, voice weight, touch persistence weight, and physiological weight, respectively. This is the second weight matrix. It is a 4×4 matrix, and each element in the 4×4 matrix This represents the strength of the attribute association between the i-th feature dimension and the j-th feature dimension; The comprehensive vector acquisition submodule is used to obtain the feature vector and the first weight matrix. Second weight matrix Determine the comprehensive defect feature vector Z of the scene object; The dimension reduction processing submodule is used to perform dimension reduction processing on the comprehensive defect feature vector Z to obtain the comprehensive defect value Sc of the scene object; The comprehensive defect map generation submodule is used to generate a comprehensive defect map of the scene object based on the comprehensive curve value Sc and the relative weights of each dimension in the comprehensive defect feature vector Z. The comprehensive defect map includes at least a defect intensity dimension and a defect type distribution dimension.
6. The interactive experience cultural tourism value-added data-driven system according to claim 1, characterized in that, The IP content processing module includes: The dimension extraction submodule is used to extract the core attribute dimensions of each element point. The core attribute dimensions include cultural connotation dimension, expression form dimension, and interaction adaptation dimension. The priority determination submodule is used to map and match the core attribute dimensions with each defect dimension in the comprehensive defect map of each scene object involved, calculate the matching degree with each scene object involved, and obtain the priority coefficient of each element point. The dynamic binding submodule is used to dynamically bind element points to the comprehensive defect map of each scene object in order of priority coefficient. For the defect dimension - IP element point adaptation gap that exists after binding, it calls the related element points in the IP knowledge architecture to supplement the binding and obtain the initial interactive content.
7. The interactive experience cultural tourism value-added data-driven system according to claim 1, characterized in that, The content optimization module includes: The mapping and association submodule is used to extract the initial feature set of the initial interactive content, and according to the feature attributes of each initial feature in the initial feature set, match the derivative conditions that are consistent with the feature attributes from the attribute-derivative lookup table, and perform mapping and association analysis on each derivative condition with the item attributes of each sub-item of sub-content to obtain the item subset of each sub-content involved in the corresponding initial feature. The array construction submodule is used to determine the number of entries in each subset of entries, the total weight of entries, and the main attributes pointing to entries, and to construct a three-dimensional array; The position analysis submodule is used to analyze the three-dimensional array of each derived condition under the same initial feature, and determine the placement position of the corresponding entry subset of each three-dimensional array based on the corresponding initial feature. The placement position includes: placing it on the adjacent left, adjacent right, or any randomly divided middle position of the corresponding initial feature. The new feature submodule is used to extract new features corresponding to each initial feature after all item subsets have been placed in their respective positions, and to obtain optimized interactive content.
8. The interactive experience cultural tourism value-added data-driven system according to claim 1, characterized in that, The drive value-added module includes: The State Field Construction Submodule is used to collect and analyze user behavior data in real time during the interaction process to construct a dynamic user immersion state field. The strategy generation submodule is used to generate real-time experience optimization strategies based on the immersion field and the optimized interaction content. The guidance submodule is used to output personalized value-added service guidance signals based on the experience optimization strategy and user behavior profile.