Augmented reality content generation method
By combining AI and AR technologies, personalized virtual characters and storyline data are generated, solving the problem of insufficient intelligent and immersive experience in the cultural and tourism field of existing AR technology. This enables intelligent, personalized and immersive cultural and tourism experiences and builds an economic ecosystem for virtual assets.
Patent Information
- Application Number
- CN202511075258.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-07
AI Technical Summary
Existing AR technology applications in the cultural and tourism sector are mainly focused on scenic spot navigation and information display, lacking intelligent and personalized immersive experiences, and failing to meet the diverse needs of users.
By combining AI and AR technologies, personalized virtual characters and storyline data are generated by acquiring user location information, personalized input parameters, and physical motion data. Virtual asset data is stored in the blockchain, enabling real-time rendering of virtual characters and execution of storylines in AR scenes.
It provides a more intelligent, personalized, and immersive cultural and tourism experience, enhances user interactivity and participation, and builds an economic ecosystem in augmented reality scenarios.
Smart Images

Figure CN120912828A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of augmented reality technology, and in particular to a method for generating augmented reality content. BACKGROUND
[0002] With the rapid development of information technology, augmented reality (AR) technology is increasingly widely used in the field of tourism. Tourism activities have cultural, travel, sightseeing and entertainment attributes, and AR technology can better fit local culture, history and natural scenery to provide immersive experiences for tourists. However, existing AR tourism applications still have many limitations.
[0003] Currently, AR technology applications in the field of tourism mainly focus on site guide and information display. For example, in an invention patent named "Digital interactive smart tourism system" (application number CN202210713808.9), a digital interactive smart tourism system is provided, which performs feature recognition through AR scanning and overlays virtual interactive information on the actual scene to realize interactive scene tour interaction.
[0004] Moreover, with the development of artificial intelligence (AI) technology, AI technology is increasingly combined with the field of tourism. The combination of the advantages of AI technology and AR technology in tourism projects can create better interactive experiences for tourism users and better meet the needs of users in terms of intelligent and personalized generation of tourism scenes.
[0005] Therefore, there is an urgent need in the prior art for a method for generating augmented reality content that combines AI technology and AR technology, which is used in the field of tourism to provide more intelligent, personalized and immersive tourism experiences. SUMMARY
[0006] The embodiments of the present application provide a method for generating augmented reality content, which realizes a method for generating augmented reality content that combines AI technology and AR technology, which is used in the field of tourism to provide more intelligent, personalized and immersive tourism experiences.
[0007] In a first aspect, the embodiments of the present application provide a method for generating augmented reality content, comprising: obtaining first augmented reality model data of a travel scene according to position information of a user in the travel scene; obtaining personalized input parameters of the user, inputting the personalized input parameters into an augmented reality content large model to generate second virtual character model data and virtual plot data corresponding to the user; obtaining physical motion data of the user, controlling the second virtual character model to be rendered and loaded in the first augmented reality model data in real time, and realizing triggering and executing of the virtual plot data; and generating and saving virtual asset data corresponding to execution result data of the virtual plot data in a blockchain according to the execution result data.
[0008] Optionally, in the above method, the first augmented reality model data of the travel scene is obtained according to the position information of the user in the travel scene, comprising: pre-caching augmented reality model data resources of candidate scenic spots on a travel route of the user; obtaining real-time position information of the user in the travel scene; calculating real-time distance data between the user and the scenic spots, and dynamically adjusting a first model precision of the first augmented reality model data of the travel scene; and obtaining augmented reality model data corresponding to the first model precision from the augmented reality model data resources of the candidate scenic spots.
[0009] Optionally, the personalized input parameters of the user include at least one of user identity data, user appearance feature parameters, user interest and behavior preference parameters, and user scene and state parameters.
[0010] Optionally, in the above method, the second virtual character model data corresponding to the user is generated by the augmented reality content large model, comprising: obtaining a fusion feature vector formed by role facial feature data set, body feature data set, and interest preference-style mapping data set; training an augmented reality 3D character model generative large model according to fusion feature vector and user personalized image constraint conditions; the augmented reality 3D character model generative large model generates a 3D face mesh based on input user facial feature vector, and matches user face contour through a Poisson fusion algorithm; the augmented reality 3D character model generative large model generates a body mesh based on appearance feature parameters of the user; the augmented reality 3D character model generative large model generates costumes, textures and corresponding role actions based on interest preference and style parameters of the user, and realizes real-time role action driving of the 3D character model through skeleton binding.
[0011] Optionally, the method further comprises: fine-tuning a first plot text large language model related to the user's travel scene; training a second scene knowledge graph storing structured data of historical events and / or cultural knowledge of the user's travel scenic spot; mixing the first plot text large language model and the second scene knowledge graph to form the augmented reality content large model; retrieving first plot theme content from the second scene knowledge graph according to geographical location data of the travel scene and cultural preference data of the user; generating a second plot structure design through the first plot text large language model according to the user's plot cultural type preference; embedding a third augmented reality interaction instruction in the plot through the first plot text large language model according to the user's interaction mode preference; and triggering generation of a fourth plot branch switch in the plot through the first plot text large language model according to the user's real-time emotional parameters.
[0012] Optionally, the user's physical motion data comprises at least one of body posture data, motion trajectory data, and interaction force data.
[0013] Optionally, the method further comprises: mapping body motion data in the user's physical motion data to a skeletal coordinate system of the second virtual character; converting a three-dimensional coordinate system of the user's physical motion data to a spatial coordinate system of the first augmented reality model data; binding the user's physical motion data to the second virtual character's skeleton and automatically calculating associated joint posture data for supplementation; rendering the virtual character in real time and updating the relative position of the virtual character and the spatial scene of the first augmented reality model data.
[0014] Optionally, the method further comprises: setting a plot trigger condition template according to the virtual plot data, wherein the plot trigger condition comprises a virtual character action type, an action parameter threshold, and a travel scene context; comparing the user's real-time motion data with the template, and determining virtual plot trigger execution when the trigger condition is met.
[0015] Optionally, the virtual plot data execution result data comprises at least one of virtual plot data execution result data generated by the augmented reality content large model according to the user's virtual character model data, context data when the user triggers the virtual plot, travel scene data of the user, and user interest and behavior preference data of the user.
[0016] Optionally, in the method, the generating and saving, in the blockchain, of virtual asset data corresponding to the virtual plot data execution result data according to the virtual plot data execution result data comprises: performing signature calculation on the virtual plot data execution result data using the private key data of the user, combining the user public key, the signature calculation result and the virtual plot data execution result data, and performing on-chain saving to form the virtual asset data of the user; wherein the virtual plot data execution result data comprises an augmented reality virtual asset loaded and displayed in the augmented reality tourism scene.
[0017] According to the present application, the first augmented reality model data of the tourism scene is obtained, and the second virtual role model data and the corresponding virtual plot data are generated for the user by the augmented reality content large model according to the user personalized data; then the second virtual role model can be rendered and loaded in the first augmented reality model data in real time, and the triggering and execution of the virtual plot data can be realized; in this way, in the augmented reality scene of the present application, the user virtual role generated by the AI large model can be fused into the AR scene, and the triggering and execution of the virtual plot data in the augmented reality scene, so that the intelligent, personalized and immersive experience of the augmented reality tourism experience is stronger. Moreover, according to the virtual plot data execution result, the virtual asset data corresponding to the virtual plot data execution result is generated and saved in the blockchain, and then the economic ecology of the tourism activity in the augmented reality scene is formed. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0019] Figure 1 A flowchart of a generation method of augmented reality content provided for the first embodiment of the present application; Figure 2 A structure diagram realized in the augmented reality content generation platform system provided for the second embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0021] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plural" generally includes at least two.
[0022] Depending on the context, the word "if" as used herein can be interpreted to mean "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined" or "if (a stated condition or event) is detected" can be interpreted to mean "when it is determined" or "in response to determining" or "when (a stated condition or event) is detected" or "in response to detecting (a stated condition or event)".
[0023] In addition, the step sequence in each of the following method embodiments is only an example and is not strictly limited.
[0024] In the embodiments of the present application, the first augmented reality model data of the tourist scene is acquired, and the second virtual role model data and the corresponding virtual plot data are generated for the user by the augmented reality content large model according to the user personalized data; then the second virtual role model can be controlled to be rendered and loaded in the first augmented reality model data in real time, and the trigger execution of the virtual plot data is realized; in this way, in the augmented reality scene of the embodiments of the present application, the user virtual role generated by the AI large model can be fused into the AR scene, and the trigger execution of the virtual plot data is performed in the augmented reality scene, so that the travel and tourism experience in the augmented reality is more intelligent, personalized and immersive. In addition, according to the execution result of the virtual plot data, the virtual asset data corresponding to the execution result of the virtual plot data is generated and saved in the blockchain, and then an economic ecology of the travel and tourism activities in the augmented reality scene is formed.
[0025] Embodiment one: Figure 1 A flowchart of a generation method of augmented reality content provided by the first embodiment of the present application. The generation method of augmented reality content in the first embodiment includes the following steps: Step 100: Based on the user's location information in the tourism scenario, obtain the first augmented reality model data of the tourism scenario. In this embodiment, the first augmented reality model data of the tourism scene can be obtained based on the user's location information in the tourism scene, and then the AR model data of the tourism scene where the user is located can be rendered in real time to load the AR scene for immersive interaction.
[0026] Preferably, in this embodiment, augmented reality model data resources of candidate attractions along the user's travel route can be pre-cached; real-time location information of the user in the travel scenario can be obtained; real-time distance data between the user's real-time location and the attractions can be calculated, and the first model precision of the first augmented reality model data of the travel scenario can be dynamically adjusted; and augmented reality model data corresponding to the first model precision can be obtained from the augmented reality model data resources of the candidate attractions.
[0027] In this embodiment, this step achieves accurate matching between the location information of the tourism scene and the augmented reality (AR) model. The specific implementation process is as follows: Pre-cached candidate attraction AR model resources Based on the user's booked tour route (e.g., "Forbidden City day trip") and historical tour data, the system predicts three core attractions (e.g., Hall of Supreme Harmony, Hall of Central Harmony, and Hall of Preserving Harmony) and two alternative attractions (e.g., Palace of Heavenly Purity and Imperial Garden).
[0028] Edge computing nodes (deployed on edge servers within the Forbidden City scenic area, with a computing power ≥8 TOPS) are used to pre-cache the AR model resources of the aforementioned attractions, including: High-poly model resources: ≥20,000 polygons, including texture details (such as the carved beams and painted rafters of the Hall of Supreme Harmony), file size ≤500MB; Medium model resources: 5,000-10,000 polygons, retaining core structural details, file size ≤200MB; Low-poly resources: number of polygons ≤ 1,000, only outline features are retained, file size ≤ 50MB.
[0029] Pre-caching trigger mechanism: caching is automatically started when the user is 1km away from the attraction, and the caching completion time is ≤10 seconds (based on 5G network environment, download speed ≥100Mbps).
[0030] Get user real-time location information Utilizing GPS+SLAM fusion positioning technology: The GPS module (accuracy ≤ 3 meters) obtains the user's latitude and longitude coordinates (e.g., 39.9165° N, 116.3971° E). The mobile device built-in SLAM (Simultaneous Localization and Mapping) engine (based on ARFoundation 5.0) constructs a local environment map, and corrects GPS (Global Positioning System) drift (in areas such as the Forbidden City building complex where GPS signal is weak, positioning accuracy is improved to ≤1 meter); Synchronously collect mobile device gyroscope data (sampling rate 100Hz), obtain device attitude angle (pitch angle, yaw angle, roll angle), and use it for AR model attitude matching.
[0031] Dynamically adjust the accuracy of the first augmented reality model data Calculate the straight-line distance between the user and the target scenic spot (based on the Haversine formula), for example: When the distance ≥1km, call low-mode resources (such as when the user is outside the Meridian Gate of the Forbidden City, load the Taihe Palace low-mode); When 500m < distance < 1km, call medium-mode resources (such as when the user is in the Taihe Gate area, load the Taihe Palace medium-mode); When the distance ≤500m, call high-mode resources (such as when the user enters the Taihe Palace square, load the Taihe Palace high-mode).
[0032] Precision switching mechanism: adopt a "progressive replacement" strategy to avoid model loading lag - first display the current precision model, asynchronously load higher precision models in the background, and after loading is completed (progress ≥90%), smoothly switch, with a switching delay ≤300ms.
[0033] Get AR model data of corresponding precision Call the target model from the edge node pre-cached resources. If the cache is not hit (such as when the user temporarily deviates from the route to the Imperial Garden), download it through the cloud CDN (Content Delivery Network) acceleration (delay ≤500ms).
[0034] After loading the model, convert the local coordinates of the AR model to the WGS84 coordinate system through the coordinate conversion algorithm (Bursa-Wolf model), align it with the user's real-time position, and ensure that the model is "anchored" in the real scene (position deviation ≤5cm).
[0035] Step 101: Obtain the user's personalized input parameters, input them into the augmented reality content large model to generate the second virtual role model data and virtual plot data corresponding to the user In this embodiment, the user's personalized input parameters preferably include at least one of user identity data, user appearance feature parameters, user interest and behavior preference parameters, and user scene and state parameters.
[0036] In this embodiment, this step realizes personalized travel experience by collecting multi-dimensional user personalized input data and generating exclusive virtual character and virtual plot data using an AI large model, and the specific implementation process is as follows: Obtain user personalized input parameters User identity data: including name ("Zhang San"), age (30 years old), gender (male), language preference (Chinese), input through user terminal APP form and associated with user ID (UUID: a1b2c3d4-...).
[0037] User appearance feature parameters: Facial features: take 3 multi-angle photos through the front camera, use MediaPipe Face Mesh to extract 68 facial key points, and generate a 128-dimensional feature vector (error ≤2 pixels); Hair style and clothing: user uploads daily photos, system identifies hair style ("short hair"), hair color ("black") and clothing style ("casual style") through ResNet-50 model, and converts into labeled data.
[0038] User interest and behavior preference parameters: Cultural preference: user checks "Ming and Qing history" and "architectural art" labels, and combines with user terminal APP internal history browsing records (such as viewing "the construction technology of Taihe Palace" article) to strengthen the weight; Plot type preference: select "history restoration type" and "light puzzle type" through questionnaire; Interaction preference: select "gesture interaction" and "voice command".
[0039] User scene and state parameters: Real-time emotion: capture facial expressions through camera, use FER+ model to analyze happiness (0.8 / 1.0) and arousal (0.6 / 1.0); Companion: identify 2 people in the picture and mark as "friend" relationship.
[0040] The embodiment details of constructing the augmented reality content large model and generating corresponding content are as follows: 1. The augmented reality content large model architecture for generating the virtual plot data corresponding to the user: adopts a hybrid architecture of "large language model + knowledge graph": The first plot text large language model is based on LLaMA-2-13B fine-tuning, and the training data includes 50,000 Palace-related historical documents and 30,000 travel and drama plot scripts, and the perplexity value after fine-tuning is reduced to 8.2. The second scene knowledge graph stores structured data such as Palace site entities (such as “Hall of Supreme Harmony”), historical events (such as “Kangxi Coronation Ceremony” and “Transmission of the Imperial Edict”), and cultural common sense (such as “Pavement” technology), containing more than 100,000 triples (entity-relation-entity).
[0041] The mixed architecture of the first plot text large language model and the second scene knowledge graph constitutes the augmented reality content large model.
[0042] 2. The augmented reality content large model for generating the second virtual role model data corresponding to the user obtains a fusion feature vector formed by a large number of 3D role facial feature data sets, body feature data sets, and interest preference-style mapping data sets. According to the fusion feature vector and the user personalized image constraint condition, an augmented reality 3D role model generative large model is trained. 3. Generating second virtual role model data Fusion feature vector construction: converting user personalized parameters into a 512-dimensional vector (text description parameters 384-dimensional + facial feature vector 128-dimensional).
[0043] 3D role generation process: Facial modeling: based on the facial feature vector, a 3D face mesh (8,000 vertices) is generated by StyleGAN3, and a Poisson fusion algorithm is used to match the user's face contour (similarity ≥ 90%); Body modeling: generating a body mesh (3,000 polygons) according to the user's input height (175 cm) and body type (“average”), and matching “Ming and Qing” costumes (based on user interest preferences); Action binding: based on the user's input interest preference “Ming and Qing style” and style parameter “ritual scholar”, generate costumes, textures and corresponding role actions, and realize real-time role action driving of 3D role model through bone binding, this embodiment presets 100+ action library (such as “ritual action” and “walking”), realizes action driving through bone binding (32 joint nodes), and the response delay is ≤50ms.
[0044] Output result: generate a user-specific “Ming and Qing character” virtual role (glTF format, file size ≤10MB).
[0045] 4. Generating virtual plot data For example, a plot theme search is performed: based on the user's location (the Hall of Supreme Harmony) and cultural preferences ( "Ming and Qing history" ), the theme content of "the Imperial Examination Ceremony in the Hall of Supreme Harmony" is searched from the knowledge graph.
[0046] Plot structure design: Beginning: the virtual character (user avatar) receives the task of "newly admitted scholars in the Imperial Examination Ceremony"; Development: participate in the behavior or language of the newly admitted scholars in the Imperial Examination Ceremony by gestures, actions, and language interactions (such as "unfolding the scroll") ; Climax: answer 3 historical questions of the newly admitted scholars in the Imperial Examination Ceremony (such as "the origin of the plaque in the Hall of Supreme Harmony", "the names of the first three scholars in the first rank", "recite the examination paper of the ancient top scorer"), and complete the task if the correct rate is ≥2 / 3; Interactive instruction embedding: for example, when the user performs the correct "bowing" action, trigger the dialogue interaction with the virtual character, such as the emperor.
[0047] Branch switching mechanism: if the user's emotional pleasure degree <0.3, automatically skip the puzzle section and directly enter the plot climax.
[0048] Step 102: acquire the user's physical motion data, control the second virtual character model to be rendered and loaded in the first augmented reality model data in real time, and realize the triggering and execution of the virtual plot data In this embodiment, preferably, the real-time control of the second virtual character model to be rendered and loaded in the first augmented reality model data includes: mapping the body action data in the user's physical motion data to the skeletal coordinate system of the second virtual character; converting the three-dimensional coordinate system of the user's physical motion data to the spatial coordinate system of the first augmented reality model data; binding the user's physical motion data to the skeleton of the second virtual character, and automatically calculating and supplementing the associated joint pose data; rendering the virtual character in real time, and updating the relative position of the virtual character and the spatial scene of the first augmented reality model data.
[0049] In this embodiment, preferably, the realization of the triggering and execution of the virtual plot data includes: setting a plot trigger condition template according to the virtual plot data, wherein the plot trigger condition includes: virtual character action type, action parameter threshold, and tourism scene context; comparing the user's real-time motion data with the template, and determining the triggering and execution of the virtual plot when the trigger condition is met.
[0050] In this embodiment, this step realizes the real-time linkage and rendering loading of the user's physical action and the virtual character, as well as the dynamic triggering of the plot, and the specific implementation process is as follows: Acquire the user's physical motion data Multi-sensor fusion collection (sampling frequency as shown in Table 1): IMU is an inertial measurement unit Data type Specific parameter Collecting device Sampling frequency Body posture data Head Euler angle (pitch 30°) Gyroscope + IMU 100 Hz Motion trajectory data Right hand three-dimensional coordinates (X = 0.5 m, Y = 1.2 m) Binocular camera + SLAM 30 Hz Interaction force data Screen click pressure (0.4 N) Pressure-sensitive screen 50 Hz Table 1 Data preprocessing: remove noise by Kalman filtering and retain effective motion signals (such as filtering hand tremor).
[0051] Real-time control of rendering loading of the second virtual role Bone mapping: map user motion data to virtual role bone coordinate system, for example: User head pitch 30°--》virtual role head pitch 30° synchronously; User right hand movement trajectory--》virtual role right hand moves along the same trajectory (smoothed by Bezier curve).
[0052] Coordinate system conversion: use 4x4 transformation matrix to convert user hand three-dimensional coordinates (device local coordinate system) to AR scene coordinate system (Taihe Hall scene), ensure that the virtual role hand is aligned with the "scroll" model in the AR scene (deviation ≤3cm).
[0053] Inverse kinematics completion: when the user's left arm is blocked, calculate the left arm posture by IK algorithm (Inverse Kinematics algorithm), error ≤5°, to avoid distortion of virtual role motion.
[0054] Real-time rendering optimization: Use Vulkan API rendering engine, frame rate stability ≥60fps; Enable Occlusion Culling technology, when the virtual role is blocked by the "pillar" model in the AR scene, only render the visible part.
[0055] The scheme for triggering and executing virtual plot data is as follows: Plot trigger condition template definition (take "Lai Fu Da Dian Xing Li" plot as an example): { "action type": "hand trajectory matching", "parameter threshold": {"trajectory similarity ≥90%", "duration ≥2s"}, "scene context": "user is near the Danping of Taihe Hall (GPS error ≤2m) "} Real-time comparison: use dynamic time warping (DTW) algorithm to compare user hand salute motion trajectory with template, similarity calculation result is 92% (≥90%), and duration is 2.3s (≥2s), it is judged that the trigger condition is met.
[0056] In this embodiment, preferably, the plot trigger condition template is set according to the virtual plot data. The plot trigger condition includes a virtual character action type, an action parameter threshold value, and a travel scene context. Then, the user real-time motion data is compared with the template, and when the trigger condition is met, it is determined that the virtual plot trigger is executed.
[0057] Plot execution process: Load the AR special effect of “Huang Bang” (particle quantity 1,000+, rendering time ≤100ms); Play the vocal name voice of the virtual character in the plot (TTS synthesis, voice similarity and user voice line matching degree ≥85%); Trigger vibration feedback (mobile phone motor vibration frequency 200Hz, duration 500ms).
[0058] Step 103: According to the virtual plot data execution result data, generate and save the virtual asset data corresponding to the virtual plot data execution result data in the blockchain In this embodiment, preferably, the virtual plot data execution result data includes at least one of the virtual character model data of the user, the context data when the user triggers the virtual plot, the travel scene data of the user, and the user interest and behavior preference data of the user. The virtual plot data execution result data generated by the augmented reality content large model. In this embodiment, preferably, in addition to the virtual plot data execution result data generated by the augmented reality content large model, a virtual plot data execution result data with a virtual asset attribute can also be given after triggering the preset plot execution condition.
[0059] In this embodiment, this step realizes the right and evidence of virtual assets through blockchain technology, and constructs a digital economy ecological environment for tourism, and the specific implementation process is as follows: Generation of virtual plot data execution result data Result data structure: { "result_id": "R20240715001", "user_id": "a1b2c3d4-...", "plot_id": "P_GUGONG_001", "content": { "status": "success", "score": 85, / / plot completion degree "rewards": ["Fragment of Forbidden City Cultural and Creative × 3", "Jinshi and Dizi's Ornament Skin"], "scene_data": {"gps": "39.9165°N, 116.3971°E", "timestamp": "2024-07-15T10:30:00Z"}, "hash": "0x5f...3a" / / SHA-256 Hash Value Generation Logic: Based on user's plot completion degree (e.g., 85 points), virtual character data (e.g., Jinshi and Dizi's Ornament), and scene data (e.g., the Palace of Heavenly Purity), the augmented reality content large model generates the content.
[0060] Blockchain Notarization Process Consensus Mechanism: 3-node PBFT (Practical Byzantine Fault Tolerance) consensus mechanism (Palace Museum server, tourism platform, and third-party audit node), block generation time ≤ 2s.
[0061] Signature and Chain: User Client Calls Web3.js SDK (Software Development Kit) to use user private key (stored in device security chip) to sign result data with ECDSA (Elliptic Curve Digital Signature Algorithm), generating a 64-byte signature result. Combine "user public key + signature result + result data hash" to form transaction data (size ≤ 512 bytes). Send transaction to alliance chain node and call mintAsset method of smart contract (written in Solidity): function mintAsset(string memory resultHash, address user, bytes memory signature) external returns (bool) { / / Verify signature validity require(ecrecover(hash(resultHash), v, r, s) == user, "Invalid signature"); / / Record asset ownership assetOwner[resultHash] = user; return true;} Off-chain storage: the original content of the result data is stored in IPFS (hash: QmXo...yCo), and the IPFS (InterPlanetary File System) hash is written into the blockchain to realize the association of on-chain and off-chain data.
[0062] Display and verification of virtual assets The user calls the smart contract getAsset method on the APP "my collection" page, queries the on-chain record, and obtains the IPFS hash; Download and verify the original data (local calculation hash and on-chain comparison, consistency >=100%); After clicking "display", the AR engine loads the "Hall of Supreme Harmony Cultural Fragment" virtual asset (3D model, polygon number 500), and renders and displays it in the real scene.
[0063] Embodiment two To more clearly reveal the working details of the generation process of the augmented reality content involved in the present application, the second embodiment of the present application provides a structure diagram realized in the augmented reality content generation platform system, see Figure 2 .
[0064] The second embodiment of the present application provides a preferred augmented reality content generation platform system 200, which includes the following modules: a position information acquisition module 201, an augmented reality model database 202, a rendering engine 203, a virtual character model data 204, an augmented reality content large model 205, a user physical motion data acquisition unit 206, a virtual plot data 207, a user individualized data acquisition unit 208, and a blockchain 209. Among them, the augmented reality content generation platform system 200 takes the augmented reality content large model 205 as the core, links the position information acquisition, individualized data acquisition, physical motion capture, model rendering and blockchain notarization modules, and constructs a complete process from scene perception to digital asset right protection, and focuses on strengthening the generation of virtual characters, the generation of virtual plots, the control of rendering and loading, the triggering of plots, the execution result data, and the landing implementation of the generation of blockchain assets.
[0065] The position information acquisition module 201 is configured to collect user travel scene position information, and provide a spatial reference for AR model loading and plot triggering. The position information acquisition module 201 realizes multi-source positioning technology integration, integrates a mobile phone GNSS (Global Navigation Satellite System) module (supports Beidou + GPS (Global Positioning System) dual system, positioning accuracy ≤2m), an IMU inertial sensor (sampling rate 500Hz, obtains acceleration and angular velocity data), and a visual SLAM camera (12 million pixels, frame rate 30fps). In a tourist attraction, GNSS positioning is preferentially called, and latitude and longitude coordinates (format: WGS84, example: north latitude 30.67°, east longitude 104.06°) are output. The IMU data is combined to correct the motion trajectory drift, and the trajectory error is ≤0.5m / 100m. In the indoor / occlusion scene of the scenic spot (for example, a scenic spot museum exhibition hall): start visual SLAM, identify scene feature points (based on the ORB-SLAM3 algorithm, extract 2000+ feature points per second), construct a local map and match it with a pre-stored scene model, and the positioning accuracy is ≤0.3m.
[0066] The augmented reality model database 202 is configured to store multi-precision AR scene models and dynamically output adaptive resources according to position information. The augmented reality model database 202 realizes model hierarchical storage and indexing, can pre-cache augmented reality model data resources of candidate scenic spots on the user's travel route, and can adjust the precision of augmented reality model data resources according to the real-time position information of the user in the travel scene. The specific implementation details are as follows: Storage structure: The directories are divided according to scenic spots (such as “Chengdu Wuhou Temple”), and each scenic spot contains three levels of models: Low-precision model (L0): polygon number ≤5000, texture resolution 256x256, used for fast loading when the user is more than 1km away from the scenic spot, loading time ≤0.5s (under 5G network).
[0067] Medium-precision model (L1): polygon number 10-50 thousand, texture resolution 512x512, covers a scene 500m-1km away from the scenic spot, loading time ≤1.2s.
[0068] High-precision model (L2): polygon number ≥100 thousand, texture resolution 1024x1024, contains artifact-level details (such as 3D replication of Wuhou Temple inscriptions), used when the user is within 500m, loading time ≤2.5s.
[0069] Index mechanism: Establish a "position-precision" mapping table, according to the user coordinates and the scenic spot coordinates output by the position information acquisition module 201, match the model precision through the spatial distance algorithm (time complexity O (1)), for example: when the user is 700m away from the Wuhouci Huiling, call the L1 precision model.
[0070] In the position information acquisition module 201, through the fusion of "GNSS + IMU + SLAM", the AR model precision is dynamically adjusted —— the real-time distance between the user and the scenic spot is accurately calculated (based on the Haversine formula, distance calculation error ≤0.1km), and through the hierarchical model, the balance between AR scene rendering efficiency and effect at different distances is ensured.
[0071] Among them, the rendering engine 203 is configured to realize the real-time rendering or fusion rendering of all virtual characters and AR scenes in the embodiment. Its implementation can be the engine commonly used in the prior art for AR data rendering, which will not be described here.
[0072] In the rendering engine 203, cross-coordinate system rendering mapping is realized, and in the embodiment, there are coordinate systems: User physical space: ENU coordinate system (East-North-Up Coordinate System, East-North-Up Coordinate System), origin is the initial positioning point of the user, used to describe physical motion data.
[0073] AR scene space: Local coordinate system based on scenic geographic information (such as Wuhouci coordinate system, origin is the GPS coordinate of the main door), used for AR model positioning.
[0074] Virtual character space: Skeleton animation coordinate system, with the center of gravity of the character as the origin, describing the action posture.
[0075] The conversion process of these coordinate systems is as follows: The user physical motion data acquisition unit 206 collects hand motion coordinates (ENU system: X=0.5m, Y=-0.2m, Z=1.2m), and through seven-parameter coordinate conversion (including 3 translation, 3 rotation, and 1 scaling parameter, conversion error ≤0.01m), it is mapped to the AR scene coordinate system.
[0076] The rendering engine 203 calls the inverse kinematics (IK) algorithm to convert the hand coordinates to the virtual character skeleton joint angle (such as the wrist joint rotating 30°), and drives the 3D character in the virtual character model data 204 to make the "touch the Three Kingdoms weapon model in the AR scene" action, with a motion synchronization delay ≤30ms.
[0077] Through coordinate system conversion and bone binding, the real-time rendering loading of the virtual role is realized by the rendering engine 203, ensuring the consistency of the virtual role action and the physical motion and the AR scene space.
[0078] Among them, the virtual role model data 204 in the embodiment stores user-specific virtual role resources, which are generated by the augmented reality content large model 205.
[0079] Among them, the augmented reality content large model 205 is configured as the core intelligent module of the augmented reality content generation platform system 200, and fuses multi-modal data to generate virtual roles and virtual plots.
[0080] In this embodiment, the augmented reality content large model 205 contains a hybrid architecture model 2051 for generating virtual plot data, and the training process is as follows.
[0081] Hybrid architecture model training Model structure: First plot text large language model: based on GPT-NeoX-20B architecture fine-tuning, training data contains 100,000 three kingdoms culture scripts, 50,000 travel and culture interaction plots (such as “Wuhouci decryption task”), and perplexity ≤12 after fine-tuning.
[0082] Second scene knowledge graph: taking “three kingdoms culture” as the core, a knowledge graph containing 100,000+ entities (such as “Zhugeliang” and “Muniuliuma”) and 200,000+ relationships (such as “invention - Muniuliuma” and “position - Chancellor”) is constructed and stored in a graph database (such as Neo4j, with query response time ≤50ms). In this embodiment, the second scene knowledge graph storing user travel site historical events and / or cultural common sense structured data can be further trained. For common historical events and / or corresponding cultural common sense of popular tourist attractions, their scene knowledge graphs can be trained. For specific training implementation, refer to the training of the knowledge graph of historical events and / or corresponding cultural common sense of “three kingdoms culture”.
[0083] In this embodiment, preferably, a first plot text large language model related to the user's travel scene is fine-tuned in the hybrid architecture model 2051; a second scene knowledge graph storing user travel site historical events and / or cultural common sense structured data is trained; the first plot text large language model and the second scene knowledge graph are mixed to form the hybrid architecture model 2051, which is part of the augmented reality content large model 205.
[0084] The virtual plot generation process is as follows: Topic retrieval: According to the location information "Wuhouci Huiling" output by the location information acquisition module 201, combined with the user's interest preference ("Three Kingdoms culture"), the "Liu Bei's trust" topic is retrieved from the knowledge graph (retrieval recall rate ≥ 90%).
[0085] Scenario structure design: The first scenario text large language model generates a "trust scene restoration - puzzle interaction - outcome branch" three-act scenario structure based on the topic, including "users need to complete the "read the imperial edict" action through virtual characters (knee joint bending degree ≥ 90°, duration ≥ 2s)" and other trigger conditions.
[0086] Interaction instruction embedding: In the second act of the scenario "puzzle interaction", the first scenario text large language model embeds augmented reality interaction instructions: "users use voice commands (wake-up word 'Liang has a plan') to trigger the display of the wooden ox and flowing horse 3D model", and the voice recognition accuracy is ≥ 92% (in a quiet environment).
[0087] Branch switching: Through the user's heart rate data collected by the user physical motion data acquisition unit 206 (if heart rate > 100 times / minute, "nervous emotion" is determined), the first scenario text large language model triggers the scenario branch: from "regular puzzle solving" to "simplified prompt version", and the branch switching response time is ≤ 1s.
[0088] In this embodiment, preferably, in the hybrid architecture model 2051, a scenario topic content is retrieved from a scenario knowledge graph according to geographical location data of a user in a tourism scenario and cultural preference data of the user; a scenario structure design is generated by a scenario text large language model according to a scenario cultural type preference of the user; augmented reality interaction instructions are embedded in the scenario by the scenario text large language model according to an interaction mode preference of the user; and a scenario branch switching is triggered and generated in the scenario by the scenario text large language model according to real-time emotional parameters of the user.
[0089] It can be seen that in this embodiment, through the "large language model + knowledge graph" hybrid architecture, personalized virtual scenario generation is realized, covering the whole process of topic retrieval, structure design, interaction embedding, and branch switching.
[0090] In the embodiment of the application, the augmented reality content large model 205 further comprises a 3D character model generation type large model 2052 for generating virtual character model data corresponding to the user in augmented reality.
[0091] The training process of the 3D character model generation type large model 2052 is as follows: The model adopts Multi-Modal Generative Adversarial Networks (MM-GAN) as the basic framework, integrates the Generator and the Discriminator, and introduces a conditional constraint module to realize the accurate generation of virtual characters based on multi-dimensional user parameters. Among them: the Generator is responsible for integrating multi-modal input and outputting a virtual character 3D model (including face, body, and action binding information); the Discriminator is used to distinguish between generated models and real character models to improve the generation quality; and the conditional constraint module embeds user personalized parameters (face, body, and interest preferences) to guide the generation direction.
[0092] 3D character model generative large model 2052 training data preparation: (a) Basic data set construction Face feature data set Collect 100,000 sets of facial images of different genders and ages (covering expressions such as smiles, neutral, and serious), with a resolution of 256x256.
[0093] Use the FaceNet model to extract 128-dimensional feature vectors, build a face feature library, calculate the Euclidean distance between features, and select similar feature groups with an error of ≤0.1. Finally, 80,000 valid data sets are retained.
[0094] Data format: {face image path: str, 128-dimensional feature vector: list[float], expression label: str} Body feature data set Collect 50,000 human body parameters (height range 150-200 cm, weight range 40-100 kg), classify by BMI (lean, average, obese), and convert to body mesh parameters (chest circumference, waist circumference, hip circumference, etc.).
[0095] Use 3D modeling software (such as Blender) to generate corresponding body mesh models (polygon count 3000-8000) and export them as glTF format.
[0096] Data format: {height: float, weight: float, body type label: str, body mesh model path: str, mesh parameters: dict} Interest preference-style mapping data set Define 20 categories of travel style labels (Three Kingdoms warriors, Tang Dynasty beauties, science fiction futures, etc.), and for each category, collect 3000 sets of 3D character models (including action binding) corresponding to the style to build a style library.
[0097] Establish the mapping relationship between interest preference tags (such as "Three Kingdoms culture" and "warlord style") and the style library, and form a preference-style association table.
[0098] Data format: {interest preference tag: list[str], style tag: str, 3D character model path: str, action binding data: dict} (b) Multi-modal data fusion Fuse the above multi-source data through attention mechanism (Attention Mechanism): The face feature vector, body mesh parameter, and interest preference tag are respectively mapped to a 512-dimensional space through linear transformation (Linear Layer).
[0099] Use the Multi-Head Attention module to calculate the correlation weight between different modal data (such as the correlation degree of facial features and "Three Kingdoms Warlord" style), and generate a fusion feature vector (dimension 512).
[0100] The data format after fusion: {fusion feature vector: list[float], original multi-modal data index: dict} The training process of this model is as follows: (1) Training stage division It is divided into two stages of pre-training (Pre-training) and fine-tuning (Fine-tuning), and gradually improves the model's fitting ability to user's personalized parameters.
[0101] (2) Pre-training: construction of general character generation ability Input: Randomly select face, body, and style data from the basic data set, and fuse them as input (100,000 training samples).
[0102] Generator training: Objective: Generate a 3D model similar to the real character model, and minimize the perceptual loss (Perceptual Loss) between the generated model and the real model.
[0103] Loss function: Where, is the perceptual loss (calculated by the pre-trained VGG model), is the generative adversarial loss (the probability loss output by the discriminator), = 0.7, = 0.3.
[0104] Training steps: 200,000 steps, learning rate 1e-4, optimizer Adam = 0.5, = 0.999 ).
[0105] Where the discriminator is trained as follows: Objective: To distinguish between the virtual characters generated by the generator and the real character models, maximize the difference between the true and false classification probabilities.
[0106] Loss function: Where, is the discriminant probability of the real model, is the discriminant probability of the generated model.
[0107] Training synchronization: Train the generator for 1 step and the discriminator for 2 steps to ensure stable convergence of the model.
[0108] (3) Fine-tuning: Strengthen the ability of personalized character generation Input: Focus on the "Three Kingdoms General" style, filter 30,000 relevant training samples (including user personalized parameters: facial features, body parameters, interest preferences).
[0109] Condition constraint reinforcement: Add a conditional embedding module (Conditional Embedding) to the input layer of the generator, encode user personalized parameters (such as "height 175cm, well-proportioned, Three Kingdoms General style") into a conditional vector (dimension 256), and input the generator after concatenating with the fused feature vector.
[0110] Introduce style loss (Style Loss), calculate the texture and pose difference between the generated character and the target style (Three Kingdoms General), and the loss function is: Where, G(x) is the generated character, is the target style character, = 0.8, SSIM is the structural similarity index, and L1 is the pixel error.
[0111] Training parameters: Training steps: 100,000 steps, learning rate 5e-5 (using cosine annealing scheduling).
[0112] Optimization goal: Minimize the total loss = + , so that the matching degree of the generated character and the user personalization parameter is ≥ 90% (verified by manual annotation).
[0113] In this embodiment, the inference process of the augmented reality 3D character model generation model 2052 in the augmented reality content large model 205 is as follows: The user personalization data acquisition unit 208 collects multi-dimensional parameters, and generates a 512-dimensional feature vector + 256-dimensional condition vector after fusion.
[0114] After the augmented reality content large model 205 is loaded, the fusion features and condition vectors are input to generate a virtual character 3D model (glTF format, including face mesh, body mesh, bone binding information, etc.).
[0115] The model data is output to the virtual character model data module 204 for rendering engine 203 to call rendering processing.
[0116] The following is a detailed example of the augmented reality 3D character model generation model 2052 generating character model data.
[0117] Step A: User Personalization Data Input and Analysis Input source: The user personalization data acquisition unit 208 collects three types of core parameters, which are analyzed as follows: Face feature data: Original input: 3 sets of different expression photos (256x256 pixels, RGB format), 128-dimensional feature vectors are extracted by FaceNet model (such as smile expression vector: [0.12, 0.34,..., 0.78]), and the Euclidean distance error is ≤0.1.
[0118] Analysis processing: The 128-dimensional vector is compressed to 64-dimensional (retaining 95% feature information) by PCA (Principal Component Analysis) dimension reduction algorithm as the basis for face generation.
[0119] Body feature data: Original input: User input height (175cm), weight (65kg), body type label ("symmetrical"), which is automatically converted to body mesh parameters by the system: {"height": 175.0, "bust": 90.0, / / Chest circumference (cm) "waist": 75.0, / / Waist circumference (cm) "hip": 92.0, / / Hip circumference (cm) "limb_ratio": 0.85 / / Four limbs and torso ratio Parsing process: Normalize the parameters to the [0,1] interval (such as height 175cm corresponding to the normalized value 0.65, based on the value range of 150-200cm), generate a 32-dimensional body feature vector.
[0120] Interest preference data: Original input: For example, the user checks the "Three Kingdoms culture" and "Warrior style" tags in the APP program, and the system associates the mapping relationship in the style database (such as "Three Kingdoms warrior" corresponding to the clothing texture library ID: texture_hanfu_003, action library ID: action_warrior_012).
[0121] Parsing process: Convert the tags into a 64-dimensional one-hot encoding vector (such as "Three Kingdoms culture" corresponding to the 12th bit being 1, and the rest being 0), as a style constraint condition.
[0122] Step B: Multimodal feature fusion and conditional encoding Fusion target: Integrate the feature vectors of face, body, and interest preferences into a unified generation condition to guide the model to generate a character that meets the user's needs.
[0123] The technical implementation details are as follows: Feature splicing: Splice the 64-dimensional face vector, 32-dimensional body vector, and 64-dimensional preference vector into a 160-dimensional basic feature vector.
[0124] Attention weighting: Calculate the feature weight through the trained attention module (such as the "warrior style" preference weight for clothing generation is increased to 0.8), the formula is as follows: Attention( ) = Where, is the i-th feature, is the weight parameter learned by training.
[0125] Conditional encoding: Input the weighted feature vector into the conditional embedding layer (composed of 2 layers of fully connected networks), output a 256-dimensional conditional encoding vector as the input of the generator.
[0126] Step C: Generator outputs virtual character model data in stages (Based on the trained MM-GAN generator, generate face, body, and stylized components in three stages) Stage 1: 3D face mesh generation Input: 64-dimensional face feature vector + 256-dimensional conditional encoding vector.
[0127] Generation process: The face branch of the generator adopts a U-Net structure, first generating a 128×128×128 3D voxel grid representing the three-dimensional structure of the face through a deconvolution layer.
[0128] Apply the Poisson fusion algorithm to optimize the facial contour, so that the generated mesh matches the user's face feature vector with a degree of ≥95% (calculated by Chamfer distance: d ≤ 0.02).
[0129] Output: 3D face mesh containing 8000 vertices and 16000 triangular facets (format:.obj), with accompanying expression binding information (supporting 5 basic expressions such as smiling and serious).
[0130] Phase 2: 3D body mesh generation Input: 32-dimensional body feature vector + 256-dimensional conditional encoding vector.
[0131] Generation process: The body branch of the generator adopts a Transformer architecture, predicting bone lengths (such as femur length 45cm, tibia length 40cm) and muscle distribution based on body parameters.
[0132] Generate a body mesh with 3000-5000 vertices, ensuring a height error of ≤2cm and a bust error of ≤1cm.
[0133] Output: Body mesh compatible with face mesh topology (format:.obj), containing 18 major skeletal joints (such as hip joint, shoulder joint).
[0134] Phase 3: Style component generation (clothing, texture, action) Input: 64-dimensional interest preference vector + face / body mesh data.
[0135] Generation process: Clothing generation: Call the style sub-network according to the "Three Kingdoms General" label to generate clothing meshes such as armor and battle robes (polygon count ≤2000), and automatically match the size of the body mesh (such as shoulder width adaptation to 90cm bust).
[0136] Texture mapping: Retrieve the "Xuanjia" texture (resolution 512×512) from the texture library and attach it to the clothing mesh through the UV unwrapping algorithm, ensuring a texture stretching error of ≤5%.
[0137] Action Binding: Select 10 core actions (e.g., sword swing, bow) from the "Warrior Action Library" and bind the action data to the skeleton through skinning weight calculation (weight error ≤0.05) to ensure smoothness ≥60fps.
[0138] Preferably, in the present embodiment, step D: model post-processing and precision optimization can be performed Objective: Improve the rendering efficiency and adaptability of the model in the AR scene to meet the "real-time rendering" requirement.
[0139] Processing measures: Mesh Simplification: Use the Quadric Error Metric algorithm to simplify the generated mesh, retaining 90% of the visual effect while reducing the number of polygons by 40% (e.g., facial mesh simplified from 8000 vertices to 4800 vertices).
[0140] Topology Check: Automatically detect and repair non-manifold edges, duplicate vertices, and other issues to ensure the model is free of abnormalities in the rendering engine.
[0141] LOD Generation: Generate 3 levels of detail models (high / medium / low models) corresponding to distances ≤5m, 5-10m, and >10m from the user, respectively. The high model retains all details, and the low model has a polygon count ≤1000.
[0142] Format Conversion: Convert the.obj format to the AR engine compatible glTF 2.0 format, compress the model volume (from 50MB to within 10MB), and load time ≤500ms.
[0143] Step E: Model Data Output and Storage Output Content: Packaged virtual character model data package, including: { "model_id": "char_12345", / / Unique identifier "face_mesh": "face_lod0.gltf", / / Facial mesh (high model) "body_mesh": "body_lod0.gltf", / / Body mesh (high model) "clothes_mesh": "armor_lod0.gltf", / / Clothing mesh "textures": ["armor_tex.jpg"], / / Texture files "animations": ["sword_attack.glb"], / / Action data "skeleton": "skeleton.json", / / Skeleton binding information "LOD_config": {"lod0": 5, "lod1": 10, "lod2": 20} / / LOD distance configuration (meters) Storage location: Output to the distributed database of the virtual character model data module 204 (response time ≤ 50 ms) and associated with the user ID, supporting subsequent calls and updates.
[0144] In this embodiment, for example, user "Zhang San" (facial feature vector V1, height 175 cm, preference "Three Kingdoms generals") is taken as an example, and the generation result is as follows: Face matching degree: The SSIM similarity between the generated virtual character face and the user photo is ≥0.92, and the user facial features (such as single eyelid and high nose bridge) can be clearly identified. Body adaptability: The virtual character height is 174.8 cm, and the chest circumference is 89.5 cm, with an error of ≤0.5%. Style fit degree: The clothing is of "Three Kingdoms Xuanjia" style, and the action library contains "riding a horse with a gun", "kneeling on one knee", and other exclusive actions of generals.
[0145] Among them, the user physical motion data acquisition unit 206 in this embodiment is configured to collect user body motion data, drive the virtual character, and trigger the plot.
[0146] In this embodiment, the user's multi-dimensional motion data is collected according to the following scheme: Hardware integration: 9-axis IMU (accelerometer, gyroscope, magnetometer) built-in mobile phone, pressure-sensitive screen (pressure resolution 0.1N), binocular camera (for action recognition).
[0147] Data type and collection parameters: Body posture data: Collect Euler angles (pitch angle, yaw angle, roll angle, accuracy ≤1°) through IMU, sample 50 times per second, for example: when the user makes a "bow" action, the pitch angle changes from 0° to -30°.
[0148] Action trajectory data: Binocular camera calculates hand motion trajectory based on optical flow method (Optical Flow), trajectory accuracy ≤0.05m, for example: the user's hand waving trajectory is a parabola, with vertex coordinates (X=0.8m, Y=1.5m, Z=0.0m).
[0149] Interaction intensity data: Pressure-sensitive screen collects touch pressure (range 0-5N), sampling rate 100Hz, for example: when the user clicks the "confirm the puzzle" button, the pressure value is 2.3N, and the duration is 0.5s.
[0150] The user physical motion data acquisition unit 206 in this embodiment provides input data for the plot trigger condition template, for example, by body posture data to determine whether the "user action" (knee angle ≤ 90°, elbow angle ≤ 40°) meets the trigger condition of the plot.
[0151] Preferably, in this embodiment, the virtual plot data generated according to the hybrid architecture model 2051 sets the plot trigger condition template (not shown), wherein the plot trigger condition includes: virtual character action type, action parameter threshold and tourism scene context. The user physical motion data acquisition unit 206 provides input data for the plot trigger condition template (not shown) in the augmented reality content large model 205. In the virtual plot data 207 in this embodiment, the user real-time motion data is compared with the template, and when the trigger condition is met, it is determined that the virtual plot trigger is executed.
[0152] In this embodiment, the virtual plot data 207 is configured to store the plot resources generated by the augmented reality content large model 205, including trigger conditions, processes, branches, etc.
[0153] In this embodiment, the plot data is preferably structured for storage, and most preferably includes virtual plot data execution result data. The following is an example of a plot data format: Data format: stored in JSON-LD format, example: { "plot_id": "WH_CI_001", "theme": "Liu Bei entrusts", "trigger_conditions": [ { "type": "pose", "parameter": "knee_angle≤90°", "scene_context": "user is within 5m of Huiling" } ], "execution_flow": ["scene_reconstruction", "puzzle_interaction", "branch_switch"], "branches": { "normal": "puzzle_hard.json", "easy": "puzzle_easy.json"}, "result_data_schema": { / / Data structure for the execution result of virtual storyline data "character_model": "wujiang_001.glb", "context_data": "The scene of Huiling entrusting her orphan to her father evokes tension in users", "scene_data": "Wuhou Temple Scenic Area, GPS(30.67,104.06)", "preference_data": "Three Kingdoms culture, military style" }} In this embodiment, a schema for generating the virtual storyline data execution result is defined. Referring to the `result_data_schema` section, it includes data on dimensions such as the virtual character model, scene context, travel scenario, and interest preferences, providing a foundation for subsequent blockchain-based evidence storage. In this embodiment, the augmented reality content model generates the virtual storyline data execution result based on at least one of the following: the user's virtual character model data, the context data when the user triggers the virtual storyline, the user's travel scenario data, and the user's interest and behavioral preference data.
[0154] In this embodiment, the user personalized data acquisition unit 208 is configured to collect multi-dimensional user personalized parameters to provide input for the generation of large models.
[0155] In this embodiment, the user personalized data acquisition unit 208 performs multimodal data acquisition and fusion from the following data sources.
[0156] Data source -- User identity data: Basic information such as name, age, and gender are obtained by associating with third-party accounts (such as WeChat and Alipay) through the OAuth 2.0 protocol (stored after anonymization, with field length ≤ 50 bytes).
[0157] Data source -- Appearance feature parameters: Combined with mobile phone camera (facial recognition) and album upload (clothing style recognition, using VGG16 model, recognition accuracy ≥85%), a 256-dimensional feature vector is generated.
[0158] Data Sources -- Interests and Behavioral Preferences: Analysis of in-app browsing history (e.g., reading time of articles on "Three Kingdoms Culture" ≥ 5 minutes) and search keywords ("Zhuge Liang's Inventions"), and generation of preference tags through a collaborative filtering algorithm (similarity calculation uses cosine similarity, Top-N recommendation N=5).
[0159] Data source - scene and state parameters: Collect ambient light intensity (range 0 - 10000 lux), temperature (range -20 - 50°C), user steps (daily steps ≥ 5000 steps to determine "active") through mobile phone sensors, and generate scene state vectors.
[0160] In this embodiment, the user personalized data acquisition unit 208 can be further configured to perform a fusion algorithm on the collected multi-modal data: use an attention mechanism to weight and fuse multi-modal data, generate a 512-dimensional user personalized feature vector (weight distribution: identity 0.1, appearance 0.3, preference 0.4, scene 0.2), and provide input for the augmented reality content large model 205.
[0161] Among them, the blockchain 209 in this embodiment is configured to realize the notarization of virtual plot execution result data and asset right confirmation.
[0162] In this embodiment, the blockchain architecture of the blockchain 209 is a consortium chain (such as Hyperledger Fabric), the nodes include scenic area servers, travel platforms, and third-party notarization agencies, and the consensus mechanism is PBFT (delay ≤ 2s, throughput ≥ 100TPS).
[0163] On-chain data preparation: Extract virtual plot data execution result data (such as the above result_data_schema), and calculate SHA-256 hash (hash value length 64 bytes).
[0164] The user client calls the key management module, uses the user private key (stored in the mobile phone security chip, ECC algorithm, key length 256 bits) to sign the hash value, and generates a 64-byte signature result.
[0165] On-chain operation: Combine "user public key (corresponding to private key, length 64 bytes), signature result, and execution result data hash" to build a blockchain transaction (transaction size ≤ 2KB).
[0166] Send the transaction to the consortium chain node and call the mintVirtualAsset function of the smart contract (written in Solidity, version 0.8.0+): function mintVirtualAsset( address userAddr, bytes32 resultHash, bytes memory signature public returns (bool) { require(verifySignature(userAddr, resultHash, signature), "Invalidsignature"); virtualAssets[resultHash] = userAddr; / / Asset ownership confirmation emit AssetMinted(userAddr, resultHash); return true;} The smart contract verifies the signature validity (using the ecrecover function, with a success rate of ≥99.9%). After confirmation, the asset hash is bound to the user address to complete the on-chaining of the virtual asset, ensuring the uniqueness and traceability of the virtual asset.
[0167] The method flow implementation solutions in the above embodiments can be implemented as software programs or encapsulated as software modules. Those skilled in the art can understand and implement them as various software programs, mobile apps, API interfaces, or SaaS (Software as a Service) services without any creative effort. The method flow implementation solutions embodied in the software programs or encapsulated as software modules corresponding to the method flow implementation solutions in the embodiments of this invention also fall within the spirit and scope of this invention.
[0168] The module-type embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating augmented reality content, comprising: obtaining first augmented reality model data of a travel scene according to user position information in the travel scene; obtaining user personalized input parameters and inputting them into an augmented reality content large model to generate second virtual character model data and virtual plot data corresponding to the user; obtaining user physical motion data to control real-time rendering and loading of the second virtual character model in the first augmented reality model data and to trigger execution of the virtual plot data; generating and saving virtual asset data corresponding to the virtual plot data execution result data in a blockchain according to the virtual plot data execution result data.
2. The method of claim 1, wherein, The first augmented reality model data of the travel scene is obtained according to the user position information in the travel scene, comprising: pre-caching augmented reality model data resources of candidate scenic spots on the user's travel route; obtaining real-time position information of the user in the travel scene; calculating real-time distance data between the user's real-time position and the scenic spots to dynamically adjust the first model precision of the first augmented reality model data of the travel scene; obtaining augmented reality model data corresponding to the first model precision from the augmented reality model data resources of the candidate scenic spots.
3. The method of claim 1, wherein, The user's personalized input parameters include: at least one of user identity data, user appearance feature parameters, user interest and behavior preference parameters, and user scene and state parameters.
4. The method of claim 3, wherein, The augmented reality content large model generates second virtual character model data corresponding to the user, comprising: obtaining a fusion feature vector formed by a role face feature dataset, a body feature dataset, and an interest preference-style mapping dataset; training an augmented reality 3D character model generative large model based on the fusion feature vector and user personalized image constraint conditions; The augmented reality 3D character model generative large model generates a 3D face mesh based on the input user face feature vector and matches the user's face contour through a Poisson fusion algorithm; The augmented reality 3D character model generative large model generates a body mesh based on the user's appearance feature parameters; The augmented reality 3D character model generative large model generates costumes, textures, and corresponding character actions based on the user's interest preferences and style parameters, and realizes real-time character action driving of the 3D character model through skeletal binding.
5. The method of claim 3, wherein, The augmented reality content large model generates virtual plot data corresponding to the user, comprising: fine-tuning a first plot text large language model related to the user's travel scene; training a second scene knowledge graph storing user travel scenic spot historical events and / or cultural common sense structured data; mixing the first plot text large language model and the second scene knowledge graph to form the augmented reality content large model; retrieving first plot theme content from the second scene knowledge graph according to geographic location data of the travel scene and cultural preference data of the user; generating a second plot structure design through the first plot text large language model according to the user's plot cultural type preference; embedding third augmented reality interaction instructions in the plot through the first plot text large language model according to the user's interaction mode preference; According to the real-time emotional parameters of the user, the fourth plot branch switching is triggered in the plot by the first plot text large language model.
6. The method of claim 1, wherein, The physical motion data of the user includes: At least one of the user's body posture data, motion trajectory data, and interaction force data.
7. The method of claim 1, wherein, The real-time control of the second virtual character model in the first augmented reality model data includes: Mapping the body motion data in the user's physical motion data to the skeletal coordinate system of the second virtual character; And convert the three-dimensional coordinate system of the user's physical motion data to the spatial coordinate system of the first augmented reality model data; Bind the user's physical motion data to the second virtual character's skeleton, and automatically calculate the associated joint posture data for supplementation; Real-time rendering of the virtual character and updating the relative position of the virtual character and the spatial scene of the first augmented reality model data.
8. The method of claim 1, wherein, The implementation of the virtual plot data trigger execution includes: According to the virtual plot data, set the plot trigger condition template, wherein the plot trigger condition includes: virtual character action type, action parameter threshold and tourism scene context; Compare the user's real-time motion data with the template, and determine the virtual plot trigger execution when the trigger condition is met.
9. The method of claim 1, wherein, The virtual plot data execution result data includes: According to at least one of the user's virtual character model data, the context data when the user triggers the virtual plot, the tourism scene data where the user is located, and the user's user interest and behavior preference data, the virtual plot data execution result data generated by the augmented reality content large model.
10. The method of claim 1, wherein, According to the virtual plot data execution result data, generate and save the virtual asset data corresponding to the virtual plot data execution result data in the blockchain includes: Using the user's private key data, the virtual plot data execution result data is signed and calculated, and the user's public key, the signature calculation result and the virtual plot data execution result data are combined to save the chain, forming the user's virtual asset data; wherein the virtual plot data execution result data includes augmented reality virtual assets loaded and displayed in the augmented reality tourism scene.
Citation Information
Patent Citations
Digitized interactive intelligent travel system
CN115375878A
Cited By
Customization method of personalized script task for real-time dynamic tourism scene
CN121722980A
A method for customizing personalized script tasks for real-time dynamic tourism scenarios
CN121722980B