Augmented reality and virtuality fused cartoon scene real-time construction interaction system and method

The real-time animation scene construction system that integrates augmented reality and virtual reality solves the limitations of virtual scene construction and interaction in the animation field. It enables the efficient construction of complex scenes and a multi-user shared experience that blends virtual reality with the real environment, enhancing users' immersion and interactive enjoyment.

CN121121007APending Publication Date: 2025-12-12ANHUI SHENGSHI JIAYE CULTURAL IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511280180.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing AR and VR technologies in the animation field suffer from problems such as virtual scene construction being limited by physical space and terminal computing power, inconsistencies between virtual elements and real-world lighting, unintuitive interaction methods, and difficulties in multi-user interaction, making it difficult to achieve a high-quality, multi-user shared experience that blends virtual reality with the real environment.

Method used

The system employs a real-time animation scene construction system that integrates augmented reality and virtual reality. It includes a scene construction and management module, a virtual and real rendering engine module, a multimodal interaction processing module, user terminal devices, and a cloud collaboration unit. Through SLAM, ambient lighting estimation, adaptive rendering, multimodal interaction processing, and cloud collaboration, it achieves seamless integration of virtual elements with the real environment and multi-user interaction.

Benefits of technology

It enables efficient construction and intuitive interaction of complex scenes, enhances the realism and immersion of virtual elements, lowers the learning threshold, supports multiple users to share the same high-quality scene, and provides a natural and intuitive interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121121007A_ABST
    Figure CN121121007A_ABST
Patent Text Reader

Abstract

The invention discloses an augmented reality and virtuality fused cartoon scene real-time construction interaction system and method. The system comprises a scene construction and management module, a virtuality and reality rendering engine module, a multi-mode interaction processing module, user terminal equipment and a cloud collaboration unit. The scene construction and management module is used for receiving, processing and storing scene data from a virtual and real environment and generating a unified fusion scene coordinate system in which virtual and real elements coexist; the virtual-real rendering and engine module is connected with the scene construction and management module and is used for carrying out integrated three-dimensional real-time rendering with consistent light and shadow on the virtual animation elements and the real environment according to the visual angle and the position of the user; the method has the beneficial effects that the reality sense and immersion sense of virtual elements are greatly improved based on a real-time rendering technology of real environment illumination, and through an interaction mode of combining gestures and entity props, the method is more natural and visual, and the interaction fun is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, in particular to an augmented reality and virtual reality fusion animation scene real-time construction interaction system and method. BACKGROUND

[0002] With the rapid development of computer graphics, sensor technology and display technology, augmented reality (AR) and virtual reality (VR) technologies are gradually moving from concepts to large-scale applications, especially in entertainment, education, tourism and other consumer fields, showing great potential. AR technology aims to superimpose computer-generated virtual information onto the real world where the user is located, achieving "augmentation" of the real environment; while VR technology creates a completely immersive virtual environment, allowing users to temporarily isolate themselves from the real world and enter a digitally simulated space.

[0003] In the animation industry, these two technologies have brought revolutionary opportunities for the presentation of content and the viewing experience of users. Traditional animation consumption relies on flat media (such as comic books, animated films) or offline physical exhibitions, with limited interactivity, and the audience is always in the position of an "onlooker". AR technology attempts to project animation characters into the real world, allowing users to view and interact with them through mobile screens or AR glasses, such as triggering pre-set animations by recognizing images. VR technology allows users to enter a completely fictional animation world in first-person perspective for exploration.

[0004] Existing technical solutions have obvious limitations. First, AR experiences are limited by physical space and terminal computing power, making it difficult to create large and complex virtual scenes. VR experiences are completely detached from the real environment, lacking the ability to interact with the real world, and the creation process is often carried out on traditional two-dimensional screens, which is not intuitive. Second, the rendering and lighting of virtual objects do not match the lighting conditions of the real environment, resulting in missing or incorrect shadows, which prevents virtual elements from truly "integrating" into the real world, disrupting the user's sense of immersion. Third, most interactions rely on traditional peripherals such as hand controllers, rather than gestures, voice or everyday objects that are more in line with human instincts, making the interaction process less intuitive and efficient. Finally, most experiences are isolated and one-time, making it difficult for multiple users to share the same fusion scene and interact with different perspectives (AR or VR) in the same physical space, and it is also difficult to anchor the carefully constructed virtual scene in a specific location for different batches of users to access. To address these issues, we propose an augmented reality and virtual reality fusion animation scene real-time construction interaction system and method. SUMMARY

[0005] The application aims to provide an augmented reality and virtual reality combined animation scene real-time construction interaction system and method to solve the problems in the background art.

[0006] To achieve the above-mentioned purpose, the augmented reality and virtual reality combined animation scene real-time construction interaction system comprises a scene construction and management module, a virtual reality rendering engine module, a multi-modal interaction processing module, a user terminal device and a cloud collaborative unit.

[0007] The scene construction and management module is used to receive, process and store scene data from a virtual reality environment and generate a unified, virtual reality element coexisting combined scene coordinate system.

[0008] The virtual reality rendering and engine module is connected to the scene construction and management module and is used to perform integrated, light and shadow consistent three-dimensional real-time rendering of virtual animation elements and real environment according to the user's perspective and position.

[0009] The multi-modal interaction processing module is used to capture and identify the interaction instructions issued by the user through gestures, voice, eye movement or physical props and analyze the instructions into operation commands for virtual elements or virtual reality combined elements in the combined scene.

[0010] The user terminal device comprises an AR display device or a VR display device for presenting the combined scene and a sensor group for capturing user interaction instructions.

[0011] The cloud collaborative unit is used to distribute rendering tasks with high computing load and complex interaction logic and provide support for sharing and persisting the unified combined scene for multiple users.

[0012] Preferably, the scene construction and management module comprises a SLAM unit and a scene anchor point management unit.

[0013] The SLAM unit is used to perform real-time map construction, positioning and tracking of the real environment.

[0014] The scene anchor point management unit is used to define and manage anchor points of virtual elements in real space, and the anchor points comprise image feature points, planes, objects or custom space coordinates.

[0015] Preferably, the virtual reality rendering engine module comprises an environment light estimation unit and an adaptive rendering unit.

[0016] The environment light estimation unit is used to analyze the light intensity, direction and color temperature of the real environment in real time.

[0017] The adaptive rendering unit is configured to dynamically adjust the shading, shadow and highlight of the virtual element according to the analysis result of the ambient light estimation unit, so as to match the lighting condition of the real environment.

[0018] Preferably, the multi-modal interaction processing module supports the recognition and tracking of physical props, which are physical objects with specific visual markers or shapes. The system can bind virtual special effects or models with the physical props held by the user in space, so that the physical props become the interactive tools or weapons of the user in the fusion scene.

[0019] Preferably, the system supports a "VR construction-AR experience" mode: the creator first constructs and lays out the animation scene in a completely virtual VR environment; after completion, the system maps the VR scene to the specified real physical space; when the experimenter enters the physical space through the AR device, the complete scene created in the VR can be seen and interacted with.

[0020] Preferably, the user terminal device is an MR head-mounted display, and the MR head-mounted display can seamlessly switch between the AR perspective mode and the VR immersion mode; when switching to the VR mode, the system uses the three-dimensional model reconstructed from the real environment as the basis, and virtually presents the real scene elements together with the virtual animation elements.

[0021] Preferably, the cloud collaboration unit is built-in with an AI driving unit, which is configured to analyze the behavior patterns of the user and pre-load the virtual assets that the user may need.

[0022] The real-time construction interaction method of the animation scene of augmented reality and virtual fusion is applied to any of the real-time construction interaction systems of the animation scene of augmented reality and virtual fusion described above, and includes the following steps:

[0023] S1, acquiring spatial data of a real environment and real-time pose data of a user terminal device through a sensor, constructing and maintaining a fusion scene coordinate system;

[0024] S2, importing or creating a virtual animation character and scene element in real time, and anchoring and registering the virtual animation character and scene element with the spatial features of the real environment, so as to ensure the stability of the position and pose of the virtual element in the physical space;

[0025] S3, performing real-time light and shadow calculation and rendering on the virtual animation element based on the user's perspective and the lighting condition of the physical environment, so as to seamlessly integrate the virtual animation element with the real environment visually, and generate a virtual-real fusion animation scene;

[0026] S4, continuously capturing the interactive behavior of the user through multi-modal sensors, and identifying the interactive intention of the user;

[0027] S5, according to the identified interaction intention, driving the virtual element in the fusion scene to make corresponding behavior feedback or change the scene state, realizing the natural interaction of the user and the animation scene.

[0028] Compared with the prior art, the beneficial effects of the present application are:

[0029] 1, the present application uses the immersion of VR to complete the efficient and intuitive construction of complex scenes, and the user obtains a deep interactive experience with the virtual-real fusion scene through the AR / VR device, realizing the innovative work flow of "VR construction-AR experience".

[0030] 2, the present application is based on the real-time rendering technology of real environment light, which greatly improves the reality and immersion of virtual elements, and the interactive mode combining gestures and physical props is more natural and intuitive, reduces the learning threshold, enhances the interactive fun, and ensures the smoothness and response speed when multiple users share the same high-quality scene through cloud collaboration and AI preloading. BRIEF DESCRIPTION OF DRAWINGS

[0031] Fig. 1 is the schematic diagram of the present application;

[0032] Fig. 2 is the method flow chart of the present application. DETAILED DESCRIPTION

[0033] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0034] Please refer to Figs. 1-2 The present application provides a technical solution: an animation scene real-time construction interaction system based on augmented reality and virtual reality fusion, including a scene construction and management module, a virtual-real rendering engine module, a multi-modal interaction processing module, a user terminal device and a cloud collaboration unit.

[0035] The scene construction and management module is used for receiving, processing and storing scene data from the virtual-real environment, and generating a unified fusion scene coordinate system in which virtual and real elements coexist. The scene construction and management module includes a SLAM unit and a scene anchor point management unit.

[0036] The SLAM unit is used for real-time map construction, positioning and tracking of the real environment; the scene anchor point management unit is used for defining and managing the anchor points of virtual elements in the real space, including image feature points, planes, objects or custom spatial coordinates.

[0037] It should be noted that the scene construction and management module is specifically started: the creator wears a VR headset and enters a completely virtual "fantasy forest" scene construction space; the SLAM unit does not directly map the real exhibition hall in this stage, but its core algorithm provides a basis for spatial positioning in the VR environment and tracking of virtual cameras; the creator uses a VR handle to drag pre-made 3D animated characters (such as a tree spirit, a dragon model), trees, rocks, glowing particles, and other virtual elements from the asset library, and freely layouts them in the virtual space to build a complete scene;

[0038] After the construction is completed, the creator specifies the origin and boundary of the virtual scene through the interface; then the system prompts the creator to walk around in the real exhibition hall physical space, and the sensors of the MR headset (the SLAM unit starts working) perform three-dimensional scanning and map construction of the exhibition hall; the creator aligns the "origin" of the virtual scene with a specific corner of the real exhibition hall floor (this can be achieved by clicking on the floor or recognizing a pre-placed marker). The scene anchor management unit records this correspondence and completes the mapping and registration of the virtual scene to the physical space. This process permanently anchors the entire VR scene in the real exhibition hall.

[0039] The virtual-real rendering and engine module is connected with the scene construction and management module, and is used for integrated, light-consistent three-dimensional real-time rendering of virtual animated elements and real environment according to the user's perspective and position. The virtual-real rendering engine module includes an ambient light estimation unit and an adaptive rendering unit;

[0040] The ambient light estimation unit is used for real-time analysis of the light intensity, direction and color temperature of the real environment; the adaptive rendering unit is used for dynamically adjusting the coloring, shading and highlights of the virtual elements according to the analysis results of the ambient light estimation unit, so that they match the lighting conditions of the real environment.

[0041] It should be noted that when the user experiences in the AR experience mode, the audience A wears an MR headset and enters the "fantasy forest" exhibition area. The sensor group (camera, depth sensor) on the headset immediately starts; the SLAM unit starts working in real time, quickly locates the precise position and pose (6DoF) of the headset in the previously built exhibition hall map through feature recognition and matching of the surrounding environment; the terminal device sends a request to the cloud-side collaborative unit. The cloud side streams the "fantasy forest" virtual scene data previously constructed and anchored by the creator to the user's headset device according to the user's position information;

[0042] The virtual-real rendering engine module starts to work, and the ambient light estimation unit analyzes the light color, brightness and main light source direction in the exhibition hall in real time through the RGB camera of the head-mounted display; the adaptive rendering unit receives the light data and dynamically adjusts the rendering parameters of the virtual elements accordingly. For example, the skin color of the tree spirit (shading) will become soft according to the warm-toned light in the exhibition hall, and the direction and softness of the shadow cast by the dragon will be consistent with the shadow under the real light, so that the virtual character looks like it is really standing on the floor of the exhibition hall, and visually seamlessly integrates.

[0043] The multi-modal interaction processing module is used to capture and recognize the interaction instructions issued by the user through gestures, voice, eye movement or physical props, and to parse the instructions into operation commands for virtual elements or virtual-real combined elements in the fusion scene. The multi-modal interaction processing module supports the recognition and tracking of physical props, which are physical objects with specific visual markers or shapes. The system can bind virtual special effects or models with physical props held by the user in space, so that the physical props become the user's interaction tools or weapons in the fusion scene.

[0044] It should be noted that in the multi-modal interaction, when audience A sees the tree spirit character waving his hand in front of him, he stretches out his hand and wants to shake hands with the character. The hand tracking camera of the head-mounted display captures his hand skeleton movement data, and the multi-modal interaction processing module recognizes that it is an "handshake" intention instruction. The instruction is parsed into an operation command to drive the tree spirit character program to make a "handshake response" animation feedback, and may trigger a sound effect. Audience B picks up a "magic wand" physical prop next to him. The camera of the head-mounted display recognizes the specific visual marker on the prop; the multi-modal interaction processing module continuously tracks the position and pose of the prop in space, and accurately binds a "fire magic" particle special effect virtual model to the top of the magic wand; when audience B waves the magic wand, he sees a virtual fire trail following the prop, as if he is really casting magic. He can point the magic wand at the virtual dragon in the distance to trigger the dragon's roar reaction.

[0045] The user terminal device includes an AR display device or a VR display device for presenting the fusion scene, and a sensor group for capturing user interaction instructions. The user terminal device is an MR head-mounted display, and the MR head-mounted display can seamlessly switch between an AR see-through mode and a VR immersion mode; when switching to the VR mode, the system uses a three-dimensional model reconstructed from the real environment as the basis to virtually present the real scene elements together with the virtual animation elements.

[0046] It should be noted that in the system hardware configuration, the user terminal device provides a mixed reality (MR) head-mounted device for the audience, which integrates an RGB camera, a depth sensor, an IMU, and a microphone array for environment perception, gesture tracking, and voice capture. At the same time, the device supports color see-through function and can switch between AR and VR modes; the physical prop is a specially designed "magic wand" prop with a special visual marker printed on the handle; the cloud server is a cloud server with a high-performance GPU cluster for running the cloud collaboration unit and the AI-driven unit; the creator workstation is equipped with a high-performance PC and a VR headset for content creators to build scenes in VR mode.

[0047] The cloud collaboration unit is used to distribute the rendering tasks with high computing load and complex interaction logic, and to support the sharing and persistence of the unified fusion scene for multiple users. The cloud collaboration unit has an AI-driven unit built in, which is used to analyze the user's behavior patterns and pre-load virtual assets that the user may need.

[0048] It should be noted that throughout the process, high-precision character model rendering and complex particle effect calculation are distributed by the cloud collaboration unit, and the rendering results are sent to the headset in low-latency encoded streams to ensure smooth experience on the terminal device. The AI-driven unit analyzes the user's behavior data in the cloud. For example, it finds that most users will walk towards the dragon after clapping with the tree spirit. Therefore, the AI will pre-load the higher precision model and interaction script of the dragon to the edge server, so that when the user approaches, the more delicate interaction can be triggered instantly, avoiding loading delay;

[0049] When switching modes, audience A can switch to VR mode by voice command "enter the fantasy world" or gesture. At this time, the MR headset closes the camera and the screen no longer displays the real environment. The system uses the three-dimensional grid model of the exhibition hall reconstructed by the SLAM unit in advance as the basic geometric structure of the virtual environment, and superimposes the virtual "fantasy forest" scene on it. The user feels that he has completely left the exhibition hall and entered the dense virtual forest, but the "space boundary" of the forest is completely consistent with the real exhibition hall, avoiding collision risk.

[0050] The real-time construction and interaction method of the animation scene of augmented reality and virtual fusion is applied to any one of the real-time construction and interaction system of the animation scene of augmented reality and virtual fusion, comprising the following steps:

[0051] S1, through the sensor, real environment spatial data and user terminal equipment real-time pose data are collected, a fusion scene coordinate system is constructed and maintained; S2, virtual animation characters and scene elements are imported or created in real time, and are anchored and registered with the spatial features of the real environment, so that the position and attitude of the virtual elements in the physical space are stable; S3, based on the user's perspective and the lighting conditions of the physical environment, real-time light and shadow calculation and rendering are carried out on the virtual animation elements, so that they are visually seamlessly integrated with the real environment, and a virtual-real integrated animation scene is generated; S4, the user's interactive behavior is continuously captured through multi-modal sensors, and the interactive intention is recognized; S5, according to the recognized interactive intention, the virtual elements in the fusion scene are driven to make corresponding behavior feedback or change the scene state, realizing natural interaction between the user and the animation scene.

[0052] It should be noted that when the application is used, the system first uses SLAM (simultaneous localization and mapping) technology to scan, locate and construct a three-dimensional map of the real environment in real time by using the sensors of the terminal device, and creates an accurate, digital fusion scene coordinate system, so that virtual animation characters, objects and other elements can be accurately and stably "placed" at a specific position in the real world through scene anchors; the system analyzes the lighting information (intensity, direction, color) of the real environment in real time through the environment lighting estimation unit, then the adaptive rendering unit uses these data to dynamically adjust the coloring (coloring), shadow and reflection effect of the virtual elements, ensuring that the light and shadow properties of the virtual elements completely match the real environment they are in, so that the virtual characters can cast shadows in accordance with the direction of the real light, and their materials can reflect the color of the ambient light, thereby perfectly "integrating" into the real environment visually, eliminating the "sense of detachment" and "unreal feeling"; the system continuously captures the user's gestures, voice, eye movements, and the motion of physical props and recognition markers through the sensor group (camera, microphone, etc.) of the terminal device; high-computing-load tasks (such as complex light and shadow rendering, high-quality model processing, multi-user data synchronization) are offloaded to the cloud collaborative unit for processing. The cloud also has an AI driving unit built-in, which analyzes user behavior data, so that multiple users can see consistent virtual content and cooperate with each other at the same time; the MR head-mounted display used by the user can be freely switched between AR perspective mode and VR immersion mode. In VR mode, the system uses the real environment three-dimensional grid constructed by SLAM as the basis for the virtual environment, so that the user can choose to interact with the virtual character in the real exhibition hall (AR mode) or completely immerse themselves in a fantastic animation world that is visually completely virtualized but mapped from the real space (VR mode), and the two modes share the same set of interaction logic and scene content.

[0053] To sum up, the application is no longer a simple technical superposition or switching, but deeply fuses the two from the bottom coordinate system, forms a unified "mixed reality" (MR) experience, the system perceives the environment (light, space) and the user (gesture, voice) in real time, renders the picture and processes the interaction in real time according to the environment and the user, forms a highly adaptive dynamic closed loop system, through the collaborative calculation of "cloud edge-terminal", assigns the appropriate task to the most appropriate computing unit, and realizes the high-quality experience on the mobile terminal which originally needs a large workstation.

[0054] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A real-time interactive system for constructing animation scenes that integrates augmented reality and virtuality, characterized in that: It includes a scene construction and management module, a virtual and real rendering engine module, a multimodal interaction processing module, user terminal devices, and a cloud collaboration unit; The scene construction and management module is used to receive, process and store scene data from the virtual and real environments, and generate a unified fusion scene coordinate system in which virtual and real elements coexist. The virtual and real rendering and engine module is connected to the scene construction and management module, and is used to perform integrated and consistent three-dimensional real-time rendering of virtual animation elements and real environment according to the user's perspective and position. The multimodal interaction processing module is used to capture and recognize the interaction commands issued by the user through gestures, voice, eye movement or physical props, and parse the commands into operation commands for virtual elements or virtual-real combined elements in the fused scene; The user terminal device includes an AR display device or a VR display device for presenting the fused scene, and a sensor group for capturing user interaction commands. The cloud-based collaborative unit is used to distribute and process high-computation-load rendering tasks and complex interaction logic, and to support multiple users in sharing and persisting a unified and integrated scenario.

2. The real-time interactive system for constructing animation scenes that integrates augmented reality and virtual reality according to claim 1, characterized in that, The scene construction and management module includes a SLAM unit and a scene anchor point management unit; The SLAM unit is used for real-time map building, localization, and tracking of the real environment; The scene anchor point management unit is used to define and manage the anchor points of virtual elements in real space. The anchor points include image feature points, planes, objects, or custom spatial coordinates.

3. The real-time interactive system for constructing animation scenes that integrates augmented reality and virtual reality according to claim 1, characterized in that, The virtual-real rendering engine module includes an ambient lighting estimation unit and an adaptive rendering unit; The ambient lighting estimation unit is used to analyze the light intensity, direction, and color temperature of the real environment in real time; The adaptive rendering unit is used to dynamically adjust the shading, shadows, and highlights of virtual elements based on the analysis results of the ambient lighting estimation unit, so as to match the lighting conditions of the real environment.

4. The real-time interactive system for constructing animation scenes that integrates augmented reality and virtual reality according to claim 1, characterized in that, The multimodal interaction processing module supports the recognition and tracking of physical props, which are physical objects with specific visual markers or shapes. The system can bind virtual effects or models to the physical props held by the user in space, making the physical props interactive tools or weapons for the user in the integrated scene.

5. The real-time interactive system for constructing animation scenes that integrates augmented reality and virtual reality according to claim 1, characterized in that, The system supports the "VR construction-AR experience" mode: creators first construct and lay out animation scenes in a completely virtual VR environment; after completion, the system maps the VR scene to a designated real physical space; when the experiencer enters the physical space through an AR device, they can see and interact with the complete scene created in VR.

6. The real-time interactive system for constructing animation scenes that integrates augmented reality and virtual reality according to claim 1, characterized in that, The user terminal device is an MR headset, which can seamlessly switch between AR perspective mode and VR immersive mode. When switching to VR mode, the system uses a 3D model reconstructed from the real environment as a basis to virtualize real scene elements and present them together with virtual animation elements.

7. The real-time interactive system for constructing animation scenes that integrates augmented reality and virtual reality according to claim 1, characterized in that, The cloud-based collaboration unit has a built-in AI-driven unit, which is used to analyze the user's behavior patterns and preload the virtual assets that the user may need.

8. A method for real-time construction and interaction of augmented reality and virtual reality integrated animation scenes, applied to the real-time construction and interaction system for augmented reality and virtual reality integrated animation scenes as described in any one of claims 1-7, characterized in that, Includes the following steps: S1. Collect spatial data of the real environment and real-time pose data of the user terminal device through sensors to construct and maintain a fused scene coordinate system; S2. Import or create virtual anime characters and scene elements in real time, and anchor and register them with the spatial features of the real environment; S3. Based on the user's perspective and the lighting conditions of the physical environment, perform real-time light and shadow calculation and rendering on virtual animation elements to seamlessly integrate them with the real environment and generate an animation scene that blends the virtual and the real. S4. Continuously capture user interaction behavior through multimodal sensors to identify their interaction intentions; S5. Based on the identified interaction intent, drive the virtual elements in the fusion scene to make corresponding behavioral feedback or change the scene state.