Immersive literature interaction system based on WebGL and 3DGS

By combining WebGL and 3DGS technologies, it achieves efficient generation and rendering of photorealistic 3D scenes on the browser side, supports free navigation and interaction for users, solves the problems of high cost and poor interactivity in existing technologies, and provides efficient creation and immersive experience for immersive literature.

CN121564205APending Publication Date: 2026-02-24CHONGQING THREE GORGES UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511681444.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing immersive literary works suffer from problems in 3D scene presentation, such as cumbersome and costly high-precision model production and large resource file size, making it difficult to load quickly and render smoothly in real time in a web browser environment. At the same time, panoramic image technology has poor interactivity, preventing users from moving and exploring freely, which affects the depth of immersion.

Method used

It adopts an immersive literary interaction system based on WebGL and 3DGS, combining 3D Gaussian splashing technology with the real-time graphics capabilities of WebGL. It reconstructs high-precision 3D scenes through multi-source data and performs lightweight optimization. It utilizes WebGL's hardware acceleration rendering to support dynamic lighting and viewpoint switching. Combined with an AI-driven interactive narrative module, it enables users to navigate and interact freely.

Benefits of technology

It achieves photorealistic 3D scene rendering, lowers the creative threshold and cost, provides users with an immersive reading experience of free exploration and interaction, breaks through the interactive barriers of panoramic images, and enhances the immersiveness and artistic expression of the work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564205A_ABST
    Figure CN121564205A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer graphics and human-computer interaction, in particular to an immersive literature interaction system based on WebGL and 3DGS, which comprises a scene reconstruction and data processing module based on 3DGS, a 3DGS real-time rendering engine based on WebGL and an interaction narration module fused with a 3DGS scene. The 3D GS-based scene reconstruction and data processing module is used for reconstructing a high-precision 3D scene from multi-source data and executing lightweight optimization; the WebGL-based 3DGS real-time rendering engine is used for loading and rendering a 3D Gaussian splash model at a browser end, supporting dynamic illumination, texture and viewpoint switching and realizing low-delay interaction; and the interactive narrative module fusing the 3DGS scene associates literature content through AI driving logic, triggers scene visualization response and integrates a role dialogue system to realize immersive narrative experience. The module realizes seamless fusion of literature interaction logic and a 3DGS scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer graphics and human-computer interaction, and in particular to a system based on WebGL 3D rendering technology and 3D Gaussian Splatting (3DGS) technology for realizing immersive literary reading and interaction on the browser side. Background Technology

[0002] Existing immersive literary or digital narrative works mainly rely on two technical approaches to present three-dimensional scenes: one is to use pre-generated detailed three-dimensional models, and the other is to use 360-degree panoramic images or videos.

[0003] However, the method of using pre-built 3D models has obvious drawbacks: the production process of high-precision models is cumbersome, time-consuming, and costly, and the final generated resource files are huge, making it difficult to achieve fast loading and smooth real-time rendering in a web browser environment, which seriously affects the universality and immediacy of user experience.

[0004] While panoramic imagery solves the problems of resource size and rendering efficiency, it suffers from extremely poor interactivity. Users can only view the scene from a fixed point of view, unable to move freely or explore. This passive experience greatly limits the depth of immersion, contradicting the spirit of exploration and discovery inherent in literary works.

[0005] Therefore, there is an urgent need in this field for a web-based 3D scene solution that can balance high visual fidelity, efficient network transmission and rendering performance, and support free user navigation. Summary of the Invention

[0006] The present invention aims to overcome the shortcomings of the prior art and solve the problem of how to efficiently generate, transmit and render 3D scenes with photorealistic quality and allowing users to interact freely in the lightweight, cross-platform environment of a web browser, thereby providing a solid technical foundation for immersive literary experiences.

[0007] To address the aforementioned technical problems, this invention provides an immersive literary interaction system based on WebGL and 3DGS. The core of this system lies in its innovative combination of 3D Gaussian splashing—an advanced scene representation and rendering technology—with the real-time graphics capabilities of WebGL and the logical requirements of literary interaction, thus constructing a complete technology stack. An immersive literary interaction system based on WebGL and 3DGS includes a 3DGS-based scene reconstruction and data processing module, a WebGL-based 3DGS real-time rendering engine, and an interactive narrative module that integrates 3DGS scenes.

[0008] The 3DGS-based scene reconstruction and data processing module is used to reconstruct high-precision 3D scenes from multi-source data and perform lightweight optimization. The WebGL-based 3DGS real-time rendering engine is used to load and render 3D Gaussian splash models on the browser side, supporting dynamic lighting, textures, and viewpoint switching to achieve low-latency interaction; this engine is the core component running in the user's browser. It uses WebGL to call the GPU for hardware-accelerated rendering. Its innovation lies in the deep optimization of the standard 3DGS rendering pipeline for the web environment.

[0009] The interactive narrative module, which integrates 3DGS scenes, uses AI-driven logic to connect literary content, triggering visual responses in the scenes and integrating a character dialogue system to achieve an immersive narrative experience. This module achieves seamless integration of literary interactive logic and 3DGS scenes.

[0010] Furthermore, the method for achieving immersive literary interaction through the interactive system includes the following steps: Step 1) Use ordinary digital devices to acquire multi-angle image sequences of literary scenes and input them into the 3DGS-based scene reconstruction and data processing module; Step 2) The image sequence is trained and reconstructed through the 3DGS-based scene reconstruction and data processing module to generate a scene representation with 3D Gaussian ellipsoid attributes, and a lightweight scene data package is generated after pruning and quantization optimization. Step 3) The WebGL-based 3DGS real-time rendering engine initializes the rendering context on the browser side, asynchronously downloads the lightweight scene data package, parses it and converts it into a buffer object that the GPU can recognize, and calls the GPU for hardware-accelerated rendering. Step 4) The interactive narrative module that integrates 3DGS scenes merges low-poly interactive objects with 3DGS scenes, allowing users to navigate freely in the scene. When interactive hotspots are triggered, narrative branches, text prompts, or sound effects are activated, achieving immersive literary interaction.

[0011] Furthermore, the working process of the 3DGS-based scene reconstruction and data processing module includes: Receive sequences of literary scene images captured by ordinary digital devices. The images must cover key areas of the scene and multiple perspectives of objects, and the lighting must be uniform, without overexposure or blur. Training and reconstruction are carried out on the server side using the 3DGS algorithm. The process involves initializing a set of Gaussian ellipsoids, calculating multi-view projection errors using image sequences, and updating ellipsoid attributes through backpropagation. This process completes the 3D reconstruction of the literary scene and generates a scene representation composed of 3D Gaussian ellipsoids. The attributes of the Gaussian ellipsoids include position, color, transparency, scaling, and rotation. The scene representation is optimized to generate lightweight scene data packages for easy network transmission. The optimizations include Gaussian point cloud pruning based on visual contribution and quantization and compression of attribute data.

[0012] Furthermore, the process of optimizing the scene representation involves pruning and removing Gaussian points with low visual contribution, quantizing to reduce the floating-point precision of Gaussian attributes, and finally outputting a lightweight scene data package in .glb format of 50-100MB.

[0013] Furthermore, The rendering process of the WebGL-based 3DGS real-time rendering engine includes: The browser loads the JavaScript core engine and initializes the rendering context through WebGL to ensure that the rendering environment is adapted to different terminal devices; The engine asynchronously downloads lightweight scene data packages, and after the download is complete, it parses the data in memory and converts it into a buffer object that the GPU can understand. We developed an optimized GLSL shader program. Addressing the core challenges of 3DGS in the WebGL rendering environment, we adopted a strategy of "pre-sorting by view frustum distance + fragment-level depth correction" to solve the depth sorting problem. We also used a weighted alpha blending algorithm to process the transparency decay of multiple Gaussian superpositions, thus achieving screen-space splash rendering of 3D Gaussian elements. It integrates frustum culling and detail-level control mechanisms to ensure smooth interactive frame rates on terminal devices with different performance levels.

[0014] The WebGL-based 3DGS real-time rendering engine receives and parses lightweight 3DGS scene data. By writing highly optimized GLSL shader programs, it performs screen-space splash rendering on each 3D Gaussian primitive, correctly implementing alpha blending and depth sorting. Simultaneously, the engine integrates mechanisms such as frustum culling and level-of-detail control to ensure smooth interactive frame rates on terminal devices with varying performance.

[0015] Furthermore, the interaction process of the interactive narrative module that integrates 3DGS scenes includes: In the system editing backend, the creator loads the 3DGS scene and creates an "interactive hotspot" with an interactive radius in a specific three-dimensional spatial coordinate system; Link interactive hotspots with literary content, which includes scanned images, ambient background sounds, and plot text. A hybrid rendering strategy is adopted to represent key interactive objects in literature as low-poly models, and embed them into the 3DGS scene through spatial coordinate transformation and lighting matching. When a user navigates the virtual camera close to an interactive hotspot, the system calculates the three-dimensional spatial distance between the virtual camera and the interactive hotspot in real time by combining spatial collision detection and coordinate ranging in the virtual scene. When the distance meets the preset interaction radius condition, a prompt icon is displayed. After the user triggers the interaction, scene navigation is paused, related content pops up, and the corresponding sound effect is played.

[0016] To address the limitations of 3DGS in representing dynamic objects, this invention employs a hybrid rendering strategy: key interactive objects in the literature are represented using traditional low-poly models, and through spatial coordinate transformation and lighting matching, they are naturally embedded into a static, photorealistic background constructed by 3DGS. Users can navigate freely within the scene, and when they approach interactive objects, corresponding narrative branches, text prompts, or sound effects are triggered, thereby advancing the storyline.

[0017] Furthermore, when capturing image sequences with ordinary digital devices, it is necessary to take overlapping shots from multiple angles around the literary scene, covering all corners of the scene and key objects. The number of shots should be controlled between 100 and 200. During the shooting process, the device should be kept stable, and attention should be paid to uniform lighting to avoid image blurring.

[0018] Furthermore, in the pruning operation of the 3DGS-based scene reconstruction and data processing module, Gaussian points with low visual contribution are removed. Gaussian points with low visual contribution include Gaussian points with extremely high transparency and Gaussian points that are always at the edge of the field of view. In the quantization operation, the floating-point precision of the Gaussian attribute is appropriately reduced to reduce the amount of data. The reduction in floating-point precision is based on the premise that it does not affect the visual fidelity of the scene.

[0019] Furthermore, when the WebGL-based 3DGS real-time rendering engine initializes the rendering context, it automatically adapts to the browser type and terminal GPU performance and adjusts the rendering parameters accordingly. When parsing scene data packets, it identifies key areas through predefined rules. These key areas include the core landscapes with the highest visual weight in the scene, anchor point areas that are bound to the plot interaction logic, and the initial camera frustum coverage area. It prioritizes loading data from these key areas to achieve progressive rendering and improve the user's initial experience.

[0020] Furthermore, the interactive narrative module integrating 3DGS scenes also supports AI character interaction, the process of which includes: Register 3D coordinates in the 3DGS scene as "anchor points" for AI characters, and bind AI entities to the "anchor points"; AI characters are presented through pre-made 3D cartoon images or 2D portraits. During rendering, the AI ​​characters always face the camera and are superimposed on the 3DGS scene. The browser captures the user's voice and converts it to text via the Web Speech API. The text is then sent to the backend AI big model, such as the Tongyi Qianwen big model, via WebSocket. The system constructs a character setting mechanism for each poet, which includes detailed settings in multiple dimensions such as basic information, personality traits, representative works, creative style, and life experience. The AI ​​big model generates responses that match the character settings and plays them through a TTS engine or pre-made audio to achieve multimodal scenario-based dialogue.

[0021] Compared with the prior art, the present invention has the following beneficial effects: This invention achieves a balance between photorealistic rendering and high performance on the web: utilizing 3DGS technology, it can reconstruct highly realistic 3D scenes from real-world images, with visual quality far exceeding that of traditional 3D modeling. Simultaneously, 3DGS boasts extremely high rendering efficiency; after optimization by this invention, it can achieve real-time, smooth rendering on browsers of ordinary computers and mobile phones, breaking the dependence of high-quality graphics experiences on native applications.

[0022] It significantly reduces the barriers to creation and costs: content creators do not need to master complex 3D modeling skills and can quickly generate the materials needed for literary scenes using ordinary photography equipment, which greatly simplifies the production process of immersive content, shortens the production cycle, and is conducive to the promotion and popularization of the technology.

[0023] It provides an unprecedented immersive reading experience: This invention breaks through the interactive barriers of panoramic image technology, allowing readers to freely wander and explore in highly realistic literary scenes, observe details from any angle, and interact with elements in the environment, transforming passive reading into active exploration, and greatly enhancing the immersion, engagement, and artistic expression of the work. Attached Figure Description

[0024] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram illustrating the technical implementation process of the immersive literary interaction system based on WebGL and 3DGS in this invention.

[0025] Figure 2 This is a block diagram of the overall architecture of the immersive literary interaction system based on WebGL and 3DGS in this invention.

[0026] Figure 3 This is a flowchart of the 3DGS-based scene reconstruction and data processing module in this invention.

[0027] Figure 4This is a schematic diagram illustrating the rendering principle of the WebGL-based 3DGS real-time rendering engine in this invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] like Figures 1-4 As shown, the immersive literary interaction system based on WebGL and 3DGS in this invention includes a 3DGS-based scene reconstruction and data processing module, a WebGL-based 3DGS real-time rendering engine, and an interactive narrative module integrating 3DGS scenes. The 3DGS-based scene reconstruction and data processing module reconstructs high-precision 3D scenes from multi-source data and performs lightweight optimization. The WebGL-based 3DGS real-time rendering engine loads and renders 3D Gaussian splash models on the browser side, supporting dynamic lighting, textures, and viewpoint switching to achieve low-latency interaction; this engine is the core component running in the user's browser. It uses WebGL to call the GPU for hardware-accelerated rendering. Its innovation lies in the deep optimization of the standard 3DGS rendering pipeline for the Web environment. The interactive narrative module integrating 3DGS scenes uses AI-driven logic to associate literary content, triggering scene visualization responses and integrating a character dialogue system to achieve an immersive narrative experience. This module achieves seamless integration of literary interaction logic and 3DGS scenes.

[0030] The method for implementing immersive literary interaction in the WebGL and 3DGS-based immersive literary interaction system of this invention includes the following steps: S1. Use digital devices to collect multi-angle image sequences of real scenes, and use Blender to construct high-precision 3D models of poets. S2. Upload the image to the cloud, train the Gaussian point cloud scene using the 3DGS algorithm, and perform lightweight processing such as pruning and quantization. S3. Integrate the optimized 3DGS scene, poet model, and interactive hotspots, and configure the AI ​​character attributes and dialogue style; S4. The user accesses the system through a browser, loads the WebGL engine and downloads all resources in one step, and initializes the rendering context and interaction module; S5 features a WebGL-based rendering engine that renders 3DGS scenes and poet models in real time, supporting free-view navigation and hybrid rendering. S6. The system captures user voice questions or receives text box input through the Web Speech API, and the backend calls LLM to generate text responses that conform to the poet role settings. S7 enables stereoscopic rendering and spatial interaction in XR devices such as Vision Pro, providing a realistic dialogue experience and achieving multimodal fusion of voice, vision, and interaction; S8 dynamically optimizes rendering performance to ensure smooth frame rates while maintaining dialogue context and interaction state to guarantee a consistent user experience.

[0031]

Example 1

[0032] Step 1: Data Collection Using a regular smartphone, enter the actual room that needs to be digitized. Open the phone's camera app and take multiple, overlapping photos around the room from various angles. Make sure to cover all corners of the room and multiple perspectives of key objects (such as the desk, bookshelf, and lamp). Take approximately 100-200 photos in total. Ensure even lighting during shooting to avoid overexposure or blurriness.

[0033] Step Two: Cloud-based 3DGS Reconstruction and Lightweight Processing The photos acquired in step one are uploaded to the system's cloud processing server via the internet. The server will then initiate a 3DGS reconstruction task. After model optimization training is complete, the generated 3DGS model is optimized as follows: a. Pruning: Remove Gaussian points that contribute little to visual perception (such as points with extremely high transparency or located at the edge of the field of view).

[0034] b. Quantization: Appropriately reduce the floating-point precision of the Gaussian property to reduce the amount of data.

[0035] c. Output: The server outputs a lightweight, .glb format scene data package optimized specifically for this system. The size of this data package can be controlled between 50-100MB.

[0036] Step 3: Web-side loading and rendering The browser automatically loads the system's JavaScript core engine and initializes the rendering context via WebGL. The engine asynchronously downloads the .glb data package from the server. Once downloaded, the engine parses the data in memory and converts it into a GPU-understandable buffer object, giving the user an immersive experience of freely roaming the room within the browser.

[0037]

Example 2

[0038] Step Two: Literary Content Relevance and Triggering Logic The creators associate the aforementioned hotspots with literary content, such as a scanned image of a manuscript, background sound, or text revealing the plot. On the web interface, when the user navigates the virtual camera closer to the desk and the distance to the hotspot enters the interaction radius, a prompt icon appears on the system UI. After the user interacts with the button, the system pauses scene navigation, displays the associated manuscript image and text description on the side of the screen, and plays a page-turning sound.

[0039] Step 3: Implementing Hybrid Rendering To achieve more complex interactions, this invention employs hybrid rendering. For example, a desk drawer is modeled as a simple, animable low-poly model. During rendering, the WebGL engine first renders the background 3DGS study scene, and then, in the same rendering pass, continues to render this low-poly model of the drawer. Through precise lighting and shadow calculations, the lighting environment of the drawer model is ensured to be consistent with the 3DGS scene, achieving a seamless visual integration. When the user interacts with the drawer, an opening / closing animation is triggered, thus enabling dynamic interaction in a highly realistic scene and greatly enhancing the possibilities for storytelling.

[0040]

Example 3

[0041] Step Two: Multimodal Interaction Process Visual presentation: The AI ​​character can be a pre-made 3D cartoon image or a more artistic 2D portrait. During rendering, the image is always facing the camera and drawn on top of the 3DGS scene.

[0042] Voice interaction triggered: a. The user asks: "Who are you?" b. The browser captures the user's speech and converts it into text via the Web Speech API.

[0043] c. Text data is sent in real time to the backend AI large model service via WebSocket.

[0044] d. The AI ​​model generates text responses that match the character's persona based on the role they are playing and the background of the literary story.

[0045] e. The reply text is returned to the browser, and at the same time, the system calls the browser's TTS engine or plays a pre-made AI voice audio file.

[0046] Contextualized dialogue: The entire dialogue takes place in a study environment constructed by 3DGS. The AI's responses can be strongly related to the details of the scene, achieving in-depth environmental narrative.

Claims

1. An immersive literary interaction system based on WebGL and 3D Gaussian splashing, characterized in that, This includes a 3DGS-based scene reconstruction and data processing module, a WebGL-based 3DGS real-time rendering engine, and an interactive narrative module that integrates 3DGS scenes. The 3DGS-based scene reconstruction and data processing module is used to reconstruct high-precision 3D scenes from multi-source data and perform lightweight optimization. The WebGL-based 3DGS real-time rendering engine is used to load and render 3D Gaussian splash models on the browser side, supporting dynamic lighting, textures, and viewpoint switching, and achieving low-latency interaction. The interactive narrative module that integrates 3DGS scenes uses AI-driven logic to connect literary content, trigger scene visualization responses, and integrates a character dialogue system to achieve an immersive narrative experience.

2. The immersive literary interaction system based on WebGL and 3D Gaussian splashing as described in claim 1, characterized in that, The method for achieving immersive literary interaction through the interactive system includes the following steps: Step 1) Use ordinary digital devices to collect multi-angle image sequences of literary scenes and input them into the 3DGS-based scene reconstruction and data processing module; Step 2) The image sequence is trained and reconstructed through the 3DGS-based scene reconstruction and data processing module to generate a scene representation with 3D Gaussian ellipsoid attributes, and a lightweight scene data package is generated after pruning and quantization optimization. Step 3) The WebGL-based 3DGS real-time rendering engine initializes the rendering context on the browser side, asynchronously downloads the lightweight scene data package, parses it and converts it into a buffer object that the GPU can recognize, and calls the GPU for hardware-accelerated rendering. Step 4) The interactive narrative module that integrates 3DGS scenes merges low-poly interactive objects with 3DGS scenes, allowing users to navigate freely in the scene. When interactive hotspots are triggered, narrative branches, text prompts, or sound effects are activated, achieving immersive literary interaction.

3. The immersive literary interaction system based on WebGL and 3DGS according to claim 1, characterized in that, The working process of the 3DGS-based scene reconstruction and data processing module includes: Receive sequences of literary scene images captured by ordinary digital devices. The images must cover key areas of the scene and objects from multiple perspectives, and the lighting must be uniform, without overexposure or blur. Training and reconstruction are carried out on the server side using the 3DGS algorithm. The process involves initializing a set of Gaussian ellipsoids, calculating multi-view projection errors using image sequences, and updating ellipsoid attributes through backpropagation. This process completes the 3D reconstruction of the literary scene and generates a scene representation composed of 3D Gaussian ellipsoids. The attributes of the Gaussian ellipsoids include position, color, transparency, scaling, and rotation. The scene representation is optimized to generate a lightweight scene data package. The optimizations include Gaussian point cloud pruning based on visual contribution and quantization and compression of attribute data.

4. The immersive literary interaction system based on WebGL and 3DGS according to claim 3, characterized in that, The process of optimizing the scene representation involves pruning Gaussian points with low visual contribution, quantizing to reduce the floating-point precision of Gaussian attributes, and finally outputting a lightweight scene data package in .glb format of 50-100MB.

5. The immersive literary interaction system based on WebGL and 3DGS according to claim 1, characterized in that, The rendering process of the WebGL-based 3DGS real-time rendering engine includes: The browser loads the JavaScript core engine and initializes the rendering context through WebGL to ensure that the rendering environment is adapted to different terminal devices; The engine asynchronously downloads lightweight scene data packages, and after the download is complete, it parses the data in memory and converts it into a buffer object that the GPU can understand. We wrote an optimized GLSL shader program, and addressed the core difficulties of 3DGS in the WebGL rendering environment by adopting the strategy of "pre-sorting by view frustum distance + fragment-level depth correction" to solve the depth sorting problem. We also used a weighted alpha blending algorithm to process the transparency decay of multiple Gaussian superpositions, and achieved screen space splash rendering of 3D Gaussian elements. It integrates frustum culling and detail-level control mechanisms to ensure smooth interactive frame rates on terminal devices with different performance levels.

6. The immersive literary interaction system based on WebGL and 3DGS according to claim 1, characterized in that, The interactive process of the interactive narrative module that integrates 3DGS scenes includes: In the system editing backend, the creator loads the 3DGS scene and creates an "interactive hotspot" with an interactive radius in a specific three-dimensional spatial coordinate system; Link interactive hotspots with literary content, which includes scanned images, ambient background sounds, and plot text. A hybrid rendering strategy is adopted to represent key interactive objects in literature as low-poly models, and embed them into the 3DGS scene through spatial coordinate transformation and lighting matching. When a user navigates the virtual camera close to an interactive hotspot, the system calculates the three-dimensional spatial distance between the virtual camera and the interactive hotspot in real time by combining spatial collision detection and coordinate ranging in the virtual scene. When the distance meets the preset interaction radius condition, a prompt icon is displayed. After the user triggers the interaction, scene navigation is paused, related content pops up, and the corresponding sound effect is played.

7. An immersive literary interaction system based on WebGL and 3DGS according to claim 2, characterized in that, When capturing image sequences with ordinary digital devices, it is necessary to take multiple overlapping shots around the literary scene, covering all corners of the scene and key objects. The number of shots should be controlled between 100 and 200. During the shooting process, the equipment should be kept stable, the lighting should be uniform, and image blur should be avoided.

8. An immersive literary interaction system based on WebGL and 3DGS according to claim 2, characterized in that, In the pruning operation of the 3DGS-based scene reconstruction and data processing module, Gaussian points with low visual contribution are removed. Gaussian points with low visual contribution include Gaussian points with extremely high transparency and Gaussian points that are always at the edge of the field of view. In the quantization operation, the floating-point precision of the Gaussian attribute is appropriately reduced to reduce the amount of data. The reduction in floating-point precision is based on the premise that it does not affect the visual fidelity of the scene.

9. An immersive literary interaction system based on WebGL and 3DGS according to claim 2, characterized in that, When the WebGL-based 3DGS real-time rendering engine initializes the rendering context, it automatically adapts to the browser type and terminal GPU performance and adjusts the rendering parameters accordingly. When parsing scene data packets, it identifies key areas through predefined rules. These key areas include the core landscape with the highest visual weight in the scene, the anchor point areas that are bound to the plot interaction logic, and the coverage area of ​​the initial camera frustum. It prioritizes loading data from key areas of the scene to achieve progressive rendering and improve the user's initial experience.

10. An immersive literary interaction system based on WebGL and 3DGS according to claim 5, characterized in that, The interactive narrative module that integrates 3DGS scenes also supports AI character interaction, the process of which includes: Register 3D coordinates in the 3DGS scene as "anchor points" for AI characters, and bind AI entities to the "anchor points"; AI characters are presented through pre-made 3D cartoon images or 2D portraits. During rendering, the AI ​​characters always face the camera and are superimposed on the 3DGS scene. The browser captures the user's voice through the Web Speech API and converts it into text. The text is sent to the backend AI model via WebSocket to generate a response that matches the user's role. The response is then played through a TTS engine or pre-made audio to achieve multimodal contextual dialogue.

Citation Information

Cited By

  • Hybrid rendering method, device and equipment for 3DGS scene and medium

    CN121746569A