Space computing-oriented cloud rendering method and system and terminal equipment
By compositing and transmitting 3D virtual space in real time on a cloud rendering platform, the problem of computing power and power consumption limitations of terminal devices is solved, achieving a high-performance spatial computing experience and compatibility with 2D applications on lightweight terminal devices, and providing a persistent workspace.
Patent Information
- Application Number
- CN202511453669.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies in spatial computing are limited by the computing power and power consumption of terminal devices, making it impossible to provide real-time rendering of high-fidelity 3D scenes. Furthermore, 2D applications cannot run efficiently in 3D environments, resulting in bulky and expensive devices or insufficient performance, and a fragmented application ecosystem.
It runs an operating system environment on a cloud rendering platform, synthesizes and renders 3D virtual space in real time, transforms the scene into a data stream through optimized encoding technology and transmits it to lightweight terminal devices, supports six degrees of freedom interaction and eye tracking, is compatible with the 2D application ecosystem, and achieves persistence through session management.
It achieves a high-performance spatial computing experience on lightweight terminal devices, is compatible with the 2D application ecosystem, provides a "never offline" personal space workspace, and expands application scenarios and productivity.
Smart Images

Figure CN121387154A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of cloud computing technology, computer graphics and human-computer interaction technology, and in particular to a method, system and terminal device for large-scale real-time rendering in the spatial computing scenario and streaming the results to a lightweight terminal device. Background Technology
[0002] Spatial computing technology, represented by virtual reality (VR) and augmented reality (AR), is leading the transformation of human-computer interaction from two-dimensional screens to three-dimensional space. However, its development faces two core challenges: First, the physical limitations of computing power and power consumption. Real-time rendering of high-fidelity 3D scenes places a huge burden on the chips, heat dissipation, and battery life of terminal devices, resulting in bulky and expensive high-end devices, while lightweight devices suffer from insufficient performance. Second, the fragmented application ecosystem. The massive amount of existing 2D applications cannot run directly and efficiently in a 3D environment, limiting the productivity application scenarios of spatial computing devices. Although existing cloud desktop technologies have enabled computing to the cloud, their 2D image transmission protocols cannot meet the demands of spatial computing for a three-dimensional immersive experience. Therefore, the market urgently needs a new technological solution that can integrate cloud computing power with spatial computing experience and is compatible with the existing application ecosystem. Summary of the Invention
[0003] To overcome the shortcomings of existing technologies, the present invention aims to provide a cloud rendering method, system and terminal device for spatial computing, which aims to break through the bottlenecks of computing power, power consumption and ecosystem of terminal hardware, and provide users with a borderless, high-performance and persistent personal space workspace that can be experienced on lightweight devices.
[0004] To achieve the above objectives, the present invention provides the following solution:
[0005] A cloud rendering method for spatial computing includes: running an operating system environment capable of supporting two-dimensional and three-dimensional applications on a cloud rendering platform; performing real-time synthesis and rendering of the outputs of various applications in a three-dimensional virtual space to generate a unified stereoscopic scene; receiving low-latency user interaction data from terminal devices and updating the scene accordingly; and finally using optimized encoding technology to convert the stereoscopic scene into a spatial data stream and transmit it to the terminal devices.
[0006] Preferably, the operating system environment can render the interface of a traditional two-dimensional planar application as a virtual plane and seamlessly embed it as an object into a three-dimensional virtual space, coexisting with the three-dimensional space application.
[0007] Preferably, the user interaction data includes six degrees of freedom (6DoF) position and pose data of the terminal device, eye-tracking data of the user, and gesture recognition data captured by a camera or controller.
[0008] Preferably, the optimized encoding technique is gaze-based rendering encoding, which uses high-fidelity encoding for the central area where the user's gaze is focused in the stereoscopic scene, and low-fidelity encoding for the surrounding areas of the field of vision, based on the user's eye tracking data, in order to optimize bandwidth usage.
[0009] Preferably, the method further includes a session management mechanism that can save all application layouts and states in the user's three-dimensional virtual space, allowing the user to resume the session at different times or on different devices.
[0010] A cloud rendering system for spatial computing includes a cloud rendering platform and terminal devices. The cloud rendering platform is responsible for application operation, scene compositing, rendering, and encoding. The terminal devices are responsible for collecting and uploading user interaction data, as well as decoding and displaying the received spatial data streams.
[0011] Preferably, the cloud rendering platform has a built-in spatial application runtime, a 3D spatial compositor, a real-time rendering engine, and a spatial streaming encoder.
[0012] A spatial computing terminal device, which is a lightweight VR / AR headset, whose local computing tasks are mainly data stream decoding and sensor data processing.
[0013] This invention discloses the following beneficial effects: Through a cloud-edge collaborative architecture, this invention completely transfers heavy-load tasks of spatial computing to the cloud, enabling users to obtain performance exceeding the limits of local hardware through lightweight terminal devices and smoothly run professional-grade applications. Simultaneously, this solution is compatible with the 2D application ecosystem and, through session persistence capabilities, creates a "never-disconnected" personal workspace that can be roamed anywhere, greatly expanding the application scenarios and productivity value of spatial computing. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 A flowchart of a cloud rendering method for spatial computing provided in an embodiment of the present invention;
[0016] Figure 2 This is a schematic diagram of a cloud rendering system architecture for spatial computing provided in an embodiment of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0019] Example 1: Creation and Interaction Process of Cloud Space Workspace
[0020] Reference Figure 1 This example describes the complete workflow of a designer working with lightweight AR glasses.
[0021] Step S101: Initialization and Session Resumption
[0022] The user puts on AR glasses (the terminal device), and after the glasses are powered on, they connect to the cloud rendering platform via a built-in communication module. After the user authenticates, the system detects that the user has a previously saved session state. The system then loads that session and, during the runtime of the spatial application in the cloud, automatically launches applications that were still running when the user last closed them, such as a 3D modeling software (like Blender), a 2D data collection browser (like Chrome), and a design reference board application.
[0023] Step S102: 3D spatial compositing and initial rendering
[0024] The cloud-based 3D compositor receives output from various applications. It arranges and composites the views from 3D modeling software, browser windows, and reference panels in a virtual 3D space, according to the layout saved by the user last time. For example, a designer might have a huge 3D model in front of them, with a browser window floating on the left and a reference panel pinned to a virtual wall on the right. The real-time rendering engine then renders this composited 3D scene with high quality based on a default virtual camera perspective, generating the first frame of the stereoscopic scene image (left and right eye views).
[0025] Step S103: Data Streaming and Local Presentation
[0026] The generated stereoscopic scene image is fed into a spatial streaming encoder. The encoder uses efficient video coding algorithms (such as AV1 or HEVC) combined with foveated rendering technology to compress it into a spatial data stream. This data stream is transmitted to the user's AR glasses via a low-latency network (such as 5G or Wi-Fi 6). The decoder module on the glasses receives and decodes the data stream, reconstructs the stereoscopic image, and overlays the virtual workspace scene onto the user's real-world field of view through a display rendering module (such as waveguide lenses).
[0027] Step S104: Real-time interaction and scene update
[0028] The designer turns their head to observe the side of the 3D model. The AR glasses' sensor module captures the head's six degrees of freedom (6DoF) position and pose data and immediately uploads it as user interaction data to the cloud rendering platform. Upon receiving the new pose data, the cloud's real-time rendering engine immediately updates the virtual camera's position and orientation to match the designer's new viewpoint and re-renders the 3D scene from that perspective. This closed loop of "interaction-upload-cloud rendering-feedback-display" is completed in milliseconds, allowing the designer to observe virtual objects naturally with extremely low latency.
[0029] Next, the designer used gesture recognition to "grab" and enlarge the browser window in mid-air for a clearer view of the information. The gesture data was also captured and uploaded, and the cloud-based 3D spatial compositor updated the browser window's size and position in virtual space accordingly, which was then used by the rendering engine to render the updated scene. The entire process was smooth and natural, just like manipulating a real physical object.
[0030] Step S105: Session saving
[0031] After completing their work, the designer selects "End Session." The system completely saves the layout, state, and internal data of all applications in the current 3D virtual space. The next time the user logs in on any compatible terminal device, they can seamlessly resume their work from the current state.
[0032] Example 2: System Architecture and Key Module Collaboration
[0033] Reference Figure 2 This embodiment illustrates the collaborative working method of each module in the system architecture of the present invention.
[0034] Cloud Rendering Platform:
[0035] Spatial App Runtime: This is a core virtualization environment that can be understood as a powerful "spatial operating system" running on a cloud server. It can run multiple traditional 2D applications (such as the Office suite and Photoshop) and professional 3D spatial applications (such as CAD and Unity / UE projects) simultaneously. For 2D applications, it captures their graphics output through virtual graphics card technology and "draws" it onto a virtual 2D panel.
[0036] 3D Spatial Compositor: This module acts as the "scene director" of the virtual world. It manages all virtual panels and 3D models output by the application, and performs operations such as layout, scaling, and rotation on them in a unified 3D coordinate system according to the user's interactive commands, ultimately constructing a complete and well-defined 3D virtual scene.
[0037] Real-time Rendering Engine: This is the "graphics heart" of the system. It receives scene descriptions constructed by the compositor and virtual camera parameters from the terminal (determined by the user's head pose). Utilizing the powerful GPU resources in the cloud, it performs complex graphics calculations such as ray tracing and physical lighting to render photorealistic stereoscopic images in real time (rendering for the left and right eyes separately).
[0038] Spatial Streaming Encoder: To efficiently transmit high-quality rendering results, this module employs advanced encoding techniques. A key optimization is foveated rendering encoding: it receives user eye-tracking data from the terminal to determine the user's focal point. During encoding, it allocates a higher bit rate to the focal region to ensure maximum clarity, while using a lower bit rate for compression of the peripheral region, thus significantly reducing network bandwidth requirements while maintaining subjective visual quality.
[0039] Terminal Device:
[0040] Input and sensor acquisition module: integrates a 6DoF tracking sensor, an eye-tracking camera, a gesture recognition camera, etc., to continuously and frequently acquire the user's head, eye and hand dynamics.
[0041] Communication module: responsible for packaging and uploading the collected interactive data with minimal latency, and for stably receiving the spatial data stream transmitted back from the cloud.
[0042] Decoder module: Built-in hardware decoding chip, specifically designed for efficiently decoding the received spatial data stream and restoring it into left and right eye image frames.
[0043] Display Presentation Module: Synchronously projects the decoded image frames onto the display elements, providing users with a stable, clear, and immersive spatial computing experience.
[0044] Through the precise collaboration between the cloud and terminal modules, this invention places the vast majority of the computing load in the cloud, enabling terminal devices to be made extremely thin, light, and low-power, while providing users with a high-performance spatial computing experience far exceeding the capabilities of their local computing power.
Claims
1. A cloud rendering method for spatial computing, characterized in that, Includes the following steps: a. Running one or more applications on a cloud rendering platform and compositing the output of the applications in a three-dimensional virtual space; b. Receiving user interaction data from a terminal device; c. Based on the user interaction data, update the content status of the three-dimensional virtual space or the virtual camera perspective on the cloud rendering platform, and render and generate the corresponding stereoscopic scene image in real time; d. Encode the stereoscopic scene image into a spatial data stream and transmit it to the terminal device.
2. The method according to claim 1, characterized in that, The applications in step a include two-dimensional planar applications and three-dimensional spatial applications.
3. The method according to claim 1, characterized in that, The user interaction data includes at least one of the following: six degrees of freedom (6DoF) position and pose data of the terminal device, eye tracking data of the user, or gesture recognition data of the user.
4. The method according to claim 1, characterized in that, The encoding step in step d employs foveated rendering encoding technology.
5. The method according to claim 1, characterized in that, The method further includes saving the session state of the three-dimensional virtual space for restoration in subsequent sessions.
6. A cloud rendering system for spatial computing, characterized in that, include: a. A terminal device; b. A cloud rendering platform, communicating with the terminal device, the cloud rendering platform comprising: c. a spatial application runtime, adapted to simultaneously support two-dimensional planar applications and three-dimensional spatial applications; d. a three-dimensional spatial compositor, adapted to uniformly lay out and composite the output content of each application in the spatial application runtime in a three-dimensional virtual space; e. a real-time rendering engine, adapted to render the composited scene in real time based on user interaction data received from the terminal device; f. a spatial streaming encoder, adapted to encode the rendered stereoscopic scene image into a spatial data stream and send it to the terminal device.
7. A space computing terminal device, characterized in that, include: a. An input and sensor acquisition module, suitable for collecting user interaction data; b. A communication module, adapted to upload the user interaction data to a cloud rendering platform and receive a spatial data stream therefrom; c. A decoder module, adapted to decode the spatial data stream into a stereoscopic scene image; d. A display module, adapted to display the stereoscopic scene image to a user.