Methods for providing views of a virtual world and rendering servers

DE102024102700A1Pending Publication Date: 2025-07-31JOURNEE TECH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024102700
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-07-31

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for providing views of a virtual world, comprising the following steps: a. loading data describing a 3D model of a base virtual world into a rendering engine (14); b. inputting a position and / or a viewing direction (POS) into the rendering engine (14); c. receiving, preferably from the rendering engine (14), input images (IFs) representing a view of the virtual world in a first (low) resolution; d. capturing the input images (IFs) to generate selected data inputs; e. encoding the selected data input, preferably using an autoencoder (31); f. inputting the encoded selected data input (T(IFs)) into a machine learning model, e.g. a diffusion model (33), to generate encoded output frames (T(IFs)), and g. decoding the encoded output frames (T(IFs)) using a decoder (32).
Need to check novelty before this filing date? Find Prior Art

Claims

[1] A method for providing views in a virtual world, comprising the following steps: a. Loading data describing a 3D model of a virtual base world into a rendering engine (14); b. entering a position and / or a viewing direction (POS) into the rendering engine (14); c. receiving, preferably from the rendering engine (14), input images (IFs) representing a view of the virtual world in a first (low) resolution; d. Capturing the input images (IFs) to generate selected data inputs; e. encoding the selected data input, preferably using an autoencoder (31); f. inputting the coded selected data input (T(IFs)) into a machine learning model, e.g. a diffusion model (33), to generate coded output frames (T(OFs)); g. Decoding the coded output frames (T(OFs)) using a decoder (32). [2] The method of claim 1, wherein step f comprises inputting a tuple including at least one label or descriptive text (TXT) and the encoded selected data input (T(IFs)). [3] A method according to any one of the preceding claims, including the step of receiving at least one label or descriptive text (TXT) from a database based on an event and / or a position relative to the base virtual world rendered in the rendering engine (14). [4] A method according to any one of the preceding claims, wherein the machine learning model is used to perform a guided reconstruction based on the label and / or the descriptive text. [5] A method according to any one of the preceding claims, wherein step d comprises: Extracting g-buffers, e.g. depth buffer, ambient occlusion buffer, diffuse color buffer, position buffer and / or normal buffer. [6] Method according to one of the preceding claims, wherein the virtual base world is: - automatically generated; and / or - abstract; and / or - has shades of grey and / or is black and white; and / or has box-shaped structures. [7] A method according to any one of the preceding claims, including a step of caching previous encoded selected data inputs, wherein step f comprises applying at least one previous encoded selected data input together with the current selected data input to the machine learning model. [8] A method for training a neural network comprising the following steps: - Receiving a first data set of a high fidelity virtual world; - Transforming the first data set into a second data set such that the second data models a high-fidelity (low-fidelity) virtual world corresponding to the high-fidelity worlds; - rendering the virtual world with high fidelity in a first rendering engine (14); - rendering the low fidelity virtual world in a second rendering engine (14); and - Using the output of the first and second rendering engines (14) to train the neural network. [9] A method according to claim 8, comprising the steps of (randomly) generating a control command for navigating, e.g. an avatar, in the virtual world with high fidelity; Entering the control command into the first and second rendering machines (14); Using the images output by the first and second rendering engines (14) to train the neural network. [10] A method according to claim 8 or 9, comprising the following steps: Encode the images from the first and second rendering engines and use the encoded images to train the neural network. [11] A computer-readable medium storing instructions which, when executed by at least one processor, cause the at least one processor to implement a method according to any one of the preceding claims. [12] Rendering server (10) for rendering a virtual world, particularly suitable for implementing the method according to one of claims 1 to 10, wherein the rendering server (10) comprises: - a rendering engine (14) for generating input images when rendering views in a virtual base world; - an encoder (31), in particular an autoencoder, for encoding at least some of the input images and / or selected data inputs derived from at least some of the input images; - a machine learning model, in particular a diffusion model (33) for processing coded selected data output by the encoder, in particular for enriching the respective view and for generating coded output frames; and - a decoder (32) for decoding the coded output frames. [13] Rendering server (10) according to claim 12, wherein the rendering server (10) is embedded in a virtual environment (50). [14] Rendering server (10) according to claim 12 or 13, comprising a streamer (16) for (continuously) outputting streaming data based on the decoded output frames and for transmitting the streaming data to a user client (40).