VR real-time rendering method and system for low-bit-rate cloud rendering plug flow

By combining a lightweight super-resolution model with a multi-scale discriminator, the computation and bandwidth issues of 4K real-time rendering on VR devices are solved, achieving high-quality, low-latency 4K image rendering, suitable for a variety of demanding application scenarios.

CN121074333AActive Publication Date: 2025-12-05北京渲光科技有限公司 +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511182712.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-05
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing technologies face challenges in achieving 4K real-time rendering on VR devices due to insufficient computing power and high bandwidth requirements. Traditional super-resolution technologies are not effective in terms of detail preservation, edge smoothness, and overall consistency.

Method used

A lightweight super-resolution model is adopted, combined with a scene-customized model and video coding. The lightweight model is trained in the cloud through a pre-trained variational autoencoder and a diffusion model, and the image reconstruction and rendering are performed locally using a cross-scale Transformer module and a multi-scale discriminator.

Benefits of technology

It achieves high-quality, low-latency 4K real-time rendering on low-performance VR devices, reducing bandwidth pressure and improving image quality and training efficiency, making it suitable for applications such as large-scale cloud simulation and VR tourism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074333A_ABST
    Figure CN121074333A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of artificial intelligence, computer vision and computer graphics, and discloses a VR real-time rendering method and system for low-bit-rate cloud rendering plug flow, and the method comprises the steps: obtaining a rendering frame image; a lightweight super-resolution model is constructed; inputting the rendered frame image into the lightweight super-resolution model to obtain a reconstructed 4K-resolution image; and performing real-time rendering and display on the VR equipment based on the reconstructed 4K-resolution image. According to the invention, high bandwidth pressure caused by direct transmission of the 4K video is avoided, and the hardware limitation that low configuration equipment cannot carry out complex rendering tasks is overcome. Through the combination of a scene customization model and video coding, high-quality and low-delay visual experience under limited resources is realized, and the method is suitable for various application scenes with relatively high requirements on image quality and real-time performance, such as large-scale cloud simulation, VR tourism and intelligent power grid monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of artificial intelligence, computer vision and computer graphics, and particularly relates to a method and system for VR real-time rendering of low-code-rate cloud rendering push stream. BACKGROUND

[0002] Implementing 4K real-time rendering on VR devices has revolutionary significance for large-scale real-time simulation applications. For example, in the field of smart grid fault detection, 4K real-time rendering can help engineers observe the power grid structure and its operating state more clearly, quickly locate and solve potential problems, and ensure the stable operation of the power system. In the field of VR tourism, tourists can wear VR devices to virtually tour famous historical sites around the world and enjoy an unprecedented immersive experience.

[0003] However, the existing technology faces many challenges in implementing 4K real-time rendering on VR devices. First, directly performing real-time rendering of 4K pictures on VR devices requires extremely high computing power, which most VR devices cannot meet due to hardware limitations. For this reason, cloud rendering becomes a viable solution, but transmitting 4K pictures from the cloud to VR devices requires huge bandwidth, which not only increases costs but also may cause latency problems, affecting user experience. Another solution is to first transmit lower resolution pictures (such as 1080p) and then enhance them to 4K resolution on the VR device. Currently, this super-resolution technology mainly relies on generative adversarial networks (GANs), although GANs have made significant progress in image generation, they still have deficiencies in processing speed, efficiency, and generated quality. In particular, existing methods often fail to achieve ideal results in terms of detail preservation, edge smoothness, and overall consistency, affecting the final visual experience.

[0004] To solve the above limitations, the present application proposes a method and system for VR real-time rendering of low-code-rate cloud push stream. SUMMARY

[0005] To solve the problems in the prior art, the present application provides a method and system for VR real-time rendering of low-code-rate cloud rendering push stream, which not only avoids the high bandwidth pressure caused by direct transmission of 4K video, but also overcomes the hardware limitations of low-config devices that cannot perform complex rendering tasks. Through the combination of scene customization models and video encoding, high-quality and low-latency visual experience is achieved under limited resources, making it suitable for various application scenarios that require high quality and real-time performance, such as large-scale cloud simulation, VR tourism, and smart grid monitoring.

[0006] To achieve the above purpose, the present application provides the following solutions:

[0007] A method for VR real-time rendering of low-code-rate cloud rendering streaming, the method comprising:

[0008] obtaining a rendered frame image;

[0009] constructing a lightweight super-resolution model;

[0010] inputting the rendered frame image into the lightweight super-resolution model to obtain a reconstructed 4K resolution image;

[0011] based on the reconstructed 4K resolution image, performing real-time rendering and display on a VR device.

[0012] Preferably, the lightweight super-resolution model comprises:

[0013] a U-Net generator G, wherein N CSTB =4 CSTB modules; a discriminator network D gan ; a diffusion model and a pre-trained variational autoencoder VAE, wherein the VAE is divided into an encoder E vae and a decoder D vae ;

[0014] The generator G replaces the standard convolutional layer in the U-Net with a cross-scale Transformer module, i.e., CSTB, specifically replacing all 3x3 convolutional layers in the U-Net encoder and decoder, retaining down-sampling and up-sampling operations, adding CSTB at the jump connection, enhancing cross-scale information transmission, and the structure is:

[0015] CSTB(z i ) = DeformConv(LayerNorm(z i-1 + MultiHeadAttn(z i-1 , z i-1 , z i-1 ))

[0016] wherein MultiHeadAttn represents multi-head attention, DeformConv represents deformable convolution, LayerNorm represents normalization layer, and z i-1 represents the output feature map of the i-1th CSTB;

[0017] The discriminator network D gan is designed to have a multi-scale discriminator structure: multi-scale discriminator {D1, D2, D3}, each discriminator uses the same backbone network PatchGAN, independently trains parameters, and the last layer outputs probability;

[0018] The diffusion model comprises: a forward process and a reverse process;

[0019] where the forward process: create a noisy version z of z0 at random time step t t , the formula is:

[0020]

[0021] where x0 is the latent representation of the original image, i.e., without added noise, x t is the noisy latent representation at time step t, is the state of the original image after adding t-step noise, q(x t | x t-1 ) is the conditional probability distribution, representing the process of generating x t-1 from x t , following a Gaussian distribution, a t = 1 - b t is the single-step noise retention coefficient, representing the proportion of the original information retained at the t-th step, is the cumulative noise retention coefficient, i.e., represents the conditional distribution of generating x t directly from x0, b t is a noise scheduler used to control the noise level at time step t, as t gradually increases, i.e., t→T, then x t will gradually approach pure noise N(0, I), and q(x t | x t-1 ) is re-parameterized, let a t = 1 - b t , get the formula q(x t | x0), I is the identity matrix;

[0022] The reverse process: z0 and are respectively subjected to an additional diffusion step to the same time step s, generating z s and Then use the pre-trained decoder D vae to decode the results back to the pixel space to obtain images x s and Images x s and are input into the discriminator D gan for evaluation:

[0023]

[0024] where I is the identity matrix, is the cumulative noise coefficient used to control the noise level at time step s, z s is the noisy version of the real latent representation z0 at time step s, is the noise level of the generator output at time step s.

[0025] Preferably, a meta-path controller is introduced to dynamically skip redundant CSTB computation units based on the local complexity of the feature map, thereby reducing computational overhead while maintaining generation quality.

[0026] splicing feature z init Perform local complexity awareness and output the complexity graph Ω, represented as:

[0027]

[0028] in, It is z init In the eigenvector at position (i,j), ||·||² is the L2 norm, γ is the balance factor, and Entropy(·) is the channel distribution entropy of the local 3×3 window. p k It is the probability distribution of the local window histogram in 256 bins;

[0029] Predicting the skip probability using a lightweight convolutional network is expressed as:

[0030] M skip =σ(Conv 3×3 (ReLU(Conv 1×1 (Ω))))

[0031] Where σ is the Sigmoid function, outputting a probability graph in the range [0,1]. Conv 1×1 It is a 1×1 convolutional layer, the purpose of which is to reduce the dimensionality to 8 channels;

[0032] During the reasoning phase, hard coding is used, that is... Then, at position (i,j), the current CSTB calculation is skipped, and the previous level feature is directly reused. During the training phase, a soft mask is used, that is, Gumbel-Softmax is used instead of hard coding, as shown below:

[0033]

[0034] Where G and G' are injected random noises, G and G' follow Gumbel(0,1), and τ is a temperature coefficient representing the smoothness control parameter. When τ→0 + The output approaches a hard decision, i.e., 0 or 1; when τ→∝, the output approaches a uniform distribution.

[0035] Enhance feature consistency through global dependency modeling;

[0036] Dynamic learning bias Δp∈R H×W×2N Where N is the kernel size, and the convolution sampling position is adjusted:

[0037]

[0038] where z is the input feature map, representing the input feature tensor to be convolved, p is the target position coordinate, representing the current calculated pixel position on the output feature map, p n is the predefined sampling offset, representing the fixed relative offset of the nth sampling point in the convolution kernel, Δp n is the dynamic learning offset, representing the adaptive position adjustment amount of the nth sampling point, w n is the convolution weight, representing the convolution kernel weight of the nth sampling point, and N is the total number of sampling points, representing the number of sampling points of the convolution kernel.

[0039] Preferably, the method for inputting the rendered frame image into the lightweight super-resolution model to obtain the reconstructed 4K resolution image comprises:

[0040] using a pre-trained variational autoencoder to encode the rendered frame image into a latent space, and gradually adding noise through a diffusion process, so that the generator learns to recover the latent representation of the high-resolution image from the noise, and then decodes the latent representation back to the pixel space through the decoder to obtain the super-resolution image.

[0041] The application also provides a system for low-code-rate cloud rendering streaming VR real-time rendering, which is used to implement the above method, and comprises an acquisition module, a construction module, a reconstruction module and a rendering module.

[0042] The acquisition module is used to acquire a rendered frame image.

[0043] The construction module is used to construct a lightweight super-resolution model.

[0044] The reconstruction module is used to input the rendered frame image into the lightweight super-resolution model to obtain a reconstructed 4K resolution image.

[0045] The rendering module is used to perform real-time rendering and display on a VR device based on the reconstructed 4K resolution image.

[0046] Preferably, the lightweight super-resolution model comprises:

[0047] a U-Net generator G, wherein the U-Net generator G is composed of N CSTB = 4 CSTB modules; a discriminator network D gan ; a diffusion model and a pre-trained variational autoencoder VAE, wherein the VAE is divided into an encoder E vae and a decoder D vae ;

[0048] The generator G replaces the standard convolutional layers in the U-Net with cross-scale Transformer blocks, i.e., CSTBs, specifically, replaces all 3x3 convolutional layers in the U-Net encoder and decoder, retains the down-sampling and up-sampling operations, adds CSTBs at the skip connections, enhances cross-scale information transmission, and the structure is as follows:

[0049] CSTB(z i )=DeformConv(LayerNorm(z i-1 +MultiHeadAttn(z i-1 ,z i-1 ,z i-1 ))

[0050] where MultiHeadAttn represents multi-head attention, DeformConv represents deformable convolution, LayerNorm represents a normalization layer, z i-1 represents the output feature map of the i-1th CSTB;

[0051] The discriminator network D gan The structure of the multi-scale discriminator is designed: multi-scale discriminator {D1, D2, D3}, each discriminator uses the same backbone network PatchGAN, independently trains the parameters, and the last layer outputs the probability;

[0052] The diffusion model includes: forward process and reverse process;

[0053] where the forward process: create a noise version z t of z0 at a random time step t, the formula is:

[0054]

[0055] where x0 is the latent representation of the original image, i.e., without adding noise, x t represents the noise latent representation at time step t, which is the state of the original image after adding t-step noise, q(x t |x t-1 ) is a conditional probability distribution, which represents the process of generating x t-1 from x t , which is subject to a Gaussian distribution, and t = 1-β t is a single-step noise retention coefficient, which represents the proportion of original information retained at the tth step, is the cumulative noise retention coefficient, i.e. q(x t |x0) represents the conditional distribution of generating x t from x0 directly, and tIt is a noise scheduler used to control the noise level at time step t. As t gradually increases (i.e., t→T), then x... t It will gradually approach pure noise N(0,I) for q(x) t |x t-1 Reparameterize α t =1-β t , The formula q(x) is obtained t |x0), where I is the identity matrix;

[0056] Reverse process: connect z0 and Perform additional diffusion steps to the same time step s to generate z. s and Then use the pre-trained decoder D vae Decode the result back into pixel space to obtain the image x. s and Image x s and Input to discriminator D gan During the evaluation:

[0057]

[0058] Where I is the identity matrix, It is the cumulative noise figure, used to control the noise level at time step s, z s It is a noisy version of the true latent representation z0 at time step s. It is the generator output. Noise level at time step s.

[0059] Preferably, a meta-path controller is introduced to dynamically skip redundant CSTB computation units based on the local complexity of the feature map, thereby reducing computational overhead while maintaining generation quality.

[0060] splicing feature z init Perform local complexity awareness and output the complexity graph Ω, represented as:

[0061]

[0062] in, It is z init In the eigenvector at position (i,j), ||·||² is the L2 norm, γ is the balance factor, and Entropy(·) is the channel distribution entropy of the local 3×3 window. p k It is the probability distribution of the local window histogram in 256 bins;

[0063] Predicting the skip probability using a lightweight convolutional network is expressed as:

[0064] M skip = sigma(Conv 3×3 (ReLU(Conv 1×1 (Ω))))

[0065] where sigma is the Sigmoid function, outputting [0,1] probability map, Conv 1×1 is a 1x1 convolution layer, the purpose is to reduce dimension to 8 channels;

[0066] In the inference stage, hard coding is adopted, that is then skip the current CSTB calculation at position (i,j) and directly reuse the last level feature, in the training stage, soft mask is adopted, that is, Gumbel-Softmax is used instead of hard coding, which is expressed as:

[0067]

[0068] where G and G' are injected random noise, G and G' obey Gumbel(0,1), tau is a temperature coefficient, which is a smoothing degree control parameter, when tau→0 + , the output tends to hard decision, that is, 0 or 1; when tau→∞, the output tends to uniform distribution, that is

[0069] Enhance feature consistency by modeling global dependence;

[0070] Dynamically learn bias Delta p belongs to R H×W×2N , where N is the size of the convolution kernel, adjust the convolution sampling position:

[0071]

[0072] where z is the input feature map, representing the input feature tensor to be convolved, p is the target position coordinate, representing the current pixel position calculated on the output feature map, p n is the pre-defined sampling offset, representing the fixed relative offset of the nth sampling point in the convolution kernel, Delta p n is the dynamically learned offset, representing the adaptive position adjustment amount of the nth sampling point, w n is the convolution weight, representing the convolution kernel weight of the nth sampling point, and N is the total number of sampling points, representing the number of sampling points of the convolution kernel.

[0073] Preferably, the rendered frame image is input into the lightweight super-resolution model, and the process of obtaining the reconstructed 4K resolution image includes:

[0074] The pre-trained variational autoencoder is used to encode the rendered frame image into a latent space, and noise is gradually added through a diffusion process, so that the generator learns to recover the latent representation of a high-resolution image from the noise, and the decoder decodes the latent representation back to the pixel space to obtain a super-resolution image.

[0075] The application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method according to any one of the application when executing the program.

[0076] The application further provides a computer readable storage medium, which stores a computer program, and the computer program, when executed, implements the method according to any one of the application.

[0077] Compared with the prior art, the application has the following beneficial effects:

[0078] The method provides a low-code-rate cloud push stream VR real-time rendering method and system. The method uses a pre-trained variational autoencoder (VAE) to encode an image into a latent space, and gradually adds noise through a diffusion process, so that the generator learns how to recover the latent representation of a high-resolution image from the noise, and finally decodes it back to the pixel space through the decoder to obtain a super-resolution image. In order to further optimize the training effect, a dynamic adjustment time step strategy is adopted, and the training progress between the generator and the discriminator is balanced according to the accuracy of the discriminator. In addition, the discriminator adopts a multi-scale structure design, which can accurately distinguish images of different scales, and improves the discrimination effect through spectral normalization and multi-scale weight fusion technology. Compared with the traditional method, the method solves the problem that the U-Net structure generator is difficult to capture global context information, enhances the texture consistency, and improves the generation ability of multi-scale edges and structural details. At the same time, the design of the multi-scale discriminator improves the discrimination ability of high-frequency and low-frequency structural features, and enhances the stability of the training. Experimental results show that the method not only can effectively improve the image quality indicators (such as PSNR and SSIM), but also can significantly reduce the FID value, showing higher training and inference efficiency and stronger robustness. BRIEF DESCRIPTION OF DRAWINGS

[0079] In order to more clearly illustrate the technical solutions of the application, the following briefly introduces the drawings needed to be used in the embodiments. Obviously, the drawings in the following description only constitute some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0080] Figure 1 The figure is a lightweight super-resolution model architecture diagram of the embodiment of the application.

[0081] Figure 2 This is a schematic diagram of the meta-path controller according to an embodiment of the present invention;

[0082] Figure 3 This is a schematic diagram of the reasoning process in an embodiment of the present invention;

[0083] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention.

[0085] 1010, Processor; 1020, Memory; 1030, Input / Output Interface; 1040, Communication Interface; 1050, Bus. Detailed Implementation

[0086] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0087] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0088] Example 1

[0089] This embodiment provides a method for real-time VR rendering with low bitrate cloud rendering streaming, the method comprising:

[0090] Obtain the rendered frame image;

[0091] Constructing a lightweight super-resolution model;

[0092] Input the rendered frame image into the lightweight super-resolution model to obtain the reconstructed 4K resolution image;

[0093] Based on the reconstructed 4K resolution image, it is rendered and displayed in real time on VR devices.

[0094] In this embodiment, as Figure 1 As shown, (1) the lightweight super-resolution model architecture includes a U-Net generator G (composed of N... CSTB =Composed of 4 CSTB modules), and a discriminator network D gan A diffusion model and a pre-trained variational autoencoder (VAE), wherein the VAE is divided into encoder E vae and decoder D vae .

[0095] (2) Training data processing: For each pair of high-resolution image x0and low-resolution image x low First, the encoder E of VAE is used to encode them into latent representations z0and z vae low This latent encoding enables efficient processing in a lower-dimensional space. In this invention, x represents the picture, and z represents the encoded feature map

[0096] (3) In the overall architecture diagram, the "high-quality output picture" is the output of the model streaming process, and the "true or false order" is the output of the model training process.

[0097] (4) The core goal of this technology is to achieve real-time rendering and display of 4K ultra-high definition quality on low-performance VR devices. Through cloud training of lightweight models, local inference combined with video encoding transmission, the performance bottleneck caused by hardware limitations or high bandwidth requirements in traditional solutions is solved. This technology is divided into two main stages: model training stage (cloud) and model inference stage (local + cloud collaboration).

[0098] 1) In the model training stage, the system independently models and trains each virtual scene based on high-performance cloud servers. Specifically, 1080P and 4K resolution pictures under the same scene are simultaneously rendered on the cloud as input and target output. Through this data, the image super-resolution network is trained, enabling the model to convert low-resolution images into high-quality 4K images. To further optimize the computational efficiency of the model and meet the operating requirements of resource-constrained devices, we use the "overfitting training" strategy, which involves training a highly lightweight model for each specific virtual scene. This approach sacrifices the model's generalization ability, but since each virtual scene has relatively fixed content, it can significantly compress the model parameter size while ensuring reconstruction quality, thereby improving inference efficiency. After training, all models for different virtual scenes are stored in the cloud model library for subsequent deployment.

[0099] ​2) In the model inference stage, the whole process realizes the cooperation of cloud and local device. The cloud server renders the 1080P picture in real time according to the virtual scene where the current user is located, and uses the H.265 high-efficiency coding mode for video compression to significantly reduce the transmission code rate and bandwidth occupation. The compressed video stream is transmitted to the local low-performance VR device through the network (such as WebRTC technology) for decoding and recovery to a series of frame images. Subsequently, the decoded frames will be automatically matched and input into the corresponding lightweight super-resolution model according to their virtual scene type. These models are downloaded from the cloud model library in advance and cached on the local device, and are specially used for image enhancement tasks of specific scenes. The output of the model processing is the reconstructed 4K resolution image, which is finally rendered and displayed in real time by the local VR device.

[0100] The advantage of this architecture is that it avoids the high bandwidth pressure brought by directly transmitting 4K video, and overcomes the hardware limitation of low-config devices that cannot perform complex rendering tasks. Through the combination of scene customized models and video coding, high-quality and low-latency visual experience is realized under limited resources, which is suitable for various application scenarios such as large-scale cloud simulation, VR tourism, smart grid monitoring, etc. that have high requirements on picture quality and real-time performance.

[0101] In the present embodiment, the diffusion process is:

[0102] (1) Forward process: create a noisy version z of z0 at random time step t t , the formula is as follows:

[0103]

[0104] where x0 is the latent representation of the original image (without added noise). x t represents the noisy latent representation at time step t, which is the state of the original image after adding t-step noise. q(x t |x t-1 ) is the conditional probability distribution, which represents the process of generating x t-1 from x t , subject to Gaussian distribution. α t = 1-β t is a single-step noise retention coefficient, which represents the proportion of original information retained at the t-th step. is the cumulative noise retention coefficient, i.e. q(x t |x0) represents the conditional distribution of generating x t from x0 directly, avoiding step-by-step calculation. β t is a noise scheduler that controls the noise level at time step t. As t gradually increases, i.e. t→T, x t will gradually approach pure noise N(0,I).t |x t-1 Reparameterize α t =1-β t , The formula q(x) is obtained t |x0). I is the identity matrix.

[0105] 1) Cosine noise scheduling design: If β t Linear noise scheduling, with its inconsistent rates of change at both ends of the time step, can lead to abrupt changes in high-frequency noise, affecting training stability. Therefore, linear scheduling is replaced with cosine noise scheduling because the rate of change of the cosine function along the time axis is symmetrically distributed, making β... t Changes are slow in the initial stage (t≈0) and the final stage (t≈T), and slow in the intermediate stage. The changes are relatively rapid. This scheduling method avoids the abrupt changes of linear scheduling, smooths out excessive noise, and reduces gradient oscillations during training; it accelerates noise addition in the intermediate stages, more efficiently covering the data distribution. The formula for cosine noise scheduling is:

[0106]

[0107] Where, α min and α max These are the minimum and maximum coefficients for noise scheduling, used to control the noise range (default α). min =0.001, α max =0.02). t is the current time step, t∈{1,2,…,T}. T is the total number of diffusion steps (default T=1000).

[0108] 2) Adaptive Step Size Selection Strategy: During the training phase, the model is trained using a complete T-step diffusion process to ensure that the model learns the noise distribution across the entire step size. During the inference phase, the model is selected based on the low-resolution latent representation z. low The complexity of dynamically adjusting the effective number of steps T in the reverse process. eff Only execute the first T eff The reverse process reduces redundant calculations for simple samples.

[0109]

[0110] Among them, ||z low ||2 is the low-resolution latent representation z low The L2 norm measures the complexity (a larger value indicates richer details). γ is a smoothing factor (default γ = 0.1), which prevents the denominator from being zero and controls the adjustment magnitude.

[0111] (2) Generator input: input the triplet (z) t ,t,z low) into the generator G, whose goal is to predict z0such that it is as close to the real latent representation z0as possible.

[0112] (3) Reverse process of the standard diffusion model: convert pure noise x t ~ N(0, I) to clean data x0.

[0113] 1) Formula

[0114]

[0115] 2) Model learning: the goal is to come up with a model p θ (x t-1 | x t ) whose parameters θ can minimize the KL divergence between the tractable posterior distributions at all time steps.

[0116]

[0117] where θ * represents the optimal parameters of the inverse process of the diffusion model, i.e., the generator parameters obtained by minimizing the sum of KL divergences. This can be simplified as the L2 norm between the real noise ∈ t and the predicted noise at the noise state x t .

[0118] 3) Sample generation: for the trained model p θ (x t-1 | x t ), new samples can be generated by starting from x T ~ N(0, I) and iteratively denoising using the model, i.e.,

[0119]

[0120] where x T represents the latent variable at the final time step T of the diffusion process. q(x T ) represents the marginal probability distribution of x T defined by the forward process of the diffusion model.

[0121] (4) Reverse process of the proposed method: to further optimize the prediction results, z0and are each subjected to an additional diffusion step to the same time step s, generating z s and The results are then decoded back to the pixel space using the pre-trained decoder D vae to obtain images x s and These images will be input into the discriminator D gan for evaluation.

[0122]

[0123] where I is the identity matrix, is the accumulated noise coefficient, used to control the noise level at time step s. s is the noisy version of the real latent representation z0 at time step s, is the generator output is the noise level at time step s.

[0124] (5) Dynamic adjustment of time step s: use the Exponential Moving Average (EMA) mechanism to monitor the accuracy of the discriminator, and adjust the time step s accordingly, so as to ensure that the discriminator is neither too strong nor too weak, thus maintaining the balance between the generator and the discriminator. The formula for dynamically adjusting the time step is as follows:

[0125]

[0126] where, is the EMA accuracy at the i-th training iteration, acc batch is the discriminator accuracy of the current batch, λ ema is the EMA weight (λ ema = 0.05), T is the maximum diffusion time step, and s is the dynamically adjusted time step.

[0127] It should be noted that: the forward noise adding process is basically the same as in the standard diffusion model (with only two small innovations), and the noise is added to T. The reverse denoising process is different from the standard diffusion model, and instead of T time steps, it only takes s time steps (this s can be dynamically adjusted). The subsequent process is to input z s and into the decoder, output x s and Then randomly splice them together and input them into the discriminator to let the discriminator judge which picture is the "high-quality input image" z s , the decoded picture (this is true); which picture is the "low-quality input image" z , the decoded picture (this is false). Finally, output the true-false order, with four results: true-false, false-true, true-true, and false-false.

[0128] In this embodiment, the generator G:

[0129] (1) Generator structure: If the generator G adopts a simple U-Net structure, it has limitations, on the one hand, traditional convolution is difficult to capture the global context, limited by the local receptive field, resulting in insufficient texture consistency; on the other hand, the generation ability of multi-scale edges and structural details is limited. Therefore, the standard convolution layer in the U-Net is replaced by the cross-scale Transformer module (referred to as CSTB), specifically replacing all 3x3 convolution layers in the U-Net encoder and decoder, retaining the downsampling and upsampling operations, adding CSTB at the jump connection, and enhancing the cross-scale information transmission. The structure is as follows:

[0130] CSTB(z i )=DeformConv(LayerNorm(z i-1 +MultiHeadAttn(z i-1 ,z i-1 ,z i-1 ))

[0131] Where, MultiHeadAttn represents multi-head attention, DeformConv represents deformable convolution, LayerNorm represents normalization layer, and z i-1 represents the output feature map of the i-1th CSTB.

[0132] 1) Input data of the generator: input noise latent state z t , time step t and latent representation z low of low resolution image. Input one, the current latent state z t is the noise latent representation processed by the forward diffusion process, which contains the information of the latent representation z0 encoded from the original high-resolution image x0 after noise processing. Input two, the time step t represents the current stage in the diffusion process, which provides the generator with information about the denoising process, helping the generator understand the denoising level to be achieved. Input three, the latent representation z low of the low resolution image is obtained by encoding the low resolution image x vae through the pre-trained VAE encoder E low . The latent representation z low of the low resolution image provides the generator with the structure and semantic information of the image, guiding the generator to generate a high-resolution image consistent with the low-resolution image.

[0133] 2) Feature fusion of input data: z t ∈R H×W×C and z low ∈R H×W×C are spliced through the channel dimension to form the initial feature z t-low ∈R H×W×2CWhere H×W×C represents the height, width, and number of channels of the feature map. Then, a 1×1 convolution is used to reduce the concatenated features to the original number of channels, yielding z. init ∈R H×W×C The formula is expressed as:

[0134] z init =Conv 1×1 (z t ⊕z low )

[0135] Here, ⊕ represents the concatenation of channel dimensions. init As the first input to CSTB, after N CSTB After N CSTB iterations (for example, a generator consists of N...), CSTB =Composed of 4 CSTB modules), to obtain the final output

[0136] 3) Meta-Path Controller (MPC): Traditional dynamic path computation only depends on the global L2 norm ||z low ||2 makes a decision (i.e., skips the CSTB module), but the complexity distribution is uneven across different image regions. For example, in a VR scene, a person's face requires detailed calculation, while a solid-color background can be simplified. MPC is a spatially adaptive computation scheduling mechanism. Its core idea is to dynamically skip redundant CSTB computation units based on the local complexity of the feature map, significantly reducing computational overhead while maintaining generation quality. The flowchart is as follows. Figure 2 As shown:

[0137] a) Local complexity analysis: For the splicing feature z init Perform local complexity awareness and output the complexity graph Ω.

[0138]

[0139] in, It is z init In the eigenvector at position (i,j), ||·||² is the L2 norm, and γ is the balance factor (default γ = 0.5). Entropy(·) is the channel distribution entropy of the local 3×3 window. p k It is the probability distribution of the local window histogram in 256 bins.

[0140] b) Dynamic mask generation: Predicting skip probability using a lightweight convolutional network.

[0141] M skip =σ(Conv 3×3 (ReLU(Conv 1×1 (Ω))))

[0142] where σ is the Sigmoid function, outputting [0, 1] probability map. Conv 1×1 is a 1x1 convolutional layer, aiming to reduce dimension to 8 channels.

[0143] c) Gumbel-Softmax mechanism: In the inference stage, hard coding is adopted, i.e. then skip the current CSTB calculation at position (i, j) and directly reuse the features of the previous level. In the training stage, soft masking is adopted, i.e. Gumbel-Softmax is used instead of hard coding.

[0144]

[0145] where G and G' are random noise injected, G and G' follow Gumbel(0, 1). τ is the temperature coefficient, representing the smoothness control parameter. When τ→0 + , the output tends to hard decision (0 or 1); when τ→∞, the output tends to uniform distribution

[0146] 1) Multi-Head Attention MultiHeadAttn: Enhance feature consistency by modeling global dependencies. Input feature map z i-1 ∈R H×W×C (the first input is z init ), calculate global dependencies through Query, Key, Value projection:

[0147]

[0148] where Q is the query matrix, which will be compared with the key matrix to determine which information is important. K is the key matrix, and the comparison result with the query matrix is used to calculate the attention weight, so as to determine which information is relevant. V is the value matrix, which will be weighted according to the calculated attention weight to generate the final output. d k is the dimension of the key vector, which is used to scale the attention score to prevent the problem of gradient disappearance of the Softmax function caused by too large inner product, so as to ensure the stability of the model. In addition, the multi-head mechanism is divided into h parallel heads (default h=8) to enhance multi-scale feature fusion.

[0149] 2) Deformable Convolution DeformConv: Dynamically adjust the receptive field to capture multi-scale details and improve edge sharpness. Dynamically learn the bias Δp∈R H×W×2N , where N is the size of the convolution kernel, adjust the convolution sampling position:

[0150]

[0151] where z is the input feature map, representing the input feature tensor to be convolved (usually from the output of the previous layer). p is the target position coordinate, representing the current calculated pixel position on the output feature map. p n is the predefined sampling offset, representing the fixed relative offset of the nth sampling point in the convolution kernel. Δp n is the dynamic learning offset, representing the adaptive position adjustment amount of the nth sampling point. w n is the convolution weight, representing the convolution kernel weight of the nth sampling point. N is the total number of sampling points, representing the number of sampling points in the convolution kernel.

[0152] 3) Normalization layer LayerNorm: Normalizing the attention output to stabilize the training process.

[0153] (2) Advantages of CSTB:

[0154] 1) Global context modeling: Multi-head self-attention mechanism captures long-range dependencies, solving the texture breakage problem caused by local convolution in the original method.

[0155] 2) Multi-scale detail enhancement: Deformable convolution dynamically adapts to edges and structures, with PSNR improved by about 0.8dB and SSIM improved by 5%.

[0156] 3) Training and inference efficiency (combined with inference process optimization, such as dynamic calculation path and cache attention matrix): Mixed precision training reduces GPU memory usage by 40%, and dynamic calculation path makes the inference speed reach 25FPS (1080Ti).

[0157] 4) Robustness improvement (combined with cross-scale consistency loss function): Cross-scale consistency loss suppresses multi-scale generation inconsistency, reducing FID by 12.3%.

[0158] (3) Role and training target of generator

[0159] 1) During training, the goal of the generator G is to learn how to recover the clean latent representation z0 from the noise latent state z t step by step. Through adversarial training with the discriminator D gan , the generator continuously optimizes its prediction ability to generate more realistic and high-quality high-resolution images.

[0160] 2) The training target of the generator is to minimize the difference between the generated image and the real image in the latent space, while deceiving the discriminator to make it unable to distinguish between the generated image and the real image. This is achieved through the generator's loss function, which is a combination of content loss and adversarial loss.

[0161] In this embodiment, the discriminator D ganIf the discriminator uses a single scale convolutional network, on the one hand, it only focuses on the original resolution features, and the discrimination ability for high-frequency textures (such as hair, texture details) and low-frequency structures (such as object contours) is insufficient; on the other hand, the traditional weight normalization method (such as BatchNorm) may cause gradient explosion or mode collapse, making the training unstable. To this end, the present application designs a multi-scale discriminator:

[0162] (1) The structure of the multi-scale discriminator: the multi-scale discriminator {D1, D2, D3}, each discriminator uses the same backbone network (such as PatchGAN), but trains the parameters independently, and the last layer outputs the probability.

[0163] (2) The input and output of the multi-scale discriminator: because the denoising process only performs s time steps, the discriminator input is x s and instead of the real high-resolution image x0 and the generated super-resolution image

[0164] 1) D1 input Output true and false probability. Wherein is the random arrangement of the discriminator input, the purpose is to improve the robustness of the adversarial training, the input of the discriminator is composed of the random arrangement of the real image and the generated image in the channel dimension, and the discriminator needs to predict the correct order of this arrangement.

[0165]

[0166] Where, x s is the decoding image of z s at time step s, is the decoding image of z at time step s, represents the concatenation operation in the channel dimension, represents the concatenation operation of random arrangement.

[0167] 2) D2 input Down-sampled image, that is, Output true and false probability.

[0168] 3) D3 input Down-sampled image, that is, Output true and false probability.

[0169] (3) Spectral normalization: apply spectral normalization after each layer of convolution in the discriminator, update the weight matrix before each forward propagation, the formula is

[0170]

[0171] where σ(W) is the spectral norm of the weight matrix (i.e. the largest singular value), computed by the power iteration approximation algorithm, W is the weight matrix of the convolutional layer, W SN is the normalized weight to ensure the network satisfies Lipschitz continuity (i.e. constraint the gradient magnitude).

[0172] (4) The output of each discriminator is obtained by global average pooling (Global Average Pooling) to get a scalar value, representing the "real score" of the image at this scale.

[0173] (5) Multi-scale weight fusion: combine the discriminant results of different scales by weighted average, to enhance robustness, introduce scale confidence factor a k , dynamically adjust the weight.

[0174]

[0175] where w k is the weight coefficient of the scale, satisfying (default w1=0.5, w2=0.3, w3=0.2, reflecting the emphasis on high-frequency details). a k is the scale confidence factor, if the output probability of a certain scale discriminator is high (D k →1), its confidence a k increases, otherwise it decreases. Set the threshold (such as ), if then the image is real, otherwise the image is generated (i.e. not real). In the training stage, the weight is the loss gradient of the kth discriminator, τ is the temperature coefficient (τ=0.1), used to control the weight adjustment amplitude.

[0176] In this embodiment, the loss function is:

[0177] (1) Cross-scale consistency loss:

[0178]

[0179] where Downsample k represents the kth down-sampling (the ratio is ), to constrain the alignment of multi-scale features.

[0180] (2) Generator loss function: the loss function of the generator is a weighted combination of multiple loss functions, by balancing these loss terms, the generator can improve the perceptual quality and fidelity while maintaining structural accuracy.

[0181] 1) Content loss L mse : Content loss measures the structural difference between generated and real images in latent space by mean squared error (MSE).

[0182]

[0183] where z0 is the latent representation of the real high-resolution image, is the latent representation of the high-resolution image output by the generator.

[0184] 2) Multi-scale adversarial loss L adv : Adversarial loss encourages the generator to generate images that can deceive the discriminator, making it unable to distinguish between generated and real images.

[0185]

[0186] where w k is the weight of each scale, D k (·) is the output probability map mean of the kth discriminator. E is the expected value.

[0187] 3) Perceptual loss L perc : High-level features are extracted using a pre-trained VGG19 network, and the distance between generated and real images in feature space is calculated. Perceptual loss can be constrained by high-level features, making the generated image more consistent with human visual perception and reducing the problem of excessive smoothing caused by MSE loss.

[0188]

[0189] where φ is the feature extractor of the Conv 4×4 layer in the VGG19 network, and the Conv 4×4 layer balances the representation ability of low-level texture (such as edges) and high-level semantics (such as object structure).

[0190] 4) Gradient consistency loss L grad : Calculate the difference between generated and real images in gradient space, forcing the generated image to be consistent with the real image in edge distribution, and suppressing blurred artifacts.

[0191]

[0192] where, represents the image gradient, calculated by the Sobel operator, i.e. |·||1 represents the L1 norm, which is more sensitive to edge differences and enhances sparsity constraints.

[0193] 5) Loss function of meta-path controller

[0194] LMPC = λ consist · L consist + (1 - λ consist ) L balance

[0195] where L consist is the decision consistency loss, which is used to prevent frequent switching of the computation state and maintain the spatio-temporal continuity. is the mask at position (i, j). L consist is detected by the Laplacian operator, which penalizes isolated point decisions (e.g., single-pixel skipping). L balance is the computation load balancing loss, which is used to ensure balanced computation load across regions. where μ target is the target skipping rate, with a default μ target = 0.4 and dynamic adjustment, i.e. where ||z low || 2,max is the maximum value in the L2 norm of all z low . λ consist is the continuity weight (e.g., λ consist = 0.4).

[0196] 6) Generator loss function L G

[0197] L G = L mse + λ adv · L adv + λ CS · L CS + λ perc · L perc + λ grad · L grad + λ MPC · L MPC

[0198] where λ adv is the weight of the adversarial loss (e.g., set to λ adv = 1 × 10 -3 ), λ CS is the weight of the consistency loss (e.g., λ CS = 0.1), λ perc is the weight of the perception loss (e.g., λ perc = 0.01), λ grad is the weight of the gradient consistency loss (e.g., λ grad = 0.05), and λ MPC is the weight of the MCP loss (e.g., λ MPC = 0.05), which is used to balance the influence of each loss term.

[0199] (3) Parameter update: the weights of the model are updated by alternating the use of discriminator loss and generator loss. This way of alternating training helps to maintain the balance between the generator and the discriminator during the training process, thereby improving the stability of the model and the quality of the generated images.

[0200] In this embodiment, the inference process (training process and inference process) is as shown in Figure 3 :

[0201] (1) Encode low-resolution image: encode the low-resolution input image x low into the latent space to obtain the latent representation z low .

[0202] (2) Diffusion inverse process

[0203] 1) Start from pure Gaussian noise, perform diffusion inverse process in latent space.

[0204] 2) Use the generator G to gradually reduce the noise and generate the latent representation of the high-quality high-resolution image

[0205] (3) Inference optimization

[0206] 1) Dynamic path calculation: according to the complexity of the input z low (by calculating ||z low ||2), adaptively skip part of the CSTB module calculation:

[0207]

[0208] where θ is the threshold value (default θ = 0.5), which can be dynamically adjusted according to the validation set.

[0209] 2) Cache attention matrix: cache the attention weights of fixed time step t, reduce repeated calculation, and improve the inference speed by 30%.

[0210] Decode latent representation: decode the final generated latent representation through the pre-trained decoder D vae into the pixel space to obtain the high-resolution output image

[0211] In this embodiment, 1, the generator level

[0212] (1) Global texture consistency: introduce a cross-scale Transformer module to replace traditional convolution, and the multi-head self-attention mechanism captures long-distance dependencies from a global perspective, completely solving the texture breakage problem caused by local convolution, and generating image texture that is coherent and natural.

[0213] (2) Multi-scale detail presentation: Deformable convolution dynamically adjusts the receptive field according to image content, accurately capturing multi-scale details from fine hair to macro outline, with a significant improvement in edge sharpness. Compared with conventional methods, PSNR is improved by about 0.8dB, SSIM is improved by 5%, and visual details are more abundant.

[0214] (3) Training and inference efficiency: Mixed precision training strategy reduces GPU memory usage by 40%, and dynamic calculation path intelligently skips redundant calculation nodes according to input complexity. The inference speed on 1080Ti GPU can reach 25FPS, achieving a perfect balance between efficiency and quality.

[0215] 2. Discriminator level

[0216] (1) Multi-scale feature discrimination: A multi-scale discriminator matrix is constructed to examine image authenticity from original scale to progressively down-sampled multi-dimensional view. It neither overlooks subtle flaws in high-frequency textures nor ignores overall logic in low-frequency structures, with a much higher discrimination granularity than single-scale solutions.

[0217] (2) Training stability guarantee: Replace traditional BatchNorm with spectral normalization to strictly constrain gradient amplitude, eliminating the risk of gradient explosion and mode collapse, and making the training process smooth and stable, significantly enhancing model convergence.

[0218] (3) Dynamic weight fusion: Introduce scale confidence factor and gradient-driven dynamic weight adjustment mechanism, intelligently allocate weights according to different scale reliability and discrimination contribution, making the image authenticity judgment more accurate and reliable compared with fixed weight scheme.

[0219] 3. Overall generation effect

[0220] (1) Ultra-high image quality: Cross-scale consistency loss function strictly constrains multi-scale feature alignment, supplemented by perceptual loss and gradient consistency loss for multi-dimensional optimization. FID index is reduced by 12.3%, and generated images achieve a qualitative leap in structural accuracy and perceptual quality, being both faithful and realistic.

[0221] Robustness is comprehensively improved: From the anti-noise diffusion process of the generator to the multi-scale redundant discrimination of the discriminator, the entire architecture forms a robust closed loop, which can still output high-quality super-resolution results for low-quality and blurred input images, greatly expanding the application scenario boundary.

[0222] Example Two

[0223] The embodiment provides a VR real-time rendering system for low-code-rate cloud rendering streaming, which is used to realize the method of example one, and the system comprises an acquisition module, a construction module, a reconstruction module and a rendering module.

[0224] An acquisition module is configured to acquire a rendered frame image;

[0225] A construction module is configured to construct a lightweight super-resolution model;

[0226] A reconstruction module is configured to input the rendered frame image into the lightweight super-resolution model to obtain a reconstructed 4K resolution image;

[0227] A rendering module is configured to perform real-time rendering and display on a VR device based on the reconstructed 4K resolution image.

[0228] In this embodiment, the lightweight super-resolution model includes:

[0229] a U-Net generator G, wherein N CSTB =4 CSTB modules are used; a discriminator network D gan ; a diffusion model and a pre-trained variational autoencoder VAE, wherein the VAE includes an encoder E vae and a decoder D vae ;

[0230] The generator G replaces the standard convolutional layer in the U-Net with a cross-scale Transformer module, i.e., CSTB. Specifically, all 3x3 convolutional layers in the U-Net encoder and decoder are replaced, the down-sampling and up-sampling operations are retained, and the CSTB is added at the jump connection to enhance cross-scale information transmission. The structure is as follows:

[0231] CSTB(z i )=DeformConv(LayerNorm(z i-1 +MultiHeadAttn(z i-1 ,z i-1 ,z i-1 ))

[0232] wherein MultiHeadAttn represents multi-head attention, DeformConv represents deformable convolution, LayerNorm represents a normalization layer, and z i-1 represents the output feature map of the i-1th CSTB;

[0233] The discriminator network D gan is designed to have a multi-scale discriminator structure: multi-scale discriminator {D1, D2, D3}, each discriminator uses the same backbone network PatchGAN, independently trains parameters, and the last layer outputs a probability;

[0234] The diffusion model includes a forward process and a reverse process.

[0235] The forward process is to create a noise version z0 of z0 at a random time step tt , where x0is the latent representation of the original image, i.e., without added noise, x

[0236]

[0237] where x0is the latent representation of the original image, i.e., without added noise, x t is the state of the original image after adding t-step noise, q(x t |x t-1 ) is the conditional probability distribution, representing the process of generating x t-1 from x t , following a Gaussian distribution, a t = 1 - b t is the single-step noise preservation coefficient, representing the proportion of the original information preserved at the t-th step, is the cumulative noise preservation coefficient, i.e., q(x t |x0) represents the conditional distribution of generating x t directly from x0, b t is a noise scheduler that controls the noise level at time step t, as t gradually increases, i.e., t→ T, x t will gradually approach pure noise N(0, I), and q(x t |x t-1 ) is reparameterized as a t = 1 - b t , to obtain the formula q(x t |x0), I is the identity matrix;

[0238] Reverse process: z0and are respectively subjected to an additional diffusion step to the same time step s, generating z s and Then use the pre-trained decoder D vae to decode the results back to the pixel space to obtain images x s and Images x s and are input into the discriminator D gan for evaluation:

[0239]

[0240] where I is the identity matrix, is the cumulative noise coefficient, used to control the noise level at time step s, z s is the noisy version of the true latent representation z0at time step s, is the noise level of the generator output at time step s.

[0241] In this embodiment, a meta-path controller is introduced to dynamically skip redundant CSTB computation units based on the local complexity of the feature map, thereby reducing computational overhead while maintaining generation quality.

[0242] splicing feature z init Perform local complexity awareness and output the complexity graph Ω, represented as:

[0243]

[0244] in, It is z init In the eigenvector at position (i,j), ||·||² is the L2 norm, γ is the balance factor, and Entropy(·) is the channel distribution entropy of the local 3×3 window. p k It is the probability distribution of the local window histogram in 256 bins;

[0245] The skip probability is predicted using a lightweight convolutional network, expressed as:

[0246] M skip =σ(Conv 3×3 (ReLU(Conv 1×1 (Ω))))

[0247] Where σ is the Sigmoid function, outputting a probability graph in the range [0,1]. Conv 1×1 It is a 1×1 convolutional layer, the purpose of which is to reduce the dimensionality to 8 channels;

[0248] During the reasoning phase, hard coding is used, that is... Then, at position (i,j), the current CSTB calculation is skipped, and the previous level feature is directly reused. During the training phase, a soft mask is used, that is, Gumbel-Softmax is used instead of hard coding, as shown below:

[0249]

[0250] Where G and G' are injected random noises, G and G' follow Gumbel(0,1), and τ is a temperature coefficient representing the smoothness control parameter. When τ→0 + The output approaches a hard decision, i.e., 0 or 1; when τ→∝, the output approaches a uniform distribution.

[0251] Enhance feature consistency through global dependency modeling;

[0252] Dynamic learning bias Δp∈R H×W×2N Where N is the kernel size, and the convolution sampling position is adjusted:

[0253]

[0254] wherein z is an input feature map, representing an input feature tensor to be convolved, p is a target position coordinate, representing a pixel position currently calculated on an output feature map, p n is a predefined sampling offset, representing a fixed relative offset of the n th sampling point in the convolution kernel, Δp n is a dynamic learning offset, representing an adaptive position adjustment amount of the n th sampling point, w n is a convolution weight, representing a kernel weight of the n th sampling point, and N is a total number of sampling points, representing a number of sampling points of the convolution kernel.

[0255] In this embodiment, the process of inputting the rendered frame image into the lightweight super-resolution model to obtain the reconstructed 4K resolution image includes:

[0256] The rendered frame image is encoded into a latent space using a pre-trained variational autoencoder, and noise is gradually added through a diffusion process, so that the generator learns to recover the latent representation of a high-resolution image from the noise, and then the latent representation is decoded back to the pixel space through the decoder to obtain a super-resolution image.

[0257] Embodiment Three

[0258] Based on the same inventive concept, the present disclosure also provides an electronic device corresponding to the method of any of the above embodiments, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the method of any one of the embodiments.

[0259] Figure 4 A more specific hardware structure of an electronic device provided by the present embodiment is shown, which can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 for communication within the device.

[0260] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the present embodiment.

[0261] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0262] The input / output interface 1030 is configured to connect an input / output module to realize information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0263] The communication interface 1040 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as a USB (Universal Serial Bus), a network cable, etc.) or through a wireless manner (such as a mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0264] The bus 1050 includes a channel to transmit information between various components (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.

[0265] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only include the components necessary to implement the embodiments of the present specification, and does not have to include all the components shown in the figure.

[0266] The system of the above embodiments is used to implement the three-dimensional structure recovery method of a high-quality urban renewal landscape building in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.

[0267] Embodiment Four

[0268] Based on the same inventive concept, the disclosure also provides a non-transitory computer readable storage medium storing computer instructions for causing a computer to perform the method of any of the above embodiments.

[0269] The computer readable medium of the embodiments can include permanent and non-permanent, removable and non-removable media, which can be realized by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0270] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the method of any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not described here.

[0271] Those skilled in the art should understand that the above discussion of any of the embodiments is only exemplary and is not intended to imply that the scope of the disclosure (including claims) is limited to these examples; the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of the embodiments of the disclosure as described above. In order to be brief, they are not provided in detail.

[0272] In addition, to simplify the description and discussion, and so as not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections of integrated circuit (IC) chips and other components can or can not be shown in the provided drawings. Furthermore, devices can be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details regarding the implementation of these block diagram devices are highly dependent on the platform to be implemented (i.e., these details should be well within the understanding of one of skill in the art). Where specific details (e.g., circuitry) are set forth in order to describe an illustrative embodiment of the present disclosure, it should be apparent to one skilled in the art that the present disclosure can be practiced without these specific details or with an implementation varying in details from the specific embodiments described herein. Accordingly, the description is to be considered illustrative and not restrictive.

[0273] Although the present disclosure has been described in conjunction with the specific embodiments thereof, it is to be understood that many alternatives, modifications and variations will be apparent to those skilled in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.

[0274] Accordingly, the units of the examples described in the embodiments of the present application can be realized in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled persons can utilize various methods to implement the described functions for each particular application, but such implementation should not be considered to be beyond the scope of the present application.

[0275] The above-described embodiments are merely intended to describe the preferred ways of the present application, and are not intended to limit the scope of the present application. Various modifications and improvements to the present application, which are made by those of ordinary skill in the art without departing from the design spirit of the present application, shall fall within the scope of the present application as defined by the claims.

Claims

1. A method for VR real-time rendering of low bitrate cloud rendered push streams, characterized in that, The method comprises: acquiring a rendered frame image; constructing a lightweight super-resolution model; inputting the rendered frame image into the lightweight super-resolution model to obtain a reconstructed 4K resolution image; based on the reconstructed 4K resolution image, performing real-time rendering and display on a VR device.

2. The method of claim 1, wherein, The lightweight super-resolution model comprises: One U-Net generator G, wherein G is composed of N CSTB = 4 CSTB modules; one discriminator network D gan ; one diffusion model and one pre-trained variational autoencoder VAE, wherein VAE is divided into an encoder E vae and a decoder D vae ; The generator G replaces the standard convolutional layer in the U-Net with a cross-scale Transformer module, namely CSTB, specifically replacing all 3x3 convolutional layers in the U-Net encoder and decoder, retaining downsampling and upsampling operations, and adding CSTB at the jump connection to enhance cross-scale information transmission, with the structure being: CSTB(z i ) = DeformConv(LayerNorm(z i-1 + MultiHeadAttn(z i-1 , z i-1 , z i-1 ))) wherein MultiHeadAttn denotes multi-head attention, DeformConv denotes deformable convolution, LayerNorm denotes a normalization layer, z i-1 denotes the output feature map of the i-1th CSTB; Discriminator network D gan The structure of the multi-scale discriminator is designed: multi-scale discriminator {D1, D2, D3}, each discriminator adopts the same backbone network PatchGAN, independently trains parameters, and the last layer outputs probability; The diffusion model comprises a forward process and a reverse process; where the forward process: create a noisy version z of z0 at random time step t t , is given by where x0is the latent representation of the original image, i.e., without added noise, x t represents the state of the original image after adding t-step noise, q(x t t-1 is the conditional probability distribution, representing the process of generating x t-1 from x t , which follows a Gaussian distribution, α t = 1-β t is a single-step noise retention coefficient, representing the proportion of the original information retained at the t-th step, is the cumulative noise retention coefficient, i.e. q(x t |x0) represents the conditional distribution of generating x t directly from x0, β t is a noise scheduler used to control the noise level at time step t, as t gradually increases, i.e., t→T, then x t will gradually approach pure noise N(0, I), and q(x t |x t-1 ) is reparameterized, letting α t = 1-β t , q(x t |x0) is obtained, and I is the identity matrix;​ Reverse process: z0and are generated by additional diffusion steps to the same time step s, respectively s and The result is then decoded back to pixel space using a pretrained decoder D vae to obtain the image x s and The image x s and is input into the discriminator D gan for evaluation: where I is the identity matrix, is the accumulated noise factor, used to control the noise level at time step s, z s is a noisy version of the real latent representation z0at time step s, is the generator output the noise level at time step s.

3. The method of claim 2, wherein, An element path controller is introduced to dynamically skip redundant CSTB calculation units according to the local complexity of the feature map, reducing computational overhead while maintaining generation quality: On the concatenation feature z init Local complexity perception is performed, and a complexity map Ω is output, represented as: wherein, is z init is the feature vector at position (i,j), ||·||2 is the L2 norm, γ is a balancing factor, Entropy(·) is the channel distribution entropy of a local 3x3 window, p k is the probability distribution of the local window histogram in 256 bins; A lightweight convolutional network is used to predict a skip probability, represented as: M skip = σ(Conv 3×3 (ReLU(Conv 1×1 (Ω)))) where σ is the Sigmoid function, outputting [0, 1] probability maps, Conv 1×1 is a 1 x 1 convolutional layer, with the purpose of reducing dimensionality to 8 channels; In the inference phase, hard-coding is adopted, i.e. Then, in position (i, j), the current CSTB calculation is skipped, and the previous feature is directly reused. In the training phase, soft-masking is adopted, i.e., Gumbel-Softmax is used instead of hard-coding, which is represented as: where G and G' are injected random noises, G and G' obey Gumbel(0, 1), τ is a temperature coefficient, and τ is a smoothing control parameter, when τ→0 + , the output tends to hard decision, i.e., 0 or 1; when τ→∞, the output tends to uniform distribution, i.e. Global dependency modeling is used to enhance feature consistency; Dynamic learning bias Δp ∈ R H×W×2N where N is the size of the convolution kernel, adjusting the convolution sampling position: wherein z is an input feature map, representing an input feature tensor to be convolved, p is a target position coordinate, representing a pixel position currently calculated on the output feature map, p n is a predefined sampling offset, representing a fixed relative offset of the nth sampling point in the convolution kernel, Δp n is a dynamic learning offset, representing an adaptive position adjustment amount of the nth sampling point, w n is a convolution weight, representing a kernel weight of the nth sampling point, and N is a total number of sampling points, representing a number of sampling points of the convolution kernel.

4. The method of claim 2, wherein, The method for inputting the rendered frame image into the lightweight super-resolution model to obtain the reconstructed 4K resolution image comprises: A pre-trained variational autoencoder is used to encode the rendered frame image into a latent space, and a diffusion process is used to gradually add noise, allowing the generator to learn the latent representation of a high-resolution image from the noise, and then the decoder is used to decode the latent representation back to the pixel space to obtain a super-resolution image.

5. A system for low bitrate cloud rendered streaming of VR real-time rendering, the system being configured to implement the method of any one of claims 1-4, characterized in that, The system comprises an acquisition module, a construction module, a reconstruction module, and a rendering module; The acquisition module is configured to acquire a rendered frame image; The construction module is configured to construct a lightweight super-resolution model; The reconstruction module is configured to input the rendered frame image into the lightweight super-resolution model to obtain a reconstructed 4K resolution image; The rendering module is configured to perform real-time rendering and display on a VR device based on the reconstructed 4K resolution image.

6. The system of claim 5, wherein, The lightweight super-resolution model comprises: One U-Net generator G, wherein G is composed of N CSTB = 4 CSTB modules; one discriminator network D gan ; one diffusion model and one pre-trained variational autoencoder VAE, wherein VAE is divided into an encoder E vae and a decoder D vae ; The generator G replaces the standard convolutional layer in the U-Net with a cross-scale Transformer module, namely CSTB, specifically replacing all 3x3 convolutional layers in the U-Net encoder and decoder, retaining downsampling and upsampling operations, and adding CSTB at the jump connection to enhance cross-scale information transmission, with the structure being: CSTB(z i ) = DeformConv(LayerNorm(z i-1 + MultiHeadAttn(z i-1 , z i-1 , z i-1 ))) where MultiHeadAttn denotes multi-head attention, DeformConv denotes deformable convolution, LayerNorm denotes a normalization layer, z i-1 denotes the output feature map of the i-1th CSTB. Discriminator network D gan The structure of the multi-scale discriminator is designed: multi-scale discriminator {D1, D2, D3}, each discriminator adopts the same backbone network PatchGAN, independently trains parameters, and the last layer outputs probability; The diffusion model comprises a forward process and a reverse process; where the forward process: create a noisy version z of z0 at random time step t t , is given by where x0is the latent representation of the original image, i.e., without added noise, x t q(xt) represents the state of the original image after adding t-step noise, q(x t t-1 is the conditional probability distribution, representing the process of generating x t-1 from x t , subject to a Gaussian distribution, α t = 1-β t is a single-step noise retention coefficient, representing the proportion of original information retained at the t-th step, is the cumulative noise retention coefficient, i.e. q(x t |x0) represents the conditional distribution of generating x t directly from x0, β t is a noise scheduler used to control the noise level at time step t, as t gradually increases, i.e., t→T, x t will gradually approach pure noise N(0, I), and q(x t |x t-1 ) is reparameterized, letting α t = 1-β t , q(x t |x0) is obtained, and I is the identity matrix;​ Reverse process: z0and are generated by additional diffusion steps to the same time step s, respectively s and The result is then decoded back to pixel space using a pretrained decoder D vae to obtain an image x s and The image x s and is input into the discriminator D gan for evaluation: where I is the identity matrix, is the accumulated noise factor, used to control the noise level at time step s, z s is a noisy version of the true latent representation z0at time step s, is the generator output the noise level at time step s.

7. The system of claim 6, wherein, An element path controller is introduced to dynamically skip redundant CSTB calculation units according to the local complexity of the feature map, reducing computational overhead while maintaining generation quality: On the concatenation feature z init Local complexity perception is performed, and a complexity map Ω is output, represented as: wherein, is z init is the feature vector at position (i,j), ||·||2 is the L2 norm, γ is a balancing factor, Entropy(·) is the channel distribution entropy of a local 3x3 window, p k is the probability distribution of the local window histogram in 256 bins; A lightweight convolutional network is used to predict a skip probability, represented as: M skip = σ(Conv 3×3 (ReLU(Conv 1×1 (Ω)))) where σ is the Sigmoid function, outputting [0, 1] probability maps, Conv 1×1 is a 1 x 1 convolutional layer, with the purpose of reducing dimensionality to 8 channels; In the inference phase, hard-coding is adopted, i.e. Then, in position (i, j), the current CSTB calculation is skipped, and the previous feature is directly reused. In the training phase, soft-masking is adopted, i.e., Gumbel-Softmax is used instead of hard-coding, which is represented as: where G and G' are injected random noises, G and G' obey Gumbel(0, 1), τ is a temperature coefficient, and τ denotes a smoothness control parameter, when τ→0 + , the output tends to hard decision, i.e., 0 or 1; when τ→∞, the output tends to uniform distribution, i.e. Global dependency modeling is used to enhance feature consistency; Dynamic learning bias amount Δp ∈ R H×W×2N where N is the size of the convolution kernel, adjusting the convolution sampling position: wherein z is an input feature map, representing an input feature tensor to be convolved, p is a target position coordinate, representing a pixel position currently calculated on an output feature map, p n is a predefined sampling offset, representing a fixed relative offset of the nth sampling point in the convolution kernel, Δp n is a dynamic learning offset, representing an adaptive position adjustment amount of the nth sampling point, w n is a convolution weight, representing a kernel weight of the nth sampling point, and N is a total number of sampling points, representing a number of sampling points of the convolution kernel.

8. The system of claim 6, wherein, The process of inputting the rendered frame image into the lightweight super-resolution model to obtain the reconstructed 4K resolution image comprises: A pre-trained variational autoencoder is used to encode the rendered frame image into a latent space, and a diffusion process is used to gradually add noise, allowing the generator to learn the latent representation of a high-resolution image from the noise, and then the decoder is used to decode the latent representation back to the pixel space to obtain a super-resolution image.

9. An electronic device, comprising: A computer program product comprising a computer readable storage medium having stored thereon computer program means, the computer program product being configured such that, upon execution by a processor, the processor is caused to perform the method according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program which, when executed, implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Video transmission method, server, terminal and video transmission system

    CN114584805A

  • Super-resolution reconstruction method for single image with potential features

    CN117078510A

  • Image super-resolution reconstruction method and system based on diffusion model

    CN117522694A

  • Video prediction method, device and equipment based on dynamic routing and readable medium

    CN119893132A

  • Three-dimensional model reconstruction method based on neural radiation field

    CN120298570A