4K real-time rendering method and system for efficient cloud plug flow

By constructing a lightweight single-image super-resolution network model, and combining prior feature extraction and a backbone network, the network bandwidth and computing resource issues of 4K real-time rendering on low-end devices are solved, achieving efficient 4K image reconstruction, improving image quality and processing efficiency, and making it suitable for applications such as distance education and virtual tourism.

CN121037619APending Publication Date: 2025-11-28北京渲光科技有限公司 +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511182736.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-28

Smart Images

  • Figure CN121037619A_ABST
    Figure CN121037619A_ABST
Patent Text Reader

Abstract

The invention discloses a 4K real-time rendering method and system for efficient cloud plug flow, and the method comprises the steps: carrying out the real-time rendering of a picture through a cloud end, recognizing a key frame, carrying out the coding compression of the key frame, and transmitting the key frame to a client; constructing a 4K real-time rendering model, and downloading the 4K real-time rendering model to the client; and in the client, converting the key frame to 4K resolution by using a 4K real-time rendering model to finish rendering. According to the method, the problems, such as block artifacts, ringing effects and edge detail loss, existing in the traditional super-resolution technology are effectively solved, and meanwhile, the overall processing efficiency and the image quality are improved. The method not only provides powerful technical support for cloud simulation, but also opens up new possibilities for other application scenes depending on high-quality video streams, such as distance education, virtual tourism, cloud games and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of high-definition image rendering, specifically to a high-efficiency cloud streaming 4K real-time rendering method and system. Background Technology

[0002] Achieving 4K real-time rendering on low-end devices is significant for cloud simulation applications. For example, sports simulations typically require high-definition visuals to ensure realism and accuracy, facilitating training and evaluating athletes' performance. 4K resolution allows for more detailed rendering of training environments and key details such as athlete movements, enhancing participant immersion and decision-making accuracy. Furthermore, this technology supports remote collaboration and training, enabling team members in different geographical locations to interact within the same virtual environment, greatly improving training efficiency and flexibility.

[0003] However, existing technologies face numerous challenges in transmitting 4K ultra-high-definition video from high-end cloud devices to low-end devices. First, directly transmitting 4K video streams consumes significant network bandwidth, increasing costs and potentially causing latency, impacting user experience. To overcome this, a common approach is to transmit lower-resolution footage first and then upscale it on the local device. However, this method also presents challenges: converting low-resolution video to 4K ultra-high-definition on a low-end device requires powerful GPU support, which most low-end devices lack. Even with appropriate software solutions, complex network structures and numerous parameters slow down processing speed, failing to meet the demands of real-time video processing. Furthermore, existing super-resolution algorithms often amplify block artifacts and ringing effects in decoded frames and ignore edge details, resulting in suboptimal reconstruction results. More importantly, these methods fail to fully utilize prior coding information to guide the super-resolution network process and do not consider the impact of compression distortion and video content features on super-resolution network technology, further limiting their effectiveness. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes a high-efficiency cloud-based 4K real-time rendering method and system. This method constructs a lightweight single-image super-resolution network model on low-spec devices, aiming to extract features from encoded prior information and fuse them with features from the backbone network to enhance the quality of the reconstructed image. Specifically, this method introduces a spectral weight matrix to weight each frequency component, guiding the model to focus on the recovery of high-frequency information, which is crucial for maintaining image sharpness and detail. Furthermore, to achieve model lightweighting while maintaining performance, this method employs reparameterization techniques and pixel-adaptive convolution modules. These techniques not only reduce computational resource requirements but also enable the model to better adapt to resource-constrained terminal devices, ensuring its real-time operation.

[0005] To achieve the above objectives, this invention provides an efficient cloud streaming method for 4K real-time rendering, comprising the following steps:

[0006] The system renders the image in real time using the cloud, then identifies keyframes, encodes and compresses the keyframes, and sends them to the client.

[0007] Build a 4K real-time rendering model and download the 4K real-time rendering model to the client;

[0008] In the client, the 4K real-time rendering model is used to convert keyframes to 4K resolution to complete the rendering.

[0009] Preferably, the 4K real-time rendering model includes a prior feature extraction network and a backbone network;

[0010] The prior feature extraction network is used to extract prior features from keyframes;

[0011] The backbone network is used to render keyframes based on the extracted features.

[0012] Preferably, the prior feature extraction network includes a multi-convolutional layer structure, which consists of two 3×3 convolutional layers and a convolutional block containing two 1×1 convolutional layers and one 3×3 convolutional layer, to reduce computational cost.

[0013] F Rep =Conv 3×3 (Conv 3×3 (F in )+Conv 1×1 _Conv 3×3 (F in ))

[0014] Among them, F in F represents the input features, Conv represents the convolutional layer, and F represents the input features. Rep This indicates multiple convolutional layers.

[0015] Preferably, the prior feature extraction network further includes pixel-adaptive convolutional blocks, which are used to generate spatially adaptive convolutional kernels. The pixel-adaptive convolution formula is as follows:

[0016]

[0017] Where, p i ∈(x i x j ) T These are pixel coordinates, Ω(i) defines an s×s convolution window, and the image features v = (v1, v2, ..., v...). n ), Representing the features of each pixel, filter weights and bias v j Let v′ represent the feature vector at position j in the input feature map. i This represents the feature vector at position i in the output feature map;

[0018] The convolution kernel weights of PAC are binarized to ±1, retaining only the sign information, and a scaling factor ∝ is introduced to compensate for quantization error:

[0019] W′=∝·Sign(W)

[0020] Where Sign is the binarization operation, W is the original weight of the PAC convolution kernel, W′ is the new weight of the PAC convolution kernel, and ∝ is the scaling factor.

[0021] Preferably, the backbone network includes: a downsampling module, an upsampling module, a channel attention module, and a bidirectional pyramid;

[0022] The downsampling module is used to extract deep features;

[0023] The upsampling module is used to restore spatial resolution;

[0024] The channel attention module is used to enhance the expressive power of features and highlight important features;

[0025] The bidirectional pyramid is used for feature fusion via a cross-scale feature pyramid.

[0026] Preferably, global modeling is performed on the downsampled low-resolution features:

[0027] F trans =MobileViT(F down )

[0028]

[0029] in, Indicates element-wise addition. MobileViT is a lightweight Transformer network module. trans F represents the feature map output by the MobileViT module. down F represents the feature map after the downsampling operation. skip This represents the skip connection feature map.

[0030] Preferably, the feature fusion step of the bidirectional pyramid includes:

[0031] First, a top-down path operation is performed to upsample deep, high-semantic features layer by layer and fuse them with shallow features. Simultaneously, a bottom-up path is used to downsample shallow, detailed features layer by layer and pass them to deeper layers to supplement local information. Then, dynamic weights are applied to each layer of fused features to allocate attention, suppress redundant information, and enhance key features. Finally, the dynamically weighted features at each scale are concatenated, and channel information is integrated through convolution to achieve the fusion of multi-scale features.

[0032] Preferably, the backbone network also incorporates a spatiotemporal attention module, the workflow of which includes:

[0033] By aligning the feature maps of adjacent frames with the current frame through optical flow-guided feature alignment, and using a multi-head spatiotemporal attention mechanism, spatial local details and temporal global motion patterns are captured respectively.

[0034] By fusing the current frame with spatiotemporal enhancement features through a learnable gating mechanism, dynamic detail modeling is enhanced and the stability of dynamic scene detail recovery in video sequences is improved.

[0035] Preferably, the 4K real-time rendering model is optimized using meta-learning dynamic hyperparameter tuning:

[0036]

[0037] η = MetaNet(State(F) in ))

[0038] Where α, β, γ, δ, and ρ are the weights of the loss function, η is the learning rate, and F in These are the input features of the current batch. It represents the gradient of the current batch of data. MetaNet is a lightweight meta-learning network consisting of 3 layers of MLP. State is the extracted state function.

[0039] The present invention also provides a high-efficiency cloud streaming 4K real-time rendering system, the system being used to implement the above method, comprising: a cloud module, a building module, and a client module;

[0040] The cloud module is used to render the image in real time, then identify key frames, encode and compress the key frames and send them to the client module.

[0041] The building module is used to build a 4K real-time rendering model and download the 4K real-time rendering model to the client module;

[0042] In the client module, the 4K real-time rendering model is used to convert keyframes to 4K resolution to complete the rendering.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] This invention effectively solves the problems existing in traditional super-resolution technologies, such as block artifacts, ringing effects, and loss of edge details, while improving overall processing efficiency and image quality. This not only provides strong technical support for cloud simulation but also opens up new possibilities for other application scenarios that rely on high-quality video streams, such as distance education, virtual tourism, and cloud gaming. Attached Figure Description

[0045] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a schematic diagram of the system structure according to an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of the prior feature extraction network structure according to an embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of the backbone network structure according to an embodiment of the present invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] First, let's introduce the technical terms used in this invention.

[0052] Cloud: High-performance computing devices.

[0053] Client: Devices with low computing power, such as mobile phones.

[0054] Keyframes: These are images rendered in real-time by cloud applications and are mainly low-quality images.

[0055] Example 1

[0056] This embodiment provides an efficient cloud streaming method for 4K real-time rendering, the steps of which include:

[0057] S1. Render the image in real time using the cloud, then identify keyframes, encode and compress the keyframes, and send them to the client.

[0058] The system renders the scene in real time in the cloud (using graphics engines such as Unreal Engine, Unity 3D, and Lumverse 3D), then identifies keyframes, encodes and compresses them, and sends them to the client via WebRTC. Subsequently, the client converts the low-quality "input keyframes" (e.g., 1080p) to 4K resolution. In this embodiment, an overfitted network model is trained for each virtual scene to reduce network parameters, making it more lightweight, lowering latency, and reducing the hardware requirements of the client, allowing it to run on low-configuration clients (such as mobile phones).

[0059] In this embodiment, a model is trained for each scene, so there is no need to transmit the model to the client every frame or at intervals. This embodiment uses a virtual scene to transmit the model once, and both the cloud and the client have a "model library". Generally, there is no need to transmit the model during real-time streaming, thereby reducing the amount of data transmitted over the network. This is equivalent to achieving "efficient cloud streaming" in the same cloud streaming process.

[0060] S2. Build a 4K real-time rendering model and download it to the client.

[0061] like Figure 1 The diagram shown is a structural schematic of the 4K real-time rendering model in this embodiment. The model runs on the client and consists of two parts: a "prior feature extraction network" and a "backbone network".

[0062] (1) Prior Feature Extraction Network

[0063] like Figure 2The diagram shown illustrates the prior feature extraction network structure of this embodiment. Encoding prior information refers to data generated during video encoding. This data reflects the features of the video content and changes during the encoding process, providing the model with rich contextual information. This allows the model to more accurately understand the features of the video content and the artifacts introduced during encoding, thereby achieving better results in super-resolution reconstruction. The generation of encoding prior information is mainly based on existing video coding standards (such as H.264 / AVC, H.265 / HEVC, H.266 / VVC) and their encoders.

[0064] Encoding prior information specifically includes the following aspects:

[0065] Prediction Signal: In video coding, the prediction signal represents the difference between the current block and the reference block. It helps the model understand the temporal continuity and spatial relevance of video content. The prediction signal provides useful information about the motion and changes in video content, thus guiding the model to better recover high-frequency details.

[0066] Quantization Parameter (QP): The quantization parameter determines the size of the quantization step, thus affecting the encoded video quality and compression ratio. A higher QP value results in a higher compression ratio, but the video quality will decrease accordingly. Through QP, the model can understand the degree of compression introduced during the encoding process, thereby better handling artifacts caused by quantization.

[0067] Residual: The residual is the difference between the original signal and the predicted signal. In video coding, the residual is typically transformed and quantized before being stored. Residual information helps the model understand the specific details of each block's changes, thus enabling more accurate recovery of these details during reconstruction.

[0068] Partition Map: A partition map describes how a video frame is divided into blocks, reflecting the texture complexity and object shapes of the video content. By analyzing the partition map, the model can understand the size and shape of each block, thereby better capturing the structural information of the video content and guiding the model to avoid or reduce the generation of block artifacts during reconstruction.

[0069] The contribution of different encoding priors to super-resolution reconstruction may change dynamically with the input content, so "channel attention" is used for dynamic weight allocation.

[0070] Specifically, an independent attention branch is designed for each encoded prior feature (predicted signal, QP, residual, partition map). A lightweight channel attention (Squeeze-and-Excitation Block) is used to generate weight vectors, dynamically adjusting the importance of each prior feature.

[0071]

[0072] in, σ is the i-th type of encoded prior feature (the prior feature before it is input to the channel attention), σ is the Sigmoid activation function, GAP is global average pooling, and FC is fully connected. These are the prior features after channel attention.

[0073] Then, a 3x3 convolutional layer with 32 channels is used to extract features from the encoded prior:

[0074] F0=γ☉F in +β

[0075] Among them, F in represents the input features, i.e., the features extracted from the input keyframe after passing through the convolutional layer. ⊙ represents element-wise multiplication. γ and β are a pair of affine transformation parameters generated based on the input frame and encoding prior characteristics, used to generate different depth compression domain information.

[0076] Figure 2 The "multiple convolutional layers" F Rep It consists of two 3×3 convolutional layers and a convolutional block containing two 1×1 convolutional layers and one 3×3 convolutional layer. The outputs of these convolutional layers are summed during training. During inference, these complex convolutional structures are reparameterized into a simple 3×3 convolutional layer, thus significantly reducing computational cost.

[0077] F Rep =Conv 3×3 (Conv 3×3 (F in )+Conv 1×1 _Conv 3×3 (F in ))

[0078] Here, Conv represents a convolutional layer.

[0079] Dynamic parameterization: Existing "multiple convolutional layers" are fixed as a single convolution during inference, which cannot adapt to complex scenarios. Therefore, a dynamic path selection mechanism is introduced during training, selecting different convolutional branches (3×3, 1×1, 5×5) based on the input content. This improves global feature modeling capabilities and reduces computational redundancy during inference. During inference, these branches are merged into dynamically weighted convolutional kernels.

[0080] W merged =∑β i ·W i

[0081] Where, β i W represents the statistical probability of path selection during training. i The weights for each branch.

[0082] Dynamic channel pruning: Different inputs have different channel requirements, and static channel allocation is redundant. Therefore, a differentiable channel pruning mechanism is introduced in the "multiple convolutional layers," and the channel importance score s is learned during training. c :

[0083] s c =Sigmoid(g c ·θ)

[0084] Among them, g c Here, θ is the channel gradient, and θ is a learnable parameter. During inference, channels with scores above a threshold (e.g., 0.5) are retained, while the rest are set to zero.

[0085] Feature fusion: Three "multi-convolutional layers" are used to extract compressed domain information at different depths. Two pixel-adaptive convolutional blocks with partition maps are used to generate compressed domain information at different depths.

[0086] F1 = f Rep (F0)

[0087]

[0088] F2=f PAC (f down (F′1), Part)

[0089]

[0090] F3 = f PAC (f down (F′2), AvgPool(Part, 2))

[0091]

[0092] Among them, f down This indicates a downsampling operation implemented using a 3x3 convolutional block with a stride of 2. Rep It is a "multiple convolutional layer", f PAC It is a pixel-adaptive convolutional block. Part is the partition map of the input frame, and AvgPool(·) is the average pooling function.

[0093] To fully utilize multi-scale information, enhance the dynamic utilization of prior information, and improve the recovery of high-frequency details, this embodiment introduces dilated convolution to extract multi-scale features. The dilation rate is set to [1, 2, 4], yielding F′1, F′2, and F′3 respectively (e.g., ...). Figure 2 As shown), α 1,k α 2,k and α 3,k These are learnable weight coefficients. and These are characteristics of different void ratios.

[0094] The pixel-adaptive convolutional block (PAC) described above is used to generate spatially adaptive convolutional kernels. The partition map provides useful information about the object's shape and texture complexity. The pixel-adaptive convolution formula is:

[0095]

[0096] Where, p i ∈(x i x j ) T These are pixel coordinates, Ω(i) defines an s×s convolution window, and the image features v = (v1, v2, ..., v...). n ), Representing the features of each pixel, filter weights and bias v j Let v′ represent the feature vector at position j in the input feature map. i This represents the feature vector at position i in the output feature map.

[0097] It is a function that generates the relative positional convolutional kernel weights based on partition map information:

[0098]

[0099] Where ∈ is the threshold (∈=16), and the clamp(·) function sets the negative weights to 0.

[0100] PAC's floating-point convolution computation is computationally expensive. Therefore, the convolution kernel weights of PAC are binarized to ±1, retaining only the sign information, and a scaling factor ∝ is introduced to compensate for quantization errors.

[0101] W′=∝·Sign(W)

[0102] Where Sign is the binarization operation, W is the original weight of the PAC convolution kernel, W′ is the new weight of the PAC convolution kernel, and ∝ is the scaling factor. The updated PAC convolution kernel weight matrix improves inference speed by 30%, reduces the number of parameters by 40%, and maintains performance.

[0103] (2) Backbone Network

[0104] The backbone network is responsible for extracting and fusing deep features from low-resolution features. Its architecture is as follows: Figure 3 As shown, multi-scale features are captured through downsampling and upsampling operations, and features at different levels are fused through skip connections.

[0105] Input feature processing: The input features are first processed through a 3×3 convolutional layer with 32 channels, encoding them into a high-dimensional feature space. The purpose of this layer is to map the low-resolution features to a higher-dimensional space so that subsequent modules can extract features more effectively.

[0106] The backbone network mainly includes the following modules:

[0107] ① Downsampling module: Extracts deep features through multiple convolution and pooling operations.

[0108] (②Upsampling module: gradually restores spatial resolution through deconvolution or pixel shuffling operations.)

[0109] ③ Attention Channel Module: Composed of a channel attention module and a spatiotemporal attention module in parallel, it is used to enhance the expressive power of features and highlight important features.

[0110] (④ "Two-way pyramid": Feature fusion is performed through a cross-scale feature pyramid.)

[0111] ① Downsampling module:

[0112] downsampling module F down This is achieved using a 3×3 convolutional layer with a stride of 2, i.e.:

[0113] F down =Conv3×3(F in stride=2)

[0114] Each downsampling operation halves the spatial size of the feature map while increasing the number of channels, thereby extracting higher-level semantic information. In the backbone network, the input features undergo two downsampling operations. After each downsampling, the feature map size is halved, and the number of channels increases. Furthermore, to leverage the global modeling capabilities of the Transformer, global modeling is performed on the downsampled low-resolution features:

[0115] F trans =MobileViT(Fdown )

[0116]

[0117] in, Indicates element-wise addition. MobileViT is a lightweight Transformer network module. trans F represents the feature map output by the MobileViT module. down F represents the feature map after the downsampling operation. skip This represents the skip connection feature map.

[0118] ② Upsampling module:

[0119] Upsampling module F up This is achieved through 1×1 convolutional layers and pixel shuffling operations. The 1×1 convolutional layers adjust the number of channels in the feature map, while the pixel shuffling operation doubles the spatial size of the feature map. During upsampling, skip connections are used to fuse the features from the downsampling stage with those from the upsampling stage. This fusion method helps retain more detailed information and improves reconstruction quality.

[0120] F up =PixelShuffle(Conv 1×1 (F pyramid ))

[0121] Where PixelShuffle represents the pixel shuffling operation, F pyramid This represents the multi-scale feature map after fusion by the "bidirectional pyramid" module.

[0122] ③ Channel attention module: It uses feature gating to modulate input features and uses channel attention to extract global information, followed by the application of two fully connected layers.

[0123] Feature gating: In the downsampled feature map, a feature gating mechanism is used to modulate the input features. Feature gating learns a weight vector and weights the input features element-wise, thereby highlighting important features and suppressing unimportant features.

[0124] Channel Attention: The channel attention module extracts global information through global average pooling and fully connected layers, and generates channel weights. The specific formula is as follows:

[0125] F channel =σ(FC(FC(GAP(F) skip ))))

[0126] Where σ is the activation function, FC is fully connected, GAP is global average pooling, and F channel This represents the channel attention feature.

[0127] Attention module structure: The attention block combines feature gating and channel attention.

[0128] F channel-attention =F gating ⊙F channel

[0129] Among them, F channel-attention This represents the channel attention module, ⊙ represents element-wise multiplication, and F gating This indicates a gating feature.

[0130] In video super-resolution, adjacent frames contain rich complementary information (such as motion details and occluded regions). Existing methods only process single-frame images and do not utilize the spatiotemporal correlation between adjacent frames in a video sequence, leading to unstable detail recovery in dynamic scenes. To address these issues, this embodiment constructs a spatiotemporal attention module, as follows:

[0131] Optical flow-guided feature alignment: Using a lightweight optical flow network (such as PWC-Net Lite) to estimate the optical flow between neighboring frames and the current frame, and aligning the feature maps, can reduce motion blur.

[0132]

[0133] Where t represents the timestamp of the current frame in the video sequence, I t I represents the low-resolution input frame at the current time t. t -1 This represents the low-resolution input frame from the previous time step t-1. F t-1 Indicates from adjacent frame I t-1 Extracted deep feature map. Warp represents the spatial deformation of the feature map based on the optical flow field. Indicates that F t-1 Alignment feature map after being distorted to the current frame t coordinate system.

[0134] The aligned features are concatenated with the features of the current frame and input into the spatiotemporal attention module.

[0135] Spatiotemporal Cooperative Attention: A multi-head spatiotemporal attention mechanism is designed to capture local spatial details and global temporal motion patterns, which can enhance dynamic detail modeling.

[0136]

[0137] Where Q, K, and V are jointly generated from the features of the current frame and the aligned frame, STAC represents spatiotemporal collaborative attention, which is an extension of the multi-head attention mechanism in video super-resolution, and d represents the scaling factor dimension, which is the feature dimension of each attention head.

[0138] Dynamic weight fusion: The current frame and spatiotemporal enhancement features are fused through a learnable gating mechanism.

[0139] F time-space =σ(W g ·[F current F STCA ])⊙F current +(1-σ(W g ))⊙F STCA

[0140] Where σ is the Sigmoid activation function, ⊙ represents element-wise multiplication, and F current F represents the deep features extracted from the keyframe at the current time t. STCA The output feature of the spatiotemporal collaborative attention module is represented by F. STCA =STCA(Q, K, V), W g F represents the learnable gated weight matrix. time-space This represents the final spatiotemporal characteristics after dynamic fusion.

[0141] Integrating the "Channel Attention Module" F channel-attention And the "spatiotemporal attention module" F time-space .

[0142]

[0143] Among them, F attention This represents the total output of the "attention module". This indicates that each element is added together.

[0144] ④ "Two-way pyramid": Feature fusion is performed through a cross-scale feature pyramid.

[0145] Multi-depth feature extraction: Feature maps at different levels of the backbone network have different spatial resolutions (e.g., F′1 is the original image). Size, F′2 is the original image Size, F′3 is the original image size).

[0146] Feature alignment and upsampling: Upsampling low-resolution features (e.g., bilinear interpolation or deconvolution) aligns them with the dimensions of high-resolution features.

[0147]

[0148] Where k = 1, 2, 3, This indicates that the elements are added one by one. This represents the result of upsampling the deep features of layer k. Upsample indicates the upsampling operation, which increases the spatial resolution of the fused features to the size of the next higher level.

[0149] Cross-scale feature fusion: Constructing a pyramid structure for bidirectional information flow, with both top-down and bottom-up approaches. Top-down path: Deep, high-semantic features are upsampled layer by layer and fused with shallower features.

[0150]

[0151] in, It is the result of upsampling deep features from a higher level (k+1 layer). It involves performing a 1×1 convolution on the shallow features of the current layer (layer k) to adjust the number of channels or enhance feature representation. These are the fused features, containing deep semantic information and shallow spatial details. For example, upsampling F3 and fusing it with F2, then upsampling it again and fusing it with F1.

[0152] Bottom-up approach: Shallow detailed features are passed to deeper layers through downsampling to supplement local information.

[0153]

[0154] in, It is the deep feature of the current layer (layer k), such as high-order semantic information extracted after multiple convolutions. It is the result of downsampling detailed features from a shallower layer (layer k-1). Concat concatenates deep features with downsampled shallow features along the channel dimension, preserving local details and global semantics. 3×3 It eliminates channel redundancy and enhances spatial consistency by fusing and stitching features through 3×3 convolution. It is a refined feature that blends deep semantics with shallow details.

[0155] Top-down and bottom-up fusion: fusion features for each layer (e.g., ...) and Dynamic weighted attention (DPA) is applied to suppress redundant information and enhance key features:

[0156]

[0157] in, This represents the bidirectional fusion feature of the k-th level pyramid.

[0158] Dynamic Weighted Attention Enhancement: Dynamic weighted attention allocation (DPA) is applied to the fused features to further adjust the contributions of features at different scales. Channel attention weights are calculated independently for each scale feature, suppressing redundant information and enhancing high-frequency details.

[0159]

[0160] in, This represents the k-th level feature after dynamic weighted attention enhancement. GAP represents global average pooling, FC represents a fully connected layer, σ is the sigmoid activation function, and ⊙ represents element-wise multiplication.

[0161] Multi-scale feature concatenation and output: Dynamically weighted features from various scales are concatenated, and channel information is integrated through a 1×1 convolution.

[0162]

[0163] Among them, F pyramid The final fused feature representing multi-scale features is the output of the bidirectional pyramid structure. The three levels of features are represented by dynamic weighted attention enhancement, which are the highest resolution level features (level 1), the intermediate resolution level features (level 2), and the lowest resolution level features (level 3).

[0164] After multiple downsampling, upsampling, and feature fusion processes, the feature map is transformed into the final output feature map through a 1×1 convolutional layer. This step adjusts the number of channels in the feature map to match that of the high-resolution image.

[0165] The above design enables the backbone network to effectively extract and fuse features with limited computing resources, thereby improving the quality of super-resolution reconstruction.

[0166] The loss function of the model in this embodiment is as follows:

[0167] L1 loss: used to calculate the absolute error between the predicted and true values, effectively measuring pixel-level differences in an image.

[0168]

[0169] Where N is the total number of pixels in the image, I pred (i) is the pixel value of the image predicted by the model at position i, I gt (i) is the pixel value of the real image at position i.

[0170] Perceived loss Introducing a pre-trained VGG-19 network to extract features, and perceptual loss can improve visual quality.

[0171]

[0172] Wherein, VGG(SR) represents the feature map extracted by the pre-trained VGG-19 network from the super-resolution reconstructed image, and VGG(HR) represents the feature map extracted by the pre-trained VGG-19 network from the real high-resolution image.

[0173] Combat losses Add a lightweight discriminator (PatchGAN) to enhance the realism of details.

[0174]

[0175] Where D(HR) represents the confidence level of the discriminator (D is a binary classification neural network) in the real image, and D(SR) represents the confidence level of the discriminator in the generated image.

[0176] Partition Focus Frequency Loss (PFFL): In compressed video, high-frequency information (such as edges and details) is often lost during the encoding process (such as transform and quantization). This loss function, through frequency domain analysis, enables the model to adaptively focus on these difficult-to-synthesize high-frequency components, thereby improving the quality of the reconstructed image. The specific implementation steps are as follows:

[0177] Segmentation coding unit: The super-resolution output and high-resolution image are segmented into 32×32 pixel coding units. These coding units are used for subsequent frequency domain analysis.

[0178] Fast Fourier Transform: Perform a Fast Fourier Transform (FFT) on each CU to transform the image from the spatial domain to the frequency domain.

[0179] Frec SR =FFT(SR) CU )

[0180] Frec HR =FFT(HR) CU )

[0181] Among them, SR CU and HR CU Frec represents the coding unit for super-resolution and high-resolution images, respectively. SR and Frec HR These represent their frequency domain representations, respectively.

[0182] Learnable spectral weights: Fixed thresholds and manually designed spectral weights lack adaptability and struggle to adapt to the frequency distribution differences of various video content. Therefore, a lightweight convolutional network (2 layers of 1×1 convolutions) is used to dynamically generate the spectral weight matrix ω.

[0183] ω(x, y, z) = Softmax(Conv) 1×1 (|Frec SR (x, y, z)-Frec HR (x, y, z)|))

[0184] The loss function is:

[0185]

[0186] Here, H and W represent the height and width of each frame, respectively. SR and Frec HR Let (x, y, z) represent the frequency domain representations of the super-resolution output and the high-resolution image, respectively, where (x, y, z) represents the three-dimensional position coordinates in the frequency domain. SR (x, y, z) represents the complex value of the super-resolution image at position (x, y, z) in the frequency domain. SR (x, y, z) represents the complex value in the frequency domain of the true high-resolution image at position (x, y, z).

[0187] Multi-terminal wavelet loss: FFT is sensitive to global frequency but lacks local frequency band analysis capabilities. Therefore, Discrete Wavelet Transform (DWT) is performed on the coding unit, decomposing it into low-frequency (LL) and high-frequency (LH, HL, HH) sub-bands. Then, the L1 loss of each sub-band is calculated and weighted.

[0188]

[0189] Where, λ b Subband weights (e.g., λ) HH =0.5, λ LH =0.3). Through multi-terminal wavelet loss, the frequency region of interest can be adaptively adjusted, enhancing the robustness of high-frequency detail recovery. DWT represents Discrete Wavelet Transform, and b represents the sub-band type identifier. b (SR) represents the wavelet transform coefficients of the super-resolution image SR in subband b, DWT b (HR) represents the wavelet transform coefficients of the high-resolution image HR on subband b.

[0190] The advantages of PFFL are: Adaptive frequency focus: PFFL dynamically adjusts weights to adaptively focus on frequency components that are difficult to synthesize, thereby improving the model's ability to recover high-frequency details. High-frequency detail recovery: Through frequency domain analysis, PFFL can effectively recover high-frequency information lost during compression, such as edge and texture details. Optimized reconstruction quality: By guiding the model to focus on high-frequency information, PFFL significantly improves the quality of reconstructed images, especially in terms of texture and edge details.

[0191] Total Loss Function: The total loss function measures pixel-level differences not only in the spatial domain but also focuses on the recovery of high-frequency information in the frequency domain. This comprehensive consideration enables the model to better recover details and edge information while maintaining overall image quality.

[0192]

[0193] Here, α, β, γ, δ, and ρ are weighting coefficients (e.g., α = 0.65, β = 0.1, γ = 0.1, δ = 0.1, ρ = 0.05), used to balance the contribution of each factor to the total loss. The weights are dynamically adjusted, with α, β, and γ gradually decreasing and δ and ρ increasing in the later stages of training.

[0194] This embodiment also introduces meta-learning dynamic hyperparameter adjustment: different video scenarios (such as high-speed motion, static background) require differentiated optimization strategies. Fixed loss function weights and network hyperparameters (such as learning rate) cannot adapt to different input content. Therefore, a lightweight meta-learning network (such as a 3-layer MLP) is designed, with the statistical features of the current batch of data (such as gradient variance, loss distribution) as input.

[0195]

[0196] η = MetaNet(State(F) in ))

[0197] Where α, β, γ, δ, and ρ are the weights of the loss function, η is the learning rate, and F in These are the input features of the current batch. This represents the gradient of the current batch of data. MetaNet is a lightweight meta-learning network consisting of 3 layers of MLP, and State is the extracted state function. Through meta-learning, the loss weights and learning rate can be adjusted in real time, eliminating the need for manual adjustment of the loss function weights and hyperparameters. This improves SSIM and accelerates training convergence.

[0198] S3. On the client side, the keyframes are converted to 4K resolution using the 4K real-time rendering model to complete the rendering.

[0199] The features extracted from the encoding prior are F1, F′2 and F′3. The input keyframe is encoded in multiple layers. Each layer goes through multiple convolutional layers and an attention module. Then it is fused with the features of the encoding prior. Then it is input into a bidirectional pyramid for special fusion. Finally, it is output through an upsampling module.

[0200] Example 2

[0201] This embodiment also provides a high-efficiency cloud-based 4K real-time rendering system, including: a cloud module, a building module, and a client module; the cloud module is used to render the image in real time, then identify keyframes, encode and compress the keyframes, and send them to the client module; the building module is used to build a 4K real-time rendering model and download the 4K real-time rendering model to the client module; in the client module, the 4K real-time rendering model is used to convert the keyframes to 4K resolution to complete the rendering.

[0202] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for efficient cloud streaming of 4K real-time rendering, characterized by the following steps: include: The system renders the image in real time using the cloud, then identifies keyframes, encodes and compresses the keyframes, and sends them to the client. Build a 4K real-time rendering model and download the 4K real-time rendering model to the client; In the client, the 4K real-time rendering model is used to convert keyframes to 4K resolution to complete the rendering.

2. The efficient cloud streaming 4K real-time rendering method according to claim 1, characterized in that, The 4K real-time rendering model includes a prior feature extraction network and a backbone network; The prior feature extraction network is used to extract prior features from keyframes; The backbone network is used to render keyframes based on the extracted features.

3. The efficient cloud streaming 4K real-time rendering method according to claim 2, characterized in that, The prior feature extraction network includes a multi-convolutional layer structure, which consists of two 3×3 convolutional layers and a convolutional block containing two 1×1 convolutional layers and one 3×3 convolutional layer, to reduce computational cost. F Rep =Conv 3×3 (Conv 3×3 (F in )+Conv 1×1 _Conv 3×3 (F in )) Among them, F in F represents the input features, Conv represents the convolutional layer, and F represents the input features. Rep This indicates multiple convolutional layers.

4. The efficient cloud streaming 4K real-time rendering method according to claim 2, characterized in that, The prior feature extraction network further includes pixel-adaptive convolutional blocks, which are used to generate spatially adaptive convolutional kernels. The pixel-adaptive convolution formula is as follows: Where, p i ∈(x i ,x j ) T These are pixel coordinates, Ω(i) defines an s×s convolution window, and the image features v = (v1, v2, ..., v...). n ), Representing the features of each pixel, filter weights and bias v j Let v′ represent the feature vector at position j in the input feature map. i This represents the feature vector at position i in the output feature map; The convolution kernel weights of PAC are binarized to ±1, retaining only the sign information, and a scaling factor ∝ is introduced to compensate for quantization error: W′=∝·Sign(W) Where Sign is the binarization operation, W is the original weight of the PAC convolution kernel, W′ is the new weight of the PAC convolution kernel, and ∝ is the scaling factor.

5. The efficient cloud streaming 4K real-time rendering method according to claim 2, characterized in that, The backbone network includes: a downsampling module, an upsampling module, a channel attention module, and a bidirectional pyramid; The downsampling module is used to extract deep features; The upsampling module is used to restore spatial resolution; The channel attention module is used to enhance the expressive power of features and highlight important features; The bidirectional pyramid is used for feature fusion via a cross-scale feature pyramid.

6. The efficient cloud streaming 4K real-time rendering method according to claim 5, characterized in that, Global modeling of the downsampled low-resolution features: F trans =MobileViT(F down ) in, Indicates element-wise addition. MobileViT is a lightweight Transformer network module. trans F represents the feature map output by the MobileViT module. down F represents the feature map after the downsampling operation. skip This represents the skip connection feature map.

7. The efficient cloud streaming 4K real-time rendering method according to claim 5, characterized in that, The steps for feature fusion in the bidirectional pyramid include: First, a top-down path operation is performed to upsample deep, high-semantic features layer by layer and fuse them with shallow features. Simultaneously, a bottom-up path is used to downsample shallow, detailed features layer by layer and pass them to deeper layers to supplement local information. Then, dynamic weights are applied to each layer of fused features to allocate attention, suppress redundant information, and enhance key features. Finally, the dynamically weighted features at each scale are concatenated, and channel information is integrated through convolution to achieve the fusion of multi-scale features.

8. The efficient cloud streaming 4K real-time rendering method according to claim 5, characterized in that, The backbone network also incorporates a spatiotemporal attention module, the workflow of which includes: By aligning the feature maps of adjacent frames with the current frame through optical flow-guided feature alignment, and using a multi-head spatiotemporal attention mechanism, spatial local details and temporal global motion patterns are captured respectively. By fusing the current frame with spatiotemporal enhancement features through a learnable gating mechanism, dynamic detail modeling is enhanced and the stability of dynamic scene detail recovery in video sequences is improved.

9. The efficient cloud streaming 4K real-time rendering method according to claim 1, characterized in that, The 4K real-time rendering model is optimized using meta-learning dynamic hyperparameter tuning: η=MetaNet(State(F in )) Where α, β, γ, δ, and ρ are the weights of the loss function, η is the learning rate, and F in These are the input features of the current batch. It represents the gradient of the current batch of data. MetaNet is a lightweight meta-learning network consisting of 3 layers of MLP. State is the extracted state function.

10. A high-efficiency cloud streaming 4K real-time rendering system, the system being used to implement the method according to any one of claims 1-9, characterized in that, include: Cloud module, build module, and client module; The cloud module is used to render the image in real time, then identify key frames, encode and compress the key frames and send them to the client module. The building module is used to build a 4K real-time rendering model and download the 4K real-time rendering model to the client module; In the client module, the 4K real-time rendering model is used to convert keyframes to 4K resolution to complete the rendering.

Citation Information

Patent Citations

  • Home decoration roaming animation cloud rendering method and system based on specific path

    CN111640173A

  • Video transmission method, server, terminal and video transmission system

    CN114584805A

  • Video super-resolution reconstruction method and system

    CN120013766A

  • Lightweight embedded real-time image enhancement system and method

    CN120163751A

  • Lance asiabell leaf fermented ripening tea and preparation method therof

    KR1020250115658A