Mobile terminal high-quality image rendering method and system based on AI and cloud rendering
By introducing AI deep learning models and generative adversarial networks into cloud rendering solutions, intelligently partitioning and scheduling rendering tasks, the problems of high latency and bandwidth consumption in existing cloud rendering solutions are solved, and high-quality and efficient rendering effects are achieved.
Patent Information
- Application Number
- CN202411952978.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing cloud rendering solutions lack intelligent division and scheduling, resulting in high network latency and bandwidth consumption, affecting user experience, and making it difficult to achieve high-quality rendering effects, especially when the device performance is low or the network condition is poor.
The rendering scene is analyzed using AI-based deep learning model, scene feature data is generated, and rendering task evaluation model is built. The rendering task is divided into local real-time rendering and cloud-based high-quality rendering. Combined with generation adversarial networks and deep learning noise reduction processing, the rendering image data is optimized, and the rendering parameters are dynamically adjusted according to device performance and network conditions through adaptive rendering parameter adjustment strategies.
It significantly improves the quality and efficiency of rendering images on mobile terminals, reduces power consumption of mobile devices, avoids lag and delays, and provides a smoother rendering experience.
Smart Images

Figure CN119963709A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, and in particular to a method and system for high-quality image rendering on a mobile terminal based on AI and cloud rendering. Background Art
[0002] The graphics rendering capabilities of mobile devices are constantly improving, but due to the limitations of hardware resources and power consumption, it is still challenging to achieve high-quality, high-fidelity real-time rendering on mobile devices. Especially in dealing with complex scenes, global illumination, advanced materials, etc., the performance bottleneck of mobile devices is particularly prominent. In order to improve the rendering effect of mobile terminals, various optimization technologies have been widely studied and applied, such as LOD (level of detail)-based technology, occlusion culling technology, etc. The rise of cloud gaming and cloud rendering technology has provided new ideas for high-quality rendering on mobile terminals. By offloading complex rendering tasks to cloud servers, powerful server resources can be used for high-quality image rendering, and the rendering results can be transmitted to mobile devices for display.
[0003] Existing cloud rendering solutions usually lack intelligent division and scheduling of rendering tasks. Simply offloading all rendering tasks to the cloud will result in high network latency and bandwidth consumption, affecting user experience. Relying entirely on local rendering cannot fully utilize the advantages of cloud resources and it is difficult to achieve high-quality rendering effects. Existing cloud rendering solutions lack adaptability to the hardware performance and network conditions of mobile devices. In the case of poor network conditions or low device performance, the efficiency and stability of cloud rendering will be affected, resulting in problems such as screen freezes and delays. Existing image quality assessment methods often rely on manual evaluation or simple indicator calculations, which are difficult to accurately reflect the user's subjective feelings. This makes the adjustment of rendering parameters lack effective guidance and it is difficult to achieve the best rendering effect. Summary of the invention
[0004] The embodiments of the present invention provide a method and system for high-quality image rendering on a mobile terminal based on AI and cloud rendering, which can solve the problems in the prior art.
[0005] According to a first aspect of the embodiments of the present invention,
[0006] Provides a high-quality mobile image rendering method based on AI and cloud rendering, including:
[0007] Obtain a rendering scene request from a mobile terminal; use a deep learning model to analyze the rendering scene request to generate scene feature data, wherein the scene feature data includes geometric information features, material features, and lighting requirement features; construct a rendering task evaluation model based on the scene feature data, wherein the rendering task evaluation model divides the rendering task into a local real-time rendering task sequence and a cloud-based high-quality rendering task sequence; transmit the local real-time rendering task sequence, the cloud-based high-quality rendering task sequence, and the scene feature data to a task processing module;
[0008] The task processing module receives the local real-time rendering task sequence, the cloud high-quality rendering task sequence and the scene feature data; distributes the local real-time rendering task sequence to the mobile terminal processor, and performs frustum culling and scene update based on the geometric information features in the scene feature data; distributes the cloud high-quality rendering task sequence to the cloud rendering cluster, and performs global illumination calculation and material rendering according to the material features and lighting requirement features in the scene feature data; optimizes the cloud rendering results by using a generative adversarial network to generate high-quality rendering image data; performs deep learning-based denoising processing on the high-quality rendering image data to obtain optimized cloud rendering data;
[0009] The mobile terminal receives the optimized cloud rendering data, synthesizes the optimized cloud rendering data with the processing result of the local real-time rendering task to obtain an initial rendered image; establishes an image quality assessment model based on deep learning, the image quality assessment model receives the initial rendered image and generates image quality assessment data; constructs an adaptive rendering parameter adjustment strategy based on the image quality assessment data, the adaptive rendering parameter adjustment strategy dynamically adjusts the rendering parameters according to the device performance status and network conditions, and the rendering parameters include texture accuracy, rendering resolution and frame rate; applies the adjusted rendering parameters to the initial rendered image to generate a final rendered image; and feeds the image quality assessment data back to the scene analysis and preprocessing steps to optimize the generation of the next frame of scene feature data.
[0010] Building a rendering task evaluation model based on the scene feature data, wherein the rendering task evaluation model divides the rendering task into a local real-time rendering task sequence and a cloud high-quality rendering task sequence; transmitting the local real-time rendering task sequence, the cloud high-quality rendering task sequence and the scene feature data to a task processing module comprises:
[0011] Constructing a scene feature acquisition module, wherein the scene feature acquisition module acquires geometric information feature data, material feature data, and lighting requirement feature data in the scene; constructing a rendering task evaluation model based on the geometric information feature data, the material feature data, and the lighting requirement feature data, wherein the rendering task evaluation model includes a complexity evaluation unit and a device capability evaluation unit;
[0012] The complexity evaluation unit calculates a GPU load prediction value based on the geometric information feature data, calculates a memory occupancy prediction value based on the material feature data, and calculates a bandwidth demand prediction value based on the lighting demand feature data, and generates computing resource demand evaluation data;
[0013] The device capability evaluation unit collects GPU computing capability data, available memory capacity data, and network bandwidth data of the local device to generate device performance evaluation data;
[0014] Inputting the computing resource demand evaluation data and the device performance evaluation data into a task allocation decision unit, the task allocation decision unit generates a task allocation strategy, and the task allocation strategy divides the rendering task into a local real-time rendering task sequence and a cloud high-quality rendering task sequence;
[0015] The task allocation decision unit allocates a frustum culling task and a scene update task to the local real-time rendering task sequence, and allocates a global illumination calculation task and a complex material rendering task to the cloud high-quality rendering task sequence;
[0016] A task scheduling management module is constructed, and the task scheduling management module generates a task dependency graph, which includes task execution priorities and synchronization point strategies, and coordinates the execution of the local real-time rendering task sequence and the cloud high-quality rendering task sequence based on the task dependency graph.
[0017] Allocating the local real-time rendering task sequence to the mobile terminal processor, and performing frustum culling and scene update based on the geometric information features in the scene feature data; allocating the cloud high-quality rendering task sequence to the cloud rendering cluster, and performing global illumination calculation and material rendering according to the material features and lighting requirement features in the scene feature data, including:
[0018] Allocate the local real-time rendering task sequence to the mobile terminal processor, allocate the cloud high-quality rendering task sequence to the cloud rendering cluster, and generate a task allocation mapping table;
[0019] Based on the task allocation mapping table, a hierarchical bounding box tree structure is constructed on the mobile terminal processor, and the hierarchical bounding box tree structure is divided into nodes by a surface area heuristic algorithm according to the vertex coordinate data in the geometric information feature to generate scene space division data; according to the scene space division data, a frustum culling is performed on the mobile terminal processor, and the intersection relationship between the frustum plane and the node bounding box is calculated by using the separating axis theorem, and the visible object set data is output;
[0020] Based on the visible object set data, a scene graph structure is constructed in the mobile terminal processor, a transformation matrix of a node in the scene graph structure is updated according to the object transformation information, and the bounding box data of the corresponding node in the hierarchical bounding box tree structure is synchronously updated to generate scene state update data; the scene state update data is transmitted to the cloud rendering cluster, and the cloud rendering cluster extracts the material features according to the task allocation mapping table and constructs a material attribute descriptor;
[0021] Based on the material attribute descriptor, tasks are assigned to the graphics processing units in the cloud rendering cluster to generate a rendering resource allocation scheme, wherein the rendering resource allocation scheme determines a computing resource allocation weight according to material complexity;
[0022] According to the rendering resource allocation scheme and the scene status update data, global illumination calculation is performed in the cloud rendering cluster, and direct illumination components, one indirect illumination components and multiple reflection illumination components are calculated through a layered ray tracing module to generate global illumination rendering data; based on the global illumination rendering data and the material attribute descriptor, material rendering calculation is performed in the cloud rendering cluster.
[0023] The cloud rendering result is optimized by using a generative adversarial network to generate high-quality rendered image data; the high-quality rendered image data is subjected to a denoising process based on deep learning, and the optimized cloud rendering data obtained includes:
[0024] Constructing a generator of a generative adversarial network, wherein the generator adopts a U-Net architecture and includes an encoder module and a decoder module, wherein the encoder module extracts multi-level feature data of the cloud-rendered input image through five downsampling blocks, each of which includes two three-by-three convolutional layers, a batch normalization layer, and a LeakyReLU activation function layer, and the decoder module reconstructs enhanced feature data through five upsampling blocks and a residual connection module;
[0025] Constructing a discriminator of the generative adversarial network according to the multi-level feature data, wherein the discriminator adopts a PatchGAN structure, and divides the cloud-rendered input image into a plurality of overlapping local image blocks through five convolution modules, each of which includes a convolution layer, an instance normalization layer, and a LeakyReLU activation function layer to generate a discriminant feature map;
[0026] Calculate the adversarial loss function value based on the discriminant feature map, calculate the content loss function value and the perceptual loss function value based on the enhanced feature data, and use the adversarial loss function value, the content loss function value and the perceptual loss function value as training targets to optimize the network parameters of the generative adversarial network to generate initial optimized image data;
[0027] Constructing a multi-scale denoising network according to the initial optimized image data, wherein the multi-scale denoising network comprises a feature extraction sub-network, a noise detection sub-network and an image reconstruction sub-network, and the feature extraction sub-network extracts multi-scale feature representation data of the initial optimized image data through a plurality of residual blocks;
[0028] Building an adaptive noise estimation module based on the multi-scale feature representation data, calculating mean data, variance data and gradient data of the local image block based on the adaptive noise estimation module, predicting the noise intensity value of each local image block through a multi-layer perceptron, and generating a noise intensity distribution map;
[0029] The noise intensity distribution map is input into the noise detection subnetwork, the noise detection subnetwork generates a noise distribution feature map through an attention mechanism module, and the image reconstruction subnetwork fuses the multi-scale feature representation data and the noise distribution feature map using a dense connection structure;
[0030] Based on the multi-scale feature representation data and the noise distribution feature map, a mean square error loss function value, a structural similarity loss function value and a perceptual loss function value are calculated, and the mean square error loss function value, the structural similarity loss function value and the perceptual loss function value are used as training targets to optimize the network parameters of the multi-scale denoising network, and generate optimized cloud rendering data.
[0031] The mobile terminal receives the optimized cloud rendering data, synthesizes the optimized cloud rendering data with the processing result of the local real-time rendering task to obtain an initial rendered image; establishes an image quality assessment model based on deep learning, the image quality assessment model receives the initial rendered image, and generates image quality assessment data including:
[0032] The mobile terminal receives the optimized cloud rendering data, obtains texture feature data, geometric detail data and lighting distribution data in the optimized cloud rendering data; collects the processing result of the local real-time rendering task, and extracts scene depth data, material attribute data and real-time lighting data in the processing result;
[0033] Constructing a feature extraction network, wherein the feature extraction network includes a cloud feature branch and a local feature branch, each of the feature branches includes four layers of hole convolution layers, respectively extracting multi-scale features from the optimized cloud rendering data and the processing results of the local real-time rendering task, and generating a cloud feature map and a local feature map;
[0034] A feature fusion network is constructed based on the cloud feature map and the local feature map, wherein the feature fusion network calculates a spatial attention map through a non-local module, calculates a channel attention map through an SE structure, and weightedly fuses the spatial attention map and the channel attention map with the original feature map to generate a fused feature map;
[0035] Constructing a progressive upsampling network according to the fused feature map, wherein the progressive upsampling network comprises a deconvolution module and a residual enhancement module, wherein the residual enhancement module comprises four residual blocks, and the feature map output by each residual block is processed by a PReLU activation function to generate an initial rendered image;
[0036] Constructing a dual-stream quality assessment network, wherein the global feature stream of the dual-stream quality assessment network extracts high-level semantic features of the initial rendered image through a ResNet-50 network, and the local feature stream of the dual-stream quality assessment network extracts local texture features of the initial rendered image through a four-layer dense connection block to generate global feature data and local feature data;
[0037] Calculating a saliency map based on the global feature data and the local feature data, wherein the saliency map indicates a weight distribution of a visual attention area in the initial rendered image, and weighting the global feature data and the local feature data according to the saliency map to generate weighted feature data;
[0038] Inputting the weighted feature data into a feature enhancement network, wherein the feature enhancement network comprises a self-attention module and a feature pyramid module, wherein the self-attention module calculates the correlation between feature points, and the feature pyramid module fuses feature representations of different scales to generate enhanced feature data;
[0039] A quality assessment module is constructed based on the enhanced feature data, and the quality assessment module calculates the brightness mean, contrast variance and structural similarity of the image block, and outputs image quality assessment data through a fully connected layer.
[0040] An adaptive rendering parameter adjustment strategy is constructed based on the image quality evaluation data, and the adaptive rendering parameter adjustment strategy dynamically adjusts rendering parameters according to device performance status and network conditions, and the rendering parameters include texture accuracy, rendering resolution and frame rate.
[0041] The image quality assessment data includes an overall quality score, a detail fidelity score, and a quality distribution map, and an adaptive rendering parameter adjustment strategy model based on deep reinforcement learning is constructed, wherein the state space of the adaptive rendering parameter adjustment strategy model includes the image quality assessment data, device performance status data, and network status data, and the action space includes an adjustment plan for rendering parameters;
[0042] Collecting the device performance status data, the device performance status data including central processing unit utilization, graphics processing unit utilization, and memory occupancy, normalizing the device performance status data and inputting the data into the adaptive rendering parameter adjustment strategy model;
[0043] Monitoring the network status data, wherein the network status data includes network delay time, network bandwidth data, and data packet loss rate, and normalizing the network status data and inputting it into the adaptive rendering parameter adjustment strategy model;
[0044] The adaptive rendering parameter adjustment strategy model generates a rendering parameter adjustment plan based on the input image quality evaluation data, the device performance status data, and the network status data. The rendering parameter adjustment plan includes a texture accuracy adjustment strategy, a rendering resolution adjustment strategy, and a frame rate adjustment strategy.
[0045] According to a second aspect of the embodiments of the present invention,
[0046] Provides a high-quality mobile image rendering system based on AI and cloud rendering, including:
[0047] The first unit is used to obtain a rendering scene request from a mobile terminal; use a deep learning model to analyze the rendering scene request to generate scene feature data, wherein the scene feature data includes geometric information features, material features, and lighting requirement features; construct a rendering task evaluation model based on the scene feature data, wherein the rendering task evaluation model divides the rendering task into a local real-time rendering task sequence and a cloud-based high-quality rendering task sequence; and transmit the local real-time rendering task sequence, the cloud-based high-quality rendering task sequence, and the scene feature data to a task processing module;
[0048] The second unit is used for the task processing module to receive the local real-time rendering task sequence, the cloud high-quality rendering task sequence and the scene feature data; distribute the local real-time rendering task sequence to the mobile terminal processor, and perform frustum culling and scene update based on the geometric information features in the scene feature data; distribute the cloud high-quality rendering task sequence to the cloud rendering cluster, and perform global illumination calculation and material rendering according to the material features and lighting requirement features in the scene feature data; optimize the cloud rendering results by using a generative adversarial network to generate high-quality rendering image data; perform denoising based on deep learning on the high-quality rendering image data to obtain optimized cloud rendering data;
[0049] The third unit is used for receiving the optimized cloud rendering data on the mobile terminal, synthesizing the optimized cloud rendering data with the processing result of the local real-time rendering task to obtain an initial rendered image; establishing an image quality assessment model based on deep learning, the image quality assessment model receives the initial rendered image and generates image quality assessment data; constructing an adaptive rendering parameter adjustment strategy based on the image quality assessment data, the adaptive rendering parameter adjustment strategy dynamically adjusts the rendering parameters according to the device performance status and network conditions, and the rendering parameters include texture accuracy, rendering resolution and frame rate; applying the adjusted rendering parameters to the initial rendered image to generate a final rendered image; and feeding back the image quality assessment data to the scene analysis and preprocessing steps to optimize the generation of the next frame of scene feature data.
[0050] According to a third aspect of the embodiments of the present invention,
[0051] An electronic device is provided, comprising:
[0052] processor;
[0053] a memory for storing processor-executable instructions;
[0054] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0055] A fourth aspect of the embodiments of the present invention is:
[0056] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0057] The beneficial effects of this application are as follows:
[0058] 1. Improve the quality of mobile rendering images: Perform global illumination calculation and material rendering through cloud rendering clusters, and combine generative adversarial networks and deep learning denoising processing to generate high-quality rendering image data, significantly improving the visual effects of mobile rendering images.
[0059] 2. Optimize rendering efficiency and performance: Analyze rendering scenes based on deep learning models, build rendering task evaluation models, and reasonably distribute rendering tasks to mobile devices and the cloud, effectively balancing rendering efficiency and performance. At the same time, the adaptive rendering parameter adjustment strategy dynamically adjusts rendering parameters according to device performance and network conditions, further optimizing rendering performance and avoiding freezes and delays.
[0060] 3. Reduce mobile rendering power consumption: Offloading complex rendering tasks to the cloud reduces the computing burden on the mobile terminal, thereby reducing the power consumption of mobile devices and extending battery life. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 A schematic diagram of a process of a high-quality image rendering method for a mobile terminal based on AI and cloud rendering according to an embodiment of the present invention;
[0062] Figure 2 This is a structural diagram of a high-quality image rendering system for a mobile terminal based on AI and cloud rendering according to an embodiment of the present invention. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0064] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0065] Figure 1 FIG. 1 is a flow chart of a method for high-quality image rendering on a mobile terminal based on AI and cloud rendering according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0066] S11. Obtain a rendering scene request from a mobile terminal; use a deep learning model to analyze the rendering scene request to generate scene feature data, wherein the scene feature data includes geometric information features, material features, and lighting requirement features; construct a rendering task evaluation model based on the scene feature data, wherein the rendering task evaluation model divides the rendering task into a local real-time rendering task sequence and a cloud-based high-quality rendering task sequence; transmit the local real-time rendering task sequence, the cloud-based high-quality rendering task sequence, and the scene feature data to a task processing module;
[0067] S12. The task processing module receives the local real-time rendering task sequence, the cloud high-quality rendering task sequence and the scene feature data; distributes the local real-time rendering task sequence to the mobile terminal processor, and performs frustum culling and scene update based on the geometric information features in the scene feature data; distributes the cloud high-quality rendering task sequence to the cloud rendering cluster, and performs global illumination calculation and material rendering according to the material features and lighting requirement features in the scene feature data; optimizes the cloud rendering results using a generative adversarial network to generate high-quality rendering image data; performs deep learning-based denoising on the high-quality rendering image data to obtain optimized cloud rendering data;
[0068] S13. The mobile terminal receives the optimized cloud rendering data, synthesizes the optimized cloud rendering data with the processing result of the local real-time rendering task to obtain an initial rendered image; establishes an image quality assessment model based on deep learning, the image quality assessment model receives the initial rendered image and generates image quality assessment data; constructs an adaptive rendering parameter adjustment strategy based on the image quality assessment data, the adaptive rendering parameter adjustment strategy dynamically adjusts the rendering parameters according to the device performance status and network conditions, the rendering parameters including texture accuracy, rendering resolution and frame rate; applies the adjusted rendering parameters to the initial rendered image to generate a final rendered image; feeds the image quality assessment data back to the scene analysis and preprocessing steps to optimize the generation of the next frame of scene feature data.
[0069] In an optional implementation, a rendering task evaluation model is constructed based on the scene feature data, wherein the rendering task evaluation model divides the rendering task into a local real-time rendering task sequence and a cloud high-quality rendering task sequence; and transmitting the local real-time rendering task sequence, the cloud high-quality rendering task sequence and the scene feature data to a task processing module comprises:
[0070] Constructing a scene feature acquisition module, wherein the scene feature acquisition module acquires geometric information feature data, material feature data, and lighting requirement feature data in the scene; constructing a rendering task evaluation model based on the geometric information feature data, the material feature data, and the lighting requirement feature data, wherein the rendering task evaluation model includes a complexity evaluation unit and a device capability evaluation unit;
[0071] The complexity evaluation unit calculates a GPU load prediction value based on the geometric information feature data, calculates a memory occupancy prediction value based on the material feature data, and calculates a bandwidth demand prediction value based on the lighting demand feature data, and generates computing resource demand evaluation data;
[0072] The device capability evaluation unit collects GPU computing capability data, available memory capacity data, and network bandwidth data of the local device to generate device performance evaluation data;
[0073] Inputting the computing resource demand evaluation data and the device performance evaluation data into a task allocation decision unit, the task allocation decision unit generates a task allocation strategy, and the task allocation strategy divides the rendering task into a local real-time rendering task sequence and a cloud high-quality rendering task sequence;
[0074] The task allocation decision unit allocates a frustum culling task and a scene update task to the local real-time rendering task sequence, and allocates a global illumination calculation task and a complex material rendering task to the cloud high-quality rendering task sequence;
[0075] A task scheduling management module is constructed, and the task scheduling management module generates a task dependency graph, which includes task execution priorities and synchronization point strategies, and coordinates the execution of the local real-time rendering task sequence and the cloud high-quality rendering task sequence based on the task dependency graph.
[0076] An intelligent rendering task allocation method can intelligently allocate rendering tasks to local devices or cloud servers based on scene characteristics and device capabilities to achieve optimal rendering effects and performance.
[0077] First, build a scene feature acquisition module. This module is responsible for collecting geometric information feature data, material feature data, and lighting requirement feature data in the scene. Geometric information feature data includes the number of objects in the scene, the number of vertices, the number of faces, etc. For example, a scene contains 1,000 objects, and each object contains an average of 5,000 vertices and 8,000 faces. Material feature data includes material type, texture resolution, lighting model, etc. For example, an object uses PBR material, the texture resolution is 4096x4096, and the lighting model is GGX. Lighting requirement feature data includes the number of light sources, lighting type, shadow quality, etc. For example, the scene contains 3 directional lights, 1 ambient light, and the shadow quality is set to high.
[0078] Next, a rendering task evaluation model is constructed based on the collected scene feature data. The model includes a complexity evaluation unit and a device capability evaluation unit. The complexity evaluation unit calculates the GPU load prediction value, the memory occupancy prediction value, and the bandwidth requirement prediction value based on the geometric information feature data, the material feature data, and the lighting requirement feature data, and generates computing resource requirement evaluation data. For example, according to the above scene feature data, the predicted GPU load is 80%, the memory occupancy is 12GB, and the bandwidth requirement is 100Mbps. The device capability evaluation unit collects the GPU computing capability data, available memory capacity data, and network bandwidth data of the local device to generate device performance evaluation data. For example, the GPU computing capability of the local device is 10TFLOPS, the available memory capacity is 16GB, and the network bandwidth is 200Mbps.
[0079] Then, the computing resource demand assessment data and the device performance assessment data are input into the task allocation decision unit. The task allocation decision unit allocates the rendering task to the cloud according to the preset strategy. For example, if the GPU load prediction value exceeds 70% of the local device GPU computing power, or the memory occupancy prediction value exceeds 90% of the local device's available memory capacity, the rendering task is allocated to the local device. According to the above case data, since the GPU load prediction value of 80% exceeds the threshold of 70% of the local device GPU computing power, the rendering task is allocated to the cloud. The task allocation decision unit divides the rendering tasks into local real-time rendering task sequences and cloud high-quality rendering task sequences.
[0080] After the task allocation strategy is determined, the task allocation decision unit allocates frustum culling tasks and scene update tasks to the local real-time rendering task sequence, and allocates global illumination calculation tasks and complex material rendering tasks to the cloud high-quality rendering task sequence. For example, the local real-time rendering task sequence performs frustum culling and scene update to ensure the rendering frame rate; the cloud high-quality rendering task sequence performs global illumination calculation and complex material rendering to generate high-quality rendering results.
[0081] Finally, build a task scheduling management module. The task scheduling management module generates a task dependency graph, which contains task execution priorities and synchronization point strategies. For example, the global illumination calculation task needs to be completed before the complex material rendering task, and the frustum culling task needs to be completed before the scene update task. The task scheduling management module coordinates the execution of the local real-time rendering task sequence and the cloud high-quality rendering task sequence based on the task dependency graph to ensure that the tasks are executed in the correct order and perform necessary synchronization. For example, after the local device completes the frustum culling, the culling result is synchronized to the cloud for the cloud to perform subsequent rendering calculations.
[0082] The solution of this application can:
[0083] Improve rendering efficiency: By assigning computing-intensive tasks to the cloud, the powerful computing resources of the cloud can be fully utilized to speed up rendering and shorten rendering time. Optimize resource utilization: Dynamically assign tasks according to scene complexity and device capabilities to avoid overloading local devices, while making full use of cloud resources to achieve reasonable allocation and efficient use of resources. Enhance rendering effects: Cloud servers can handle complex rendering tasks, such as global illumination and high-quality material rendering, thereby improving the quality and realism of the final rendered image.
[0084] In an optional implementation, the local real-time rendering task sequence is assigned to a mobile terminal processor, and frustum culling and scene updating are performed based on the geometric information features in the scene feature data; the cloud high-quality rendering task sequence is assigned to a cloud rendering cluster, and global illumination calculation and material rendering are performed based on the material features and lighting requirement features in the scene feature data, including:
[0085] Allocate the local real-time rendering task sequence to the mobile terminal processor, allocate the cloud high-quality rendering task sequence to the cloud rendering cluster, and generate a task allocation mapping table;
[0086] Based on the task allocation mapping table, a hierarchical bounding box tree structure is constructed on the mobile terminal processor, and the hierarchical bounding box tree structure is divided into nodes by a surface area heuristic algorithm according to the vertex coordinate data in the geometric information feature to generate scene space division data; according to the scene space division data, a frustum culling is performed on the mobile terminal processor, and the intersection relationship between the frustum plane and the node bounding box is calculated by using the separating axis theorem, and the visible object set data is output;
[0087] Based on the visible object set data, a scene graph structure is constructed in the mobile terminal processor, a transformation matrix of a node in the scene graph structure is updated according to the object transformation information, and the bounding box data of the corresponding node in the hierarchical bounding box tree structure is synchronously updated to generate scene state update data; the scene state update data is transmitted to the cloud rendering cluster, and the cloud rendering cluster extracts the material features according to the task allocation mapping table and constructs a material attribute descriptor;
[0088] Based on the material attribute descriptor, tasks are assigned to the graphics processing units in the cloud rendering cluster to generate a rendering resource allocation scheme, wherein the rendering resource allocation scheme determines a computing resource allocation weight according to material complexity;
[0089] According to the rendering resource allocation scheme and the scene status update data, global illumination calculation is performed in the cloud rendering cluster, and direct illumination components, one indirect illumination components and multiple reflection illumination components are calculated through a layered ray tracing module to generate global illumination rendering data; based on the global illumination rendering data and the material attribute descriptor, material rendering calculation is performed in the cloud rendering cluster.
[0090] The mobile terminal and the cloud-side collaborative rendering method achieves high-quality, low-latency rendering effects by assigning real-time rendering tasks to the mobile terminal and high-quality rendering tasks to the cloud.
[0091] First, generate a task allocation mapping table. Divide the rendering tasks in the scene into local real-time rendering task sequences and cloud-based high-quality rendering task sequences. The local real-time rendering task sequence includes tasks with high real-time requirements such as frustum culling and scene updates, which are assigned to the mobile processor for execution. The cloud-based high-quality rendering task sequence includes tasks with high rendering quality requirements such as global illumination calculation and material rendering, which are assigned to the cloud-based rendering cluster for execution. Generate a task allocation mapping table to record the execution location of each rendering task (mobile or cloud). For example, tasks such as updating the transformation matrix of objects in the scene and frustum culling are assigned to the mobile terminal, and tasks such as global illumination calculation and physically-based rendering are assigned to the cloud.
[0092] Then, a hierarchical bounding box tree structure is constructed on the mobile processor. According to the geometric information features in the scene feature data, especially the vertex coordinate data, the scene is spatially divided using a surface area heuristic algorithm to construct a hierarchical bounding box tree structure. For example, a scene is divided into eight subspaces, and each subspace is recursively divided until the number of objects contained in each leaf node is less than a preset threshold. Each node is represented by a bounding box, which contains the geometry of all child nodes of the node. Assuming there are 1,000 triangular patches in the scene, the depth of the constructed hierarchical bounding box tree structure is 5, and the leaf nodes contain an average of 10 triangular patches.
[0093] Next, perform frustum culling on the mobile terminal. According to the scene space division data, use the separating axis theorem to determine the intersection relationship between the frustum plane and the node bounding box. If the bounding box does not intersect the frustum, the node and all its child nodes are culled. If the bounding box intersects the frustum, continue to determine the intersection relationship between its child nodes and the frustum until the leaf node. Finally, the visible object set data is output. For example, after frustum culling, there are 200 triangle patches left in the scene that need to be rendered.
[0094] Subsequently, the scene graph structure is constructed on the mobile terminal and the scene state is updated. Based on the visible object set data, the scene graph structure is constructed to represent the hierarchical relationship between objects in the scene. According to the object transformation information, the transformation matrix of the node in the scene graph structure is updated, and the bounding box data of the corresponding node in the hierarchical bounding box tree structure is synchronously updated. These updated scene state data are transmitted to the cloud rendering cluster. For example, the translation, rotation, scaling and other transformation information of the objects in the scene are updated to the scene graph and the hierarchical bounding box tree.
[0095] After receiving the scene status update data, the cloud rendering cluster extracts material features according to the task allocation mapping table and constructs material attribute descriptors. For example, it extracts the diffuse color, specular color, roughness and other parameters of the objects in the scene to construct material attribute descriptors.
[0096] After that, the cloud rendering cluster allocates tasks to the graphics processing unit according to the material attribute descriptor and generates a rendering resource allocation plan. The computing resource allocation weight is determined according to the material complexity, and computing resources are allocated preferentially to objects with complex materials. For example, more computing resources are allocated to objects with complex lighting models and textures.
[0097] Next, the cloud rendering cluster performs global illumination calculations based on the rendering resource allocation plan and scene status update data. The layered ray tracing module calculates the direct illumination component, the primary indirect illumination component, and the multi-reflection illumination component to generate global illumination rendering data. For example, it simulates the propagation path of light in the scene and calculates the illumination intensity of each pixel.
[0098] Finally, the cloud rendering cluster performs material rendering calculations based on the global illumination rendering data and material attribute descriptors. For example, the global illumination data is combined with the material attributes to calculate the final color value of each pixel.
[0099] The solution of this application can:
[0100] Improve rendering efficiency: offload computationally intensive rendering tasks to the cloud, free up computing resources on the mobile end, and improve rendering efficiency. Enhance rendering quality: use the powerful computing power of the cloud to perform high-quality global illumination calculations and material rendering, significantly improving rendering effects. Reduce rendering latency: balance the distribution of rendering tasks through the collaborative work of the mobile end and the cloud end, reduce overall rendering latency, and achieve a smooth rendering experience.
[0101] In an optional implementation, a generative adversarial network is used to optimize cloud rendering results to generate high-quality rendered image data; a deep learning-based denoising process is performed on the high-quality rendered image data to obtain optimized cloud rendering data, including:
[0102] Constructing a generator of a generative adversarial network, wherein the generator adopts a U-Net architecture and includes an encoder module and a decoder module, wherein the encoder module extracts multi-level feature data of the cloud-rendered input image through five downsampling blocks, each of which includes two three-by-three convolutional layers, a batch normalization layer, and a LeakyReLU activation function layer, and the decoder module reconstructs enhanced feature data through five upsampling blocks and a residual connection module;
[0103] Constructing a discriminator of the generative adversarial network according to the multi-level feature data, wherein the discriminator adopts a PatchGAN structure, and divides the cloud-rendered input image into a plurality of overlapping local image blocks through five convolution modules, each of which includes a convolution layer, an instance normalization layer, and a LeakyReLU activation function layer to generate a discriminant feature map;
[0104] Calculate the adversarial loss function value based on the discriminant feature map, calculate the content loss function value and the perceptual loss function value based on the enhanced feature data, and use the adversarial loss function value, the content loss function value and the perceptual loss function value as training targets to optimize the network parameters of the generative adversarial network to generate initial optimized image data;
[0105] Constructing a multi-scale denoising network according to the initial optimized image data, wherein the multi-scale denoising network comprises a feature extraction sub-network, a noise detection sub-network and an image reconstruction sub-network, and the feature extraction sub-network extracts multi-scale feature representation data of the initial optimized image data through a plurality of residual blocks;
[0106] Building an adaptive noise estimation module based on the multi-scale feature representation data, calculating mean data, variance data and gradient data of the local image block based on the adaptive noise estimation module, predicting the noise intensity value of each local image block through a multi-layer perceptron, and generating a noise intensity distribution map;
[0107] The noise intensity distribution map is input into the noise detection subnetwork, the noise detection subnetwork generates a noise distribution feature map through an attention mechanism module, and the image reconstruction subnetwork fuses the multi-scale feature representation data and the noise distribution feature map using a dense connection structure;
[0108] Based on the multi-scale feature representation data and the noise distribution feature map, a mean square error loss function value, a structural similarity loss function value and a perceptual loss function value are calculated, and the mean square error loss function value, the structural similarity loss function value and the perceptual loss function value are used as training targets to optimize the network parameters of the multi-scale denoising network, and generate optimized cloud rendering data.
[0109] The cloud-rendered image optimization method aims to improve the quality of cloud-rendered images. Its core idea is to combine generative adversarial networks and deep learning denoising technology. This method mainly consists of two stages: image enhancement based on generative adversarial networks and image denoising based on deep learning.
[0110] First, the image enhancement stage is carried out. A generative adversarial network (GAN) is constructed, which consists of two parts: the generator and the discriminator. The generator adopts the U-Net architecture, and its encoder module contains five downsampling blocks. Each downsampling block consists of two 3x3 convolutional layers, a batch normalization layer, and a LeakyReLU activation function layer. Through these downsampling blocks, multi-level feature data can be extracted from the input cloud-rendered image. The decoder module contains five upsampling blocks and a residual connection module to reconstruct and enhance the extracted feature data and finally generate an enhanced image. The discriminator adopts the PatchGAN structure, which divides the input image into multiple overlapping local image blocks through five convolutional modules. Each convolutional module contains a convolutional layer, an instance normalization layer, and a LeakyReLU activation function layer. The discriminator discriminates each image block and generates a discriminant feature map.
[0111] Next, the enhanced feature data generated by the generator is used to calculate the content loss and perceptual loss, and the adversarial loss is calculated in combination with the discriminative feature map generated by the discriminator. The values of these three loss functions are used as training targets to optimize the parameters of the generative adversarial network through the back-propagation algorithm. For example, the content loss can be calculated using the sum of squares of pixel-by-pixel differences, the perceptual loss can be calculated using the feature differences extracted by the pre-trained image classification network, and the adversarial loss is used to measure the similarity between the generated image and the real image. By minimizing these loss functions, the generator can generate more realistic and high-quality images. The output of this stage is the initial optimized image data optimized by GAN.
[0112] Then we enter the image denoising stage. A multi-scale denoising network is constructed based on the initial optimized image data. The network consists of three parts: feature extraction sub-network, noise detection sub-network and image reconstruction sub-network. The feature extraction sub-network uses multiple residual blocks to extract multi-scale feature representation data of the initial optimized image data.
[0113] Next, the extracted multi-scale feature representation data is used to construct an adaptive noise estimation module. This module calculates statistics such as the mean, variance, and gradient of each local image block. For example, the image can be segmented into 16x16 image blocks, and the mean and variance of each image block are calculated. These statistics are then input into a multi-layer perceptron to predict the noise intensity value of each local image block, and finally a noise intensity distribution map is generated.
[0114] The generated noise intensity distribution map is input into the noise detection subnetwork. The noise detection subnetwork uses the attention mechanism module to generate a noise distribution feature map to highlight the characteristics of the noise area. The image reconstruction subnetwork adopts a dense connection structure to fuse multi-scale feature representation data and noise distribution feature maps to generate the final optimized cloud rendering data.
[0115] Finally, the parameters of the multi-scale denoising network are optimized using mean square error loss, structural similarity loss, and perceptual loss as training objectives. For example, mean square error loss calculates the square difference of pixel values between the predicted image and the target image, structural similarity loss measures the similarity of image structures, and perceptual loss is based on the difference in features extracted by the pre-trained network. By minimizing these loss functions, the denoising network can effectively remove noise and preserve image details.
[0116] The solution of this application can:
[0117] Improve image quality: By combining generative adversarial networks and deep learning denoising technology, noise and artifacts in cloud-rendered images can be effectively removed, improving image clarity, detail, and realism. Optimize rendering efficiency: Through pre-processing and post-processing technologies, the amount of computation and time of cloud-rendering can be reduced, improving rendering efficiency. Enhance user experience: High-quality rendered images can provide a better visual experience and enhance user immersion and satisfaction.
[0118] In an optional implementation, the mobile terminal receives the optimized cloud rendering data, synthesizes the optimized cloud rendering data with the processing result of the local real-time rendering task to obtain an initial rendered image; establishes an image quality assessment model based on deep learning, the image quality assessment model receives the initial rendered image, and generates image quality assessment data including:
[0119] The mobile terminal receives the optimized cloud rendering data, obtains texture feature data, geometric detail data and lighting distribution data in the optimized cloud rendering data; collects the processing result of the local real-time rendering task, and extracts scene depth data, material attribute data and real-time lighting data in the processing result;
[0120] Constructing a feature extraction network, wherein the feature extraction network includes a cloud feature branch and a local feature branch, each of the feature branches includes four layers of hole convolution layers, respectively extracting multi-scale features from the optimized cloud rendering data and the processing results of the local real-time rendering task, and generating a cloud feature map and a local feature map;
[0121] A feature fusion network is constructed based on the cloud feature map and the local feature map, wherein the feature fusion network calculates a spatial attention map through a non-local module, calculates a channel attention map through an SE structure, and weightedly fuses the spatial attention map and the channel attention map with the original feature map to generate a fused feature map;
[0122] Constructing a progressive upsampling network according to the fused feature map, wherein the progressive upsampling network comprises a deconvolution module and a residual enhancement module, wherein the residual enhancement module comprises four residual blocks, and the feature map output by each residual block is processed by a PReLU activation function to generate an initial rendered image;
[0123] Constructing a dual-stream quality assessment network, wherein the global feature stream of the dual-stream quality assessment network extracts high-level semantic features of the initial rendered image through a ResNet-50 network, and the local feature stream of the dual-stream quality assessment network extracts local texture features of the initial rendered image through a four-layer dense connection block to generate global feature data and local feature data;
[0124] Calculating a saliency map based on the global feature data and the local feature data, wherein the saliency map indicates a weight distribution of a visual attention area in the initial rendered image, and weighting the global feature data and the local feature data according to the saliency map to generate weighted feature data;
[0125] Inputting the weighted feature data into a feature enhancement network, wherein the feature enhancement network comprises a self-attention module and a feature pyramid module, wherein the self-attention module calculates the correlation between feature points, and the feature pyramid module fuses feature representations of different scales to generate enhanced feature data;
[0126] A quality assessment module is constructed based on the enhanced feature data, and the quality assessment module calculates the brightness mean, contrast variance and structural similarity of the image block, and outputs image quality assessment data through a fully connected layer.
[0127] The mobile cloud rendering image quality assessment method aims to improve the quality assessment accuracy and efficiency of mobile cloud rendering images. This method integrates cloud rendering data and local real-time rendering results, and uses a deep learning model to perform image quality assessment.
[0128] First, the mobile terminal receives optimized cloud rendering data from the cloud server, which includes texture feature data, geometric detail data, and lighting distribution data. For example, texture feature data can be the pixel value of the texture image, geometric detail data can be the vertex coordinates and normal vectors of the three-dimensional model, and lighting distribution data can be the position, color, and intensity of the light source. At the same time, the mobile terminal also collects the processing results of the local real-time rendering task, extracts scene depth data, material attribute data, and real-time lighting data. For example, scene depth data can be the distance from each pixel to the camera, material attribute data can be the reflectivity and roughness of the material, and real-time lighting data can be the real-time calculation results of the light in the scene.
[0129] Next, a feature extraction network is constructed. The network consists of two branches: a cloud feature branch and a local feature branch. Each branch uses four layers of dilated convolutional layers to extract multi-scale features from cloud rendering data and local real-time rendering results, respectively, to generate cloud feature maps and local feature maps. Dilated convolution can expand the receptive field without increasing the number of parameters, thereby better capturing the global information of the image. Assume that the size of the cloud feature map is 256x256x64 and the size of the local feature map is 256x256x64.
[0130] Then, a feature fusion network is constructed. This network fuses the cloud feature map and the local feature map. It first uses the non-local module to calculate the spatial attention map to capture the relationship between different positions in the feature map. Then the SE structure is used to calculate the channel attention map to capture the relationship between different channels. Finally, the spatial attention map and the channel attention map are weighted fused with the original feature map to generate a fused feature map. Assume that the size of the fused feature map is 256x256x128.
[0131] Next, a progressive upsampling network is constructed. This network upsamples the fused feature maps to generate the initial rendered image. It uses a deconvolution module and a residual enhancement module for upsampling. The residual enhancement module contains four residual blocks, and the feature map output by each residual block is processed by the PReLU activation function. Assume that the size of the initial rendered image is 1024x1024x3.
[0132] Then, a two-stream quality assessment network is constructed. The network contains a global feature stream and a local feature stream. The global feature stream uses the ResNet-50 network to extract high-level semantic features of the initial rendered image. The local feature stream uses a four-layer dense connection block to extract local texture features of the initial rendered image. Finally, global feature data and local feature data are generated. Assume that the dimension of the global feature data is 2048 and the dimension of the local feature data is 1024.
[0133] Next, a saliency map is calculated based on the global feature data and the local feature data. The saliency map indicates the weight distribution of the visual attention area in the initial rendered image. Then, the global feature data and the local feature data are weighted according to the saliency map to generate weighted feature data. Assume that the dimension of the weighted feature data is 3072.
[0134] After that, the weighted feature data is input into the feature enhancement network. The network contains a self-attention module and a feature pyramid module. The self-attention module calculates the correlation between feature points. The feature pyramid module fuses feature representations of different scales to generate enhanced feature data. Assume that the dimension of the enhanced feature data is 4096.
[0135] Finally, a quality assessment module is constructed based on the enhanced feature data. This module calculates the brightness mean, contrast variance, and structural similarity of the image blocks and outputs image quality assessment data through a fully connected layer. For example, the image quality assessment data can be a score between 0 and 1, indicating the quality level of the image.
[0136] The solution of this application can:
[0137] Improved the accuracy of image quality assessment. By integrating cloud rendering data and local real-time rendering results, image quality can be evaluated more comprehensively. Improved the efficiency of image quality assessment. By using deep learning models, image features can be automatically extracted and quality assessment can be performed without human intervention. Enhanced the robustness of image quality assessment. By using multi-scale features and attention mechanisms, image quality assessment under different scenes and lighting conditions can be better handled.
[0138] In an optional implementation, an adaptive rendering parameter adjustment strategy is constructed based on the image quality evaluation data, and the adaptive rendering parameter adjustment strategy dynamically adjusts rendering parameters according to device performance status and network conditions, and the rendering parameters include texture accuracy, rendering resolution and frame rate.
[0139] The image quality assessment data includes an overall quality score, a detail fidelity score, and a quality distribution map, and an adaptive rendering parameter adjustment strategy model based on deep reinforcement learning is constructed, wherein the state space of the adaptive rendering parameter adjustment strategy model includes the image quality assessment data, device performance status data, and network status data, and the action space includes an adjustment plan for rendering parameters;
[0140] Collecting the device performance status data, the device performance status data including central processing unit utilization, graphics processing unit utilization, and memory occupancy, normalizing the device performance status data and inputting the data into the adaptive rendering parameter adjustment strategy model;
[0141] Monitoring the network status data, wherein the network status data includes network delay time, network bandwidth data, and data packet loss rate, and normalizing the network status data and inputting it into the adaptive rendering parameter adjustment strategy model;
[0142] The adaptive rendering parameter adjustment strategy model generates a rendering parameter adjustment plan based on the input image quality evaluation data, the device performance status data, and the network status data. The rendering parameter adjustment plan includes a texture accuracy adjustment strategy, a rendering resolution adjustment strategy, and a frame rate adjustment strategy.
[0143] In order to implement an adaptive rendering strategy that dynamically adjusts rendering parameters based on image quality assessment data, device performance status, and network conditions, the specific implementation is as follows:
[0144] First, an image quality assessment dataset is constructed. This dataset contains rendered images of multiple scenes and their corresponding quality assessment data. The image quality assessment data of each scene contains an overall quality score, a detail fidelity score, and a quality distribution map. For example, for a scene image containing trees, houses, and the sky, its overall quality score is 85 points, its detail fidelity score is 90 points, and the quality distribution map shows that the quality score of the tree area is higher, while the quality score of the sky area is slightly lower.
[0145] Next, collect device performance status data. Use monitoring tools to collect the device's CPU utilization, GPU utilization, and memory usage under different load conditions. For example, when the device is rendering a specific scene, the CPU utilization is 60%, the GPU utilization is 80%, and the memory usage is 50%. Then, normalize these data and scale the data range to between 0 and 1. For example, normalize 60% CPU utilization to 0.6, 80% GPU utilization to 0.8, and 50% memory usage to 0.5.
[0146] At the same time, it is necessary to monitor network status data. Use the network test tool to obtain network delay time, network bandwidth data and packet loss rate. For example, the current network delay time is 50 milliseconds, the network bandwidth is 10Mbps, and the packet loss rate is 1%. Similarly, normalize these data to between 0 and 1. For example, normalize the network delay time of 50 milliseconds to 0.5, the network bandwidth of 10Mbps to 0.2, and the packet loss rate of 1% to 0.01.
[0147] Based on the above data, a deep reinforcement learning model is constructed for adaptive rendering parameter adjustment. The state space of the model contains image quality assessment data, normalized device performance status data, and normalized network status data. The action space contains rendering parameter adjustment schemes, such as reducing texture accuracy, reducing rendering resolution, and reducing frame rate.
[0148] During model training, the model selects an action based on the current state and adjusts the rendering parameters based on the action. Then, the image quality is re-evaluated based on the adjusted rendered image, and new device performance status data and network status data are obtained. The new status and the obtained reward are fed back to the model to update the model parameters. The reward function is designed to maximize the image quality score while minimizing device resource consumption and network load.
[0149] Through continuous iterative training, we eventually obtained an adaptive rendering strategy model that can dynamically adjust rendering parameters based on image quality evaluation data, device performance status, and network conditions. For example, when the device performance is good and the network is stable, the model will choose to increase the rendering parameters to obtain higher quality images; when the device performance is poor or the network is not good, the model will choose to reduce the rendering parameters to ensure smooth rendering.
[0150] The solution of this application can:
[0151] Improve user experience: This technical solution can dynamically adjust rendering parameters according to device performance and network conditions, and maximize image quality while ensuring rendering smoothness, thereby improving user experience. Optimize resource utilization: Adaptively adjust rendering parameters to avoid unnecessary resource waste. For example, lowering rendering resolution when network conditions are poor can reduce network bandwidth consumption. Enhance scene adaptability: This technical solution can adapt to different device performance and network conditions, and can provide better rendering effects in various environments.
[0152] Figure 2 Schematic diagram of the structure of a mobile terminal high-quality image rendering system based on AI and cloud rendering according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0153] The first unit is used to obtain a rendering scene request from a mobile terminal; use a deep learning model to analyze the rendering scene request to generate scene feature data, wherein the scene feature data includes geometric information features, material features, and lighting requirement features; construct a rendering task evaluation model based on the scene feature data, wherein the rendering task evaluation model divides the rendering task into a local real-time rendering task sequence and a cloud-based high-quality rendering task sequence; and transmit the local real-time rendering task sequence, the cloud-based high-quality rendering task sequence, and the scene feature data to a task processing module;
[0154] The second unit is used for the task processing module to receive the local real-time rendering task sequence, the cloud high-quality rendering task sequence and the scene feature data; distribute the local real-time rendering task sequence to the mobile terminal processor, and perform frustum culling and scene update based on the geometric information features in the scene feature data; distribute the cloud high-quality rendering task sequence to the cloud rendering cluster, and perform global illumination calculation and material rendering according to the material features and lighting requirement features in the scene feature data; optimize the cloud rendering results by using a generative adversarial network to generate high-quality rendering image data; perform denoising based on deep learning on the high-quality rendering image data to obtain optimized cloud rendering data;
[0155] The third unit is used for receiving the optimized cloud rendering data on the mobile terminal, synthesizing the optimized cloud rendering data with the processing result of the local real-time rendering task to obtain an initial rendered image; establishing an image quality assessment model based on deep learning, the image quality assessment model receives the initial rendered image and generates image quality assessment data; constructing an adaptive rendering parameter adjustment strategy based on the image quality assessment data, the adaptive rendering parameter adjustment strategy dynamically adjusts the rendering parameters according to the device performance status and network conditions, and the rendering parameters include texture accuracy, rendering resolution and frame rate; applying the adjusted rendering parameters to the initial rendered image to generate a final rendered image; and feeding back the image quality assessment data to the scene analysis and preprocessing steps to optimize the generation of the next frame of scene feature data.
[0156] According to a third aspect of the embodiments of the present invention,
[0157] An electronic device is provided, comprising:
[0158] processor;
[0159] a memory for storing processor-executable instructions;
[0160] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0161] A fourth aspect of the embodiments of the present invention is:
[0162] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0163] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A high-quality image rendering method for a mobile terminal based on AI and cloud rendering, characterized in that: include: Obtain a rendering scene request from a mobile terminal; use a deep learning model to analyze the rendering scene request to generate scene feature data, wherein the scene feature data includes geometric information features, material features, and lighting requirement features; construct a rendering task evaluation model based on the scene feature data, wherein the rendering task evaluation model divides the rendering task into a local real-time rendering task sequence and a cloud-based high-quality rendering task sequence; transmit the local real-time rendering task sequence, the cloud-based high-quality rendering task sequence, and the scene feature data to a task processing module; The task processing module receives the local real-time rendering task sequence, the cloud high-quality rendering task sequence and the scene feature data; distributes the local real-time rendering task sequence to the mobile terminal processor, and performs frustum culling and scene update based on the geometric information features in the scene feature data; distributes the cloud high-quality rendering task sequence to the cloud rendering cluster, and performs global illumination calculation and material rendering according to the material features and lighting requirement features in the scene feature data; optimizes the cloud rendering results by using a generative adversarial network to generate high-quality rendering image data; performs deep learning-based denoising processing on the high-quality rendering image data to obtain optimized cloud rendering data; The mobile terminal receives the optimized cloud rendering data, synthesizes the optimized cloud rendering data with the processing result of the local real-time rendering task, and obtains an initial rendered image; establishes an image quality assessment model based on deep learning, and the image quality assessment model receives the initial rendered image and generates image quality assessment data; constructs an adaptive rendering parameter adjustment strategy based on the image quality assessment data, and the adaptive rendering parameter adjustment strategy dynamically adjusts the rendering parameters according to the device performance status and network status, and the rendering parameters include texture accuracy, rendering resolution and frame rate; The adjusted rendering parameters are applied to the initial rendered image to generate a final rendered image; and the image quality assessment data is fed back to the scene analysis and preprocessing step to optimize the generation of the next frame of scene feature data.
2. The method according to claim 1, characterized in that Building a rendering task evaluation model based on the scene feature data, wherein the rendering task evaluation model divides the rendering task into a local real-time rendering task sequence and a cloud high-quality rendering task sequence; transmitting the local real-time rendering task sequence, the cloud high-quality rendering task sequence and the scene feature data to a task processing module comprises: Constructing a scene feature acquisition module, wherein the scene feature acquisition module acquires geometric information feature data, material feature data, and lighting requirement feature data in the scene; constructing a rendering task evaluation model based on the geometric information feature data, the material feature data, and the lighting requirement feature data, wherein the rendering task evaluation model includes a complexity evaluation unit and a device capability evaluation unit; The complexity evaluation unit calculates a GPU load prediction value based on the geometric information feature data, calculates a memory occupancy prediction value based on the material feature data, and calculates a bandwidth demand prediction value based on the lighting demand feature data, and generates computing resource demand evaluation data; The device capability evaluation unit collects GPU computing capability data, available memory capacity data, and network bandwidth data of the local device to generate device performance evaluation data; Inputting the computing resource demand evaluation data and the device performance evaluation data into a task allocation decision unit, the task allocation decision unit generates a task allocation strategy, and the task allocation strategy divides the rendering task into a local real-time rendering task sequence and a cloud high-quality rendering task sequence; The task allocation decision unit allocates a frustum culling task and a scene update task to the local real-time rendering task sequence, and allocates a global illumination calculation task and a complex material rendering task to the cloud high-quality rendering task sequence; A task scheduling management module is constructed, and the task scheduling management module generates a task dependency graph, which includes task execution priorities and synchronization point strategies, and coordinates the execution of the local real-time rendering task sequence and the cloud high-quality rendering task sequence based on the task dependency graph.
3. The method according to claim 1, characterized in that Allocating the local real-time rendering task sequence to the mobile terminal processor, and performing frustum culling and scene update based on the geometric information features in the scene feature data; allocating the cloud high-quality rendering task sequence to the cloud rendering cluster, and performing global illumination calculation and material rendering according to the material features and lighting requirement features in the scene feature data, including: Allocate the local real-time rendering task sequence to the mobile terminal processor, allocate the cloud high-quality rendering task sequence to the cloud rendering cluster, and generate a task allocation mapping table; Based on the task allocation mapping table, a hierarchical bounding box tree structure is constructed on the mobile terminal processor, and the hierarchical bounding box tree structure is divided into nodes by a surface area heuristic algorithm according to the vertex coordinate data in the geometric information feature to generate scene space division data; according to the scene space division data, a frustum culling is performed on the mobile terminal processor, and the intersection relationship between the frustum plane and the node bounding box is calculated by using the separating axis theorem, and the visible object set data is output; Based on the visible object set data, a scene graph structure is constructed in the mobile terminal processor, a transformation matrix of a node in the scene graph structure is updated according to the object transformation information, and the bounding box data of the corresponding node in the hierarchical bounding box tree structure is synchronously updated to generate scene state update data; the scene state update data is transmitted to the cloud rendering cluster, and the cloud rendering cluster extracts the material features according to the task allocation mapping table and constructs a material attribute descriptor; Based on the material attribute descriptor, tasks are assigned to the graphics processing units in the cloud rendering cluster to generate a rendering resource allocation scheme, wherein the rendering resource allocation scheme determines a computing resource allocation weight according to material complexity; According to the rendering resource allocation scheme and the scene status update data, global illumination calculation is performed in the cloud rendering cluster, and direct illumination components, one indirect illumination components and multiple reflection illumination components are calculated through a layered ray tracing module to generate global illumination rendering data; based on the global illumination rendering data and the material attribute descriptor, material rendering calculation is performed in the cloud rendering cluster.
4. The method according to claim 1, characterized in that: The cloud rendering result is optimized by using a generative adversarial network to generate high-quality rendered image data; the high-quality rendered image data is subjected to a denoising process based on deep learning, and the optimized cloud rendering data obtained includes: Constructing a generator of a generative adversarial network, the generator adopts a U-Net architecture, including an encoder module and a decoder module, the encoder module extracts multi-level feature data of the cloud-rendered input image through five downsampling blocks, each of the downsampling blocks includes two three-by-three convolutional layers, a batch normalization layer and a LeakyReLU activation function layer, and the decoder module reconstructs enhanced feature data through five upsampling blocks and a residual connection module; Constructing a discriminator of the generative adversarial network according to the multi-level feature data, wherein the discriminator adopts a PatchGAN structure, and divides the cloud-rendered input image into a plurality of overlapping local image blocks through five convolution modules, each of which includes a convolution layer, an instance normalization layer, and a LeakyReLU activation function layer to generate a discriminant feature map; Calculate the adversarial loss function value based on the discriminant feature map, calculate the content loss function value and the perceptual loss function value based on the enhanced feature data, and use the adversarial loss function value, the content loss function value and the perceptual loss function value as training targets to optimize the network parameters of the generative adversarial network to generate initial optimized image data; Constructing a multi-scale denoising network according to the initial optimized image data, wherein the multi-scale denoising network comprises a feature extraction sub-network, a noise detection sub-network and an image reconstruction sub-network, and the feature extraction sub-network extracts multi-scale feature representation data of the initial optimized image data through a plurality of residual blocks; Building an adaptive noise estimation module based on the multi-scale feature representation data, calculating mean data, variance data and gradient data of the local image block based on the adaptive noise estimation module, predicting the noise intensity value of each local image block through a multi-layer perceptron, and generating a noise intensity distribution map; The noise intensity distribution map is input into the noise detection subnetwork, the noise detection subnetwork generates a noise distribution feature map through an attention mechanism module, and the image reconstruction subnetwork fuses the multi-scale feature representation data and the noise distribution feature map using a dense connection structure; Based on the multi-scale feature representation data and the noise distribution feature map, a mean square error loss function value, a structural similarity loss function value and a perceptual loss function value are calculated, and the mean square error loss function value, the structural similarity loss function value and the perceptual loss function value are used as training targets to optimize the network parameters of the multi-scale denoising network, and generate optimized cloud rendering data.
5. The method according to claim 1, characterized in that The mobile terminal receives the optimized cloud rendering data, and synthesizes the optimized cloud rendering data with the processing result of the local real-time rendering task to obtain an initial rendering image; Establishing an image quality assessment model based on deep learning, wherein the image quality assessment model receives the initial rendered image, and generating image quality assessment data includes: The mobile terminal receives the optimized cloud rendering data, obtains texture feature data, geometric detail data and lighting distribution data in the optimized cloud rendering data; collects the processing result of the local real-time rendering task, and extracts scene depth data, material attribute data and real-time lighting data in the processing result; Constructing a feature extraction network, wherein the feature extraction network includes a cloud feature branch and a local feature branch, each of the feature branches includes four layers of hole convolution layers, respectively extracting multi-scale features from the optimized cloud rendering data and the processing results of the local real-time rendering task, and generating a cloud feature map and a local feature map; A feature fusion network is constructed based on the cloud feature map and the local feature map, wherein the feature fusion network calculates a spatial attention map through a non-local module, calculates a channel attention map through an SE structure, and weightedly fuses the spatial attention map and the channel attention map with the original feature map to generate a fused feature map; Constructing a progressive upsampling network according to the fused feature map, wherein the progressive upsampling network comprises a deconvolution module and a residual enhancement module, wherein the residual enhancement module comprises four residual blocks, and the feature map output by each residual block is processed by a PReLU activation function to generate an initial rendered image; Constructing a dual-stream quality assessment network, wherein the global feature stream of the dual-stream quality assessment network extracts high-level semantic features of the initial rendered image through a ResNet-50 network, and the local feature stream of the dual-stream quality assessment network extracts local texture features of the initial rendered image through a four-layer dense connection block to generate global feature data and local feature data; Calculating a saliency map based on the global feature data and the local feature data, wherein the saliency map indicates a weight distribution of a visual attention area in the initial rendered image, and weighting the global feature data and the local feature data according to the saliency map to generate weighted feature data; Inputting the weighted feature data into a feature enhancement network, wherein the feature enhancement network comprises a self-attention module and a feature pyramid module, wherein the self-attention module calculates the correlation between feature points, and the feature pyramid module fuses feature representations of different scales to generate enhanced feature data; A quality assessment module is constructed based on the enhanced feature data, and the quality assessment module calculates the brightness mean, contrast variance and structural similarity of the image block, and outputs image quality assessment data through a fully connected layer.
6. The method according to claim 1, characterized in that An adaptive rendering parameter adjustment strategy is constructed based on the image quality evaluation data, and the adaptive rendering parameter adjustment strategy dynamically adjusts rendering parameters according to device performance status and network conditions, and the rendering parameters include texture accuracy, rendering resolution and frame rate. The image quality assessment data includes an overall quality score, a detail fidelity score, and a quality distribution map, and an adaptive rendering parameter adjustment strategy model based on deep reinforcement learning is constructed, wherein the state space of the adaptive rendering parameter adjustment strategy model includes device performance status data and network status data, and the action space includes an adjustment plan for rendering parameters; Collecting the device performance status data, the device performance status data including central processing unit utilization, graphics processing unit utilization, and memory occupancy, normalizing the device performance status data and inputting the data into the adaptive rendering parameter adjustment strategy model; Monitoring the network status data, wherein the network status data includes network delay time, network bandwidth data, and data packet loss rate, and normalizing the network status data and inputting it into the adaptive rendering parameter adjustment strategy model; The adaptive rendering parameter adjustment strategy model generates a rendering parameter adjustment plan based on the input image quality evaluation data, the device performance status data, and the network status data. The rendering parameter adjustment plan includes a texture accuracy adjustment strategy, a rendering resolution adjustment strategy, and a frame rate adjustment strategy.
7. A mobile high-quality image rendering system based on AI and cloud rendering, used to implement the method described in any one of claims 1 to 6, characterized in that: include: The first unit is used to obtain a rendering scene request from a mobile terminal; use a deep learning model to analyze the rendering scene request to generate scene feature data, wherein the scene feature data includes geometric information features, material features, and lighting requirement features; construct a rendering task evaluation model based on the scene feature data, wherein the rendering task evaluation model divides the rendering task into a local real-time rendering task sequence and a cloud-based high-quality rendering task sequence; and transmit the local real-time rendering task sequence, the cloud-based high-quality rendering task sequence, and the scene feature data to a task processing module; The second unit is used for the task processing module to receive the local real-time rendering task sequence, the cloud high-quality rendering task sequence and the scene feature data; distribute the local real-time rendering task sequence to the mobile terminal processor, and perform frustum culling and scene update based on the geometric information features in the scene feature data; distribute the cloud high-quality rendering task sequence to the cloud rendering cluster, and perform global illumination calculation and material rendering according to the material features and lighting requirement features in the scene feature data; optimize the cloud rendering results by using a generative adversarial network to generate high-quality rendering image data; perform denoising based on deep learning on the high-quality rendering image data to obtain optimized cloud rendering data; The third unit is used for receiving the optimized cloud rendering data on the mobile terminal, synthesizing the optimized cloud rendering data with the processing result of the local real-time rendering task to obtain an initial rendered image; establishing an image quality assessment model based on deep learning, the image quality assessment model receives the initial rendered image and generates image quality assessment data; constructing an adaptive rendering parameter adjustment strategy based on the image quality assessment data, the adaptive rendering parameter adjustment strategy dynamically adjusts the rendering parameters according to the device performance status and network status, and the rendering parameters include texture accuracy, rendering resolution and frame rate; The adjusted rendering parameters are applied to the initial rendered image to generate a final rendered image; and the image quality assessment data is fed back to the scene analysis and preprocessing step to optimize the generation of the next frame of scene feature data.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Distributed real-time rendering method and system based on edge calculation and medium
CN120355829A
WEB data processing method based on collaborative rendering
CN120707723A
VUE dynamic component interface configuration and rendering method for multi-scene multiplexing
CN120910371A
Three-dimensional space data rendering method, device and system, business server and computer storage medium
CN121233343A
Task scheduling method and system based on hybrid cloud digital human processing architecture
CN122179477A