Generative AI model real-time rendering engine construction method and related equipment thereof
By constructing a real-time rendering engine for generative AI models, utilizing the CLIP model to parse multimodal semantic information and generate lightweight intermediate representations, and combining rasterization acceleration and ray stepping technology, the problems of long rendering time and poor consistency in generating 3D content by generative AI models are solved, achieving efficient real-time rendering and interaction, and is suitable for multiple platforms.
Patent Information
- Application Number
- CN202511478233.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-16
AI Technical Summary
In existing technologies, generative AI models generate 3D content that needs to be exported to an independent rendering engine for processing, which results in excessive time consumption. Furthermore, the generated content exhibits randomness, inter-frame flickering, and poor spatial consistency, affecting the visual experience.
We build a real-time rendering engine for generative AI models. By parsing multimodal semantic information through CLIP models, we generate lightweight intermediate representations. Combined with rasterization acceleration and ray stepping technology, we capture user operations in real time and optimize performance by using hybrid rendering strategies and spatiotemporal constraint technology.
It achieves deep integration of generation and rendering, compressing end-to-end latency from seconds to milliseconds, improving generation efficiency and visual consistency, supporting real-time interactive operation, adapting to different hardware platforms, and enhancing the visual experience.
Smart Images

Figure CN120953465A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, and more specifically, to a method for constructing a real-time rendering engine for generative AI models and related equipment. Background Technology
[0002] Generative AI models refer to models that automatically generate content through artificial intelligence algorithms, rather than simply editing existing content. In a rendering scene, it can proactively generate entirely new visual content (such as 3D models, textures, lighting effects, etc.) based on text descriptions, sketches, a few images, or parameters. For example, inputting "a cyberpunk-style futuristic city" can directly generate 3D scene details that match that description. The rendering engine is the core system that transforms digital models (3D geometry, materials, lighting, etc.) into visualized images. "Real-time" means that the rendering speed must match human visual perception (usually ≥30fps, i.e., generating more than 30 frames per second) and support interactive operations (such as view rotation and scene editing) without significant latency. Traditional real-time rendering relies on predefined geometry and materials, but when combined with generative AI, the engine can dynamically generate content, breaking through the limitations of fixed materials.
[0003] Shortcomings of existing technology: In traditional technologies, generative AI models generate 3D content that needs to be exported to an independent rendering engine for processing. This involves a serial process of "generating first and then rendering," which results in excessively long overall processing times. The generation of 3D content by the model is random (such as color deviation and shape distortion) and is difficult to respond to real-time parameter adjustments. Frame flickering (poor temporal consistency) or abrupt changes on the surface of objects (poor spatial consistency) are prone to occur within the AI generation, affecting the visual experience.
[0004] To address the above problems, this invention proposes a solution. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a method for constructing a real-time rendering engine for generative AI models and related equipment. By constructing a system, the AI model actively generates 3D content and transforms this content into visualized images, thereby solving the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The method for building a real-time rendering engine for generative AI models includes the following steps: The system acquires multimodal semantic information from user input, parses the multimodal semantic information using the CLIP model, and transforms it into latent vectors and constraint parameters that can be understood by the generative AI model. A lightweight generation model is constructed based on latent vectors and constraint parameters. The lightweight generation model is used to directly generate renderable intermediate representations, which include 3D Gaussian point clouds, implicit signed distance functions, and dynamic texture maps. A rendering pipeline is built based on the generated intermediate representation, and the intermediate representation is transformed into a visual image by combining rasterization acceleration and ray stepping and spatial hashing acceleration techniques. The system captures user actions in real time and dynamically updates the constraints of the generated model, triggering incremental generation. At the same time, it optimizes system performance by employing model lightweighting, hardware acceleration, and hybrid rendering strategies. By using temporal filtering and spatial constraint techniques, the generated content is made consistent in both time and space.
[0007] In a preferred embodiment, the specific steps for parsing multimodal semantic information using the CLIP model and converting it into latent vectors and constraint parameters that can be understood by the generative AI model are as follows: The preprocessed text sequence is input into the text encoder of the CLIP model to obtain the text feature vector. The CLIP text encoder adopts the Transformer structure and extracts the semantic features of the text through deep encoding of the text sequence. The preprocessed image is input into the image encoder of the CLIP model to obtain the image feature vector. The CLIP image encoder typically uses a convolutional neural network to extract the visual features of the image. The text feature vector and the image feature vector are fused to obtain the fused multimodal feature vector, as shown in the following formula:
[0008] In the formula, For multimodal feature vectors, This is the aligned text feature vector. The aligned image feature vector, denoted as b, where b is the importance weight of text features and b is the importance weight of image features. The fused multimodal feature vectors are mapped onto the latent space of the generative AI model to obtain the latent vectors; From the multimodal features extracted from the CLIP model and the parameters input by the user, the constraint parameters required for the generative AI model are extracted.
[0009] In a preferred embodiment, the fused multimodal feature vectors are mapped to the latent space of the generative AI model to obtain latent vectors, as follows: Obtain the latent space dimension of the generative AI model, and adjust the dimension of the fused feature vector based on the latent space dimension; The adjusted fused feature vectors are projected onto the latent space of the generative model through a nonlinear mapping network to obtain latent vectors, ensuring that the vector distribution is consistent with the latent space during model training.
[0010] In a preferred embodiment, the process of projecting the adjusted fused features onto the latent space of the generative model using a nonlinear mapping network to obtain the latent vector is as follows: Batch normalization is performed on the fused multimodal feature vectors to stabilize their distribution range; The batch-normalized multimodal feature vectors are input into a multilayer perceptron (MLP) for nonlinear transformation; the output of the multilayer perceptron (MLP) is regularized to ensure that it conforms to the prior distribution of the latent space, thereby obtaining the latent vector.
[0011] In a preferred embodiment, the process of directly generating a renderable intermediate representation using the lightweight generation model is as follows: Using latent vectors as input, the generator outputs an initial point set and Gaussian parameters, imposes constraints on the Gaussian parameters, and generates a 3D Gaussian point cloud by clustering and merging similar Gaussian points. A micro MLP is used to generate an implicit signed distance function. The input is 3D coordinates, latent vector and constraint parameters, and the output is the signed distance from the coordinate to the object surface. The shape constraint is converted into a spatial mask and positive distance is forced to be output for coordinates that exceed the boundary. Based on latent vectors, a lightweight GAN generator generates basic texture feature maps and converts material parameters into texture modulation factors. The texture resolution is dynamically adjusted according to rendering requirements to obtain dynamic texture maps.
[0012] In a preferred embodiment, the process of capturing user actions in real time and dynamically updating the constraints of the generated model is as follows: User actions are captured in real time via the system API, with a sampling frequency of 60Hz to ensure zero latency. User actions are transformed into incremental changes in constraint parameters, and synchronized to the generative model through an event callback mechanism to achieve dynamic updates of the constraints in the generative model.
[0013] In a preferred embodiment, the incremental generation process is triggered as follows: Generation is triggered only for the region affected by the operation; if the operation involves the entire region, incremental updates of all intermediate representations are generated. If it is a local operation, the generation range is limited by a spatial mask.
[0014] In a preferred embodiment, the process of optimizing the performance of the real-time rendering engine for generative AI models using model lightweighting, hardware acceleration, and hybrid rendering strategies is as follows: The model precision is automatically switched according to hardware performance. FP16 precision is used on high-end GPUs, and INT8 precision is automatically switched on mobile GPUs. Enable dedicated GPU acceleration units for Gaussian point cloud matrix operations, enable dedicated rendering channels for mobile GPUs, and distribute rasterization and ray stepping tasks to different computing units for parallel processing through synchronization barriers. Static areas use a pre-rendered cache, storing the result as a texture after the first rendering, and then sampling it directly in subsequent frames. Dynamic areas use real-time generation plus low-resolution rendering, and recover details through a super-resolution module, balancing speed and quality.
[0015] In a preferred embodiment, the process of ensuring that the generated visualization image is consistent in time and space through temporal filtering and spatial constraint techniques is as follows: Extract CLIP features from the current frame and the previous frame and calculate cosine similarity. If the cosine similarity is less than the threshold, backtrack and adjust the latent vector of the generation model to ensure that the semantics of the generated content are consistent with the historical frames. The image quality score is calculated using a no-reference image quality assessment algorithm to detect whether there is blurring in the image due to excessive spatiotemporal constraints. If the quality score is lower than the threshold, the smoothing coefficient is dynamically reduced or the spatial constraint threshold is increased to balance consistency and sharpness.
[0016] An electronic device, the electronic device comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer programs that can be executed by the at least one processor. The computer program is executed by the at least one processor. So that the at least one processor can execute the method for building a real-time rendering engine for generative AI models as described in any one of claims 1 to 9.
[0017] The technical effects and advantages of the generative AI model real-time rendering engine construction method and related equipment of this invention are as follows: 1. This invention constructs a deeply integrated "generation-rendering" architecture, enabling lightweight generation models to directly output renderable intermediate representations. Furthermore, the rendering pipeline and generation module share data interfaces, effectively solving the problem of separation between generation and rendering. This reduces end-to-end latency from seconds to milliseconds, meeting real-time requirements of over 30fps and significantly improving overall efficiency. Through constraint injection and incremental generation mechanisms, hard constraints are embedded into the generation process using feature modulation technology. Incremental generation is triggered based on user actions, reducing the error rate between generated content and constraint parameters and improving generation efficiency, thus overcoming the technical challenge of poor controllability in AI generation.
[0018] 2. This invention employs a multi-layered lightweight strategy, performing knowledge distillation, quantization, and pruning at the model level, and automatically adapting to different GPU computing power at the hardware level while combining a hybrid rendering strategy. This overcomes the limitations of insufficient hardware adaptability, achieving cross-platform compatibility from mobile devices to high-end workstations. Through spatiotemporal constraint technology, exponential weighted smoothing and Gamma correction are used in the time dimension, while 3D Gaussian point clouds and SDF are processed in the spatial dimension, overcoming the quality defects caused by the lack of spatiotemporal consistency and improving the visual experience. The related devices of this invention, through hardware-software co-optimization, further improve rendering efficiency and operational response speed, enhance the real-time interactive experience, and provide efficient 3D content generation and rendering solutions for multiple fields. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the structure of the generative AI model real-time rendering engine construction method of the present invention.
[0020] Figure 2 This is a schematic diagram of the device structure for building the real-time rendering engine of the generative AI model of this invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0022] Example 1, Figure 1 The present invention provides a method for constructing a real-time rendering engine for generative AI models.
[0023] The system acquires multimodal semantic information from user input, parses the multimodal semantic information using the CLIP model, and transforms it into latent vectors and constraint parameters that can be understood by the generative AI model. The multimodal semantic information (text, images, parameters, etc.) input by users is a natural way for humans to interact, but generative AI models cannot directly understand this unstructured or heterogeneous information. For example, the text "red curved object" is a natural language description, the image is a pixel matrix, and the parameters are isolated values. AI models need structured feature vectors to perform generative calculations, so they must establish connections through parsing and transformation.
[0024] The process of obtaining multimodal semantic information from user input is as follows: receiving natural language descriptions from user input through a text input interface, obtaining reference images provided by the user through an image acquisition device (such as a camera) or an image upload interface, and receiving specific parameters set by the user through the interactive interface; The input text is segmented, stop words are removed, and lemmatization is performed to obtain a standardized text sequence. The input image is resized, normalized, and denoised to unify the image format and quality, such as resizing the image to 224×224 pixels and normalizing the pixel values to the range of [0,1]. The user-input parameters are converted into standardized numerical forms to facilitate subsequent processing, such as converting the hexadecimal values of colors to RGB values, thereby obtaining multimodal semantic information.
[0025] The specific steps for parsing multimodal semantic information using the CLIP model and transforming it into latent vectors and constraint parameters that can be understood by generative AI models are as follows: The preprocessed text sequence is input into the text encoder of the CLIP model to obtain the text feature vector. The CLIP text encoder adopts a Transformer structure and extracts the semantic features of the text through deep encoding of the text sequence. The preprocessed image is input into the image encoder of the CLIP model to obtain the image feature vector. The image encoder of CLIP usually uses a convolutional neural network (such as ResNet, VisionTransformer, etc.) to extract the visual features of the image. The text feature vector and the image feature vector are fused to obtain the fused multimodal feature vector. The process is as follows: Since the feature vectors output by the text encoder and image encoder of the CLIP model may have different dimensions, a linear transformation is first needed to unify the dimensions, aligning the dimensions of the text feature vectors and image feature vectors. Weights are then assigned based on the importance of the multimodal inputs, and the aligned feature vectors are summed using a weighted average to obtain the multimodal feature vector, as shown in the following formula:
[0026] In the formula, For multimodal feature vectors, This is the aligned text feature vector. The aligned image feature vector, denoted as b, where b is the importance weight of text features and b is the importance weight of image features.
[0027] The fused multimodal feature vectors are mapped to the latent space of the generative AI model to obtain latent vectors, as follows: Obtain the latent space dimension of the generative AI model (such as diffusion model, StyleGAN), and adjust the dimension of the fused feature vector according to the latent space dimension; project the adjusted fused feature vector onto the latent space of the generative model through a nonlinear mapping network to obtain the latent vector, ensuring that the vector distribution is consistent with the latent space during model training.
[0028] The process of projecting the adjusted fused features onto the latent space of the generative model using a nonlinear mapping network to obtain the latent vector is as follows: First, the fused multimodal feature vectors are batch-normalized to stabilize their distribution range; then, the batch-normalized multimodal feature vectors are input into a multilayer perceptron (MLP) for nonlinear transformation; finally, the output of the MLP is regularized to ensure it conforms to the prior distribution of the latent space, thus obtaining the latent vector. This process is achieved through a linear mapping layer, which transforms the fused feature vectors into vectors that match the dimensions of the latent space of the generative AI model.
[0029] The constraint parameter extraction process is as follows: Constraint parameters required by the generative AI model are extracted from the multimodal features extracted from the CLIP model and the parameters input by the user. For example, the color constraint parameter "blue" (RGB value) and the shape constraint parameter "circle" are extracted from the text "blue round table"; the height and diameter of the table, and the material constraint parameters are extracted from the size parameters input by the user.
[0030] A lightweight generation model is constructed based on latent vectors and constraint parameters. The lightweight generation model is used to directly generate renderable intermediate representations, which include 3D Gaussian point clouds, implicit signed distance functions, and dynamic texture maps. The process of constructing a lightweight generative model based on latent vectors and constraint parameters is as follows: Constraint parameters are embedded into the intermediate layer of the generator through feature modulation to ensure that the generated content meets hard constraints; a large model (such as the original StableDiffusion3D) is used as the teacher model, and the lightweight model learns the generation distribution through distillation loss; the original floating-point values are mapped to the integer domain through a scaling factor, and the model weights are quantized using INT8, while the activation values are quantized using FP16; redundant channels are filtered using L1 regularization, retaining the feature channels that contribute highly to the generation results, and then... Layer channel weights ,like ( If the threshold is set, then the channel is pruned. The threshold is controlled by the accuracy loss of the validation set.
[0031] The process of directly generating a renderable intermediate representation using the aforementioned lightweight generation model is as follows: Taking the latent vector as input, the generator outputs the initial point set and Gaussian parameters, and imposes constraints on the Gaussian parameters; by clustering and merging similar Gaussian points (merging points with an intersection-union ratio ≥ 0.8), the number of points is reduced to 60% of the original, while maintaining the visual effect. The Gaussian parameters include color parameters and shape parameters. An 8-layer micro MLP (64-128 channels per layer) is used. The input is 3D coordinates, latent vectors and constraint parameters, and the output is the signed distance from the coordinates to the object surface. The shape constraint is converted into a spatial mask, and positive distance is forced to be output for coordinates that exceed the boundary. Based on latent vectors, a lightweight GAN generator is used to generate basic texture feature maps; material parameters are converted into texture modulation factors, and texture resolution is dynamically adjusted according to rendering requirements to obtain dynamic texture maps.
[0032] A rendering pipeline is built based on the generated intermediate representation, and the intermediate representation is transformed into a visual image by combining rasterization acceleration and ray stepping and spatial hashing acceleration techniques. The rendering pipeline adopts a three-stage pipeline of "input routing + parallel processing + result fusion". It uses a feature detector to determine the intermediate representation type and automatically switches the processing path according to the intermediate representation type (3D Gaussian point cloud / implicit SDF / dynamic texture). If the intermediate representation is a 3D Gaussian point cloud, the mean and covariance of the 3D Gaussian points are converted into screen space parameters, projected onto 2D screen coordinates using the camera projection matrix in 3D coordinates, and the scaling of the covariance in screen space is calculated. All Gaussian points are processed in parallel using the GPU's CUDA cores. For each Gaussian point, its coverage area on the screen (such as a 2D elliptical region) is calculated based on the scaling of the mean and covariance in screen space. The pixel color contribution is quickly generated using a pre-computed Gaussian weight table. Overlapping pixels are sorted according to their projected z-coordinates, and alpha blending is used to eliminate occlusion conflicts. If the intermediate representation type is implicit SDF, the 3D space is divided into a voxel mesh. A hash table is constructed using a hash function to store the SDF sampling points within each voxel, accelerating the coarse localization of the intersection points between the ray and the object. A ray is emitted for each pixel on the screen, and the voxels that the ray may pass through are located using the spatial hash table. Adaptive stepping is used within the voxel to quickly approximate the object surface. At the intersection point, the normal vector is calculated using the center difference, and the pixel color is calculated by combining the dynamic texture map. If the intermediate representation type is a dynamic texture map, the dynamic texture T(x,y,t) is bound to the GPU texture unit and mapped to the 3D model surface (such as a Gaussian point cloud or SDF surface) through UV coordinates. Only the changed areas of the dynamic texture are updated, reducing the amount of GPU data transfer.
[0033] The system captures user actions in real time and dynamically updates the constraints of the generated model, triggering incremental generation. At the same time, it optimizes the performance of the real-time rendering engine for generative AI models by employing model lightweighting, hardware acceleration, and hybrid rendering strategies. The process of capturing user operations in real time and dynamically updating the constraints of the generated model is as follows: user operations such as mouse / touch, keyboard, materials, and VR controllers are captured in real time through the system API, with a sampling frequency of 60Hz to ensure zero latency; user operations are converted into incremental changes in constraint parameters and synchronized to the generated model through an event callback mechanism; The incremental generation process is triggered as follows: generation is triggered only for the area affected by the operation. If the operation involves the whole world (such as view rotation), incremental updates of all intermediate representations are generated. If it is a local operation (such as modifying the local color of an object), the generation range is limited by a spatial mask (mask value 1 indicates the area to be updated).
[0034] The optimization process for the real-time rendering engine of generative AI models using model lightweighting, hardware acceleration, and hybrid rendering strategies is as follows: The model precision is automatically switched based on hardware performance (such as GPU computing power and memory). FP16 precision is used on high-end GPUs (such as RTX4090), while INT8 precision is automatically switched on mobile devices (such as Snapdragon 8 Gen3). The quantization parameter scale factor is adjusted in real time. A dedicated GPU acceleration unit is enabled. TensorCore is activated on NVIDIA GPUs to perform Gaussian point cloud matrix operations. Vulkan's dedicated rendering pass (RenderPass) is enabled on mobile GPUs. Rasterization and ray stepping tasks are distributed to different computing units for parallel processing through a synchronization fence. Static areas (such as backgrounds and fixed objects) use pre-rendering cache. After the first rendering, the result is stored as a texture and sampled directly in subsequent frames. Dynamic areas (such as moving objects) use real-time generation + low-resolution rendering. Details are restored through a super-resolution module to balance speed and quality.
[0035] By employing temporal filtering and spatial constraint techniques, the generated visualization images are made consistent in both time and space.
[0036] For the current frame And the previous frame The pixel motion vector is calculated using an optical flow algorithm, representing the pixel's movement from the current frame. Go to the previous frame The displacement is calculated; the image is divided into static and dynamic regions based on motion vector clustering; a smoothing coefficient is set for the static and dynamic regions, and the color of each pixel in the current frame is filtered based on the motion weight coefficient and the smoothed result of the previous frame; at the same time, the inter-frame brightness difference is calculated for regions with sudden brightness changes (such as changes in illumination), and if the inter-frame brightness difference is greater than a preset threshold, Gamma correction is triggered.
[0037] Spatial gradient checks are performed on the generated intermediate representations. The distance between adjacent points is calculated using 3D Gaussian point calculations; if the distance exceeds a threshold, transition points are inserted. The SDF gradient of adjacent sampling points in the SDF space is calculated; if there are abrupt changes, the SDF function is optimized using a smoothing term to make the object surface smoother. Dynamic texture maps are also processed. Calculate the local texture gradient. If the texture difference between adjacent pixels in a certain region exceeds a threshold, then process it by bilateral filtering smoothing.
[0038] Extract CLIP features from the current frame and the previous frame and calculate cosine similarity. If the cosine similarity is less than the threshold, backtrack and adjust the latent vector of the generation model to ensure that the semantics of the generated content are consistent with the historical frames. Use a no-reference image quality assessment algorithm (such as BRISQUE) to detect whether there is blurring in the image due to excessive spatiotemporal constraints. If the quality score is lower than the threshold, dynamically reduce the smoothing coefficient or increase the spatial constraint threshold to balance consistency and sharpness.
[0039] It should be noted that by using spatiotemporal constraint technology, exponential weighted smoothing and Gamma correction are applied in the time dimension, and 3D Gaussian point clouds and SDF are processed in the spatial dimension, overcoming the quality defects caused by the lack of spatiotemporal consistency and improving the visual experience.
[0040] Example 2, Figure 2 The present invention provides a device for building a real-time rendering engine for generative AI models.
[0041] An electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute the generative AI model real-time rendering engine construction method of the present invention.
[0042] Through hardware-software co-optimization, rendering efficiency and operation response speed have been further improved, real-time interactive experience has been enhanced, and efficient 3D content generation and rendering solutions have been provided for multiple fields.
[0043] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0044] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0045] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0046] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0047] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0048] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for constructing a real-time rendering engine for generative AI models, characterized in that, Includes the following steps: The system acquires multimodal semantic information from user input, parses the multimodal semantic information using the CLIP model, and transforms it into latent vectors and constraint parameters that can be understood by the generative AI model. A lightweight generation model is constructed based on latent vectors and constraint parameters. The lightweight generation model is used to directly generate renderable intermediate representations, which include 3D Gaussian point clouds, implicit signed distance functions, and dynamic texture maps. A rendering pipeline is built based on the generated intermediate representation, and the intermediate representation is transformed into a visual image by combining rasterization acceleration and ray stepping and spatial hashing acceleration techniques. The system captures user actions in real time and dynamically updates the constraints of the generated model, triggering incremental generation. At the same time, it optimizes system performance by employing model lightweighting, hardware acceleration, and hybrid rendering strategies. By using temporal filtering and spatial constraint techniques, the generated content is made consistent in both time and space.
2. The method for constructing a real-time rendering engine for generative AI models according to claim 1, characterized in that, The specific steps for parsing multimodal semantic information using the CLIP model and transforming it into latent vectors and constraint parameters that can be understood by generative AI models are as follows: The preprocessed text sequence is input into the text encoder of the CLIP model to obtain the text feature vector. The CLIP text encoder adopts the Transformer structure and extracts the semantic features of the text through deep encoding of the text sequence. The preprocessed image is input into the image encoder of the CLIP model to obtain the image feature vector. The CLIP image encoder typically uses a convolutional neural network to extract the visual features of the image. The text feature vector and the image feature vector are fused to obtain the fused multimodal feature vector, as shown in the following formula: ; In the formula, For multimodal feature vectors, This is the aligned text feature vector. The aligned image feature vector, denoted as b, where b is the importance weight of text features and b is the importance weight of image features. The fused multimodal feature vectors are mapped onto the latent space of the generative AI model to obtain the latent vectors; From the multimodal features extracted from the CLIP model and the parameters input by the user, the constraint parameters required for the generative AI model are extracted.
3. The method for constructing a real-time rendering engine for generative AI models according to claim 2, characterized in that, The fused multimodal feature vectors are mapped to the latent space of the generative AI model to obtain latent vectors, as follows: Obtain the latent space dimension of the generative AI model, and adjust the dimension of the fused feature vector based on the latent space dimension; The adjusted fused feature vectors are projected onto the latent space of the generative model through a nonlinear mapping network to obtain latent vectors, ensuring that the vector distribution is consistent with the latent space during model training.
4. The method for constructing a real-time rendering engine for generative AI models according to claim 3, characterized in that, The process of projecting the adjusted fused features onto the latent space of the generative model using a nonlinear mapping network to obtain the latent vector is as follows: Batch normalization is performed on the fused multimodal feature vectors to stabilize their distribution range; The batch-normalized multimodal feature vectors are input into a multilayer perceptron (MLP) for nonlinear transformation; the output of the multilayer perceptron (MLP) is regularized to ensure that it conforms to the prior distribution of the latent space, thereby obtaining the latent vector.
5. The method for constructing a real-time rendering engine for generative AI models according to claim 4, characterized in that, The process of directly generating a renderable intermediate representation using the aforementioned lightweight generation model is as follows: Using latent vectors as input, the generator outputs an initial point set and Gaussian parameters, and imposes constraints on the Gaussian parameters. By clustering and merging similar Gaussian points, a 3D Gaussian point cloud is generated. A micro MLP is used to generate an implicit signed distance function. The input is 3D coordinates, latent vector and constraint parameters, and the output is the signed distance from the coordinate to the object surface. The shape constraint is converted into a spatial mask and positive distance is forced for coordinates that exceed the boundary. Based on latent vectors, a lightweight GAN generator generates basic texture feature maps and converts material parameters into texture modulation factors. The texture resolution is dynamically adjusted according to rendering requirements to obtain dynamic texture maps.
6. The method for constructing a real-time rendering engine for generative AI models according to claim 5, characterized in that, The process of capturing user actions in real time and dynamically updating the constraints of the generated model is as follows: User actions are captured in real time via the system API, with a sampling frequency of 60Hz to ensure zero latency. User actions are transformed into incremental changes in constraint parameters, and synchronized to the generative model through an event callback mechanism to achieve dynamic updates of the constraints in the generative model.
7. The method for constructing a real-time rendering engine for generative AI models according to claim 6, characterized in that, The incremental generation process is triggered as follows: Generation is triggered only for the region affected by the operation; if the operation involves the entire region, incremental updates of all intermediate representations are generated. If it is a local operation, the generation range is limited by a spatial mask.
8. The method for constructing a real-time rendering engine for generative AI models according to claim 7, characterized in that, The optimization process for the real-time rendering engine of generative AI models using model lightweighting, hardware acceleration, and hybrid rendering strategies is as follows: The model precision is automatically switched according to hardware performance. FP16 precision is used on high-end GPUs, and INT8 precision is automatically switched on mobile GPUs. Enable dedicated GPU acceleration units for Gaussian point cloud matrix operations, enable dedicated rendering channels for mobile GPUs, and distribute rasterization and ray stepping tasks to different computing units for parallel processing through synchronization barriers. Static areas use a pre-rendered cache, storing the result as a texture after the first rendering, and then sampling it directly in subsequent frames. Dynamic areas use real-time generation plus low-resolution rendering, and recover details through a super-resolution module, balancing speed and quality.
9. The method for constructing a real-time rendering engine for generative AI models according to claim 8, characterized in that, The process of ensuring temporal and spatial consistency of the generated visualization image through temporal filtering and spatial constraint techniques is as follows: Extract CLIP features from the current frame and the previous frame and calculate cosine similarity. If the cosine similarity is less than the threshold, backtrack and adjust the latent vector of the generation model to ensure that the semantics of the generated content are consistent with the historical frames. The image quality score is calculated using a no-reference image quality assessment algorithm to detect whether there is blurring in the image due to excessive spatiotemporal constraints. If the quality score is lower than the threshold, the smoothing coefficient is dynamically reduced or the spatial constraint threshold is increased to balance consistency and sharpness.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer programs that can be executed by the at least one processor. The computer program is executed by the at least one processor. So that the at least one processor can execute the method for building a real-time rendering engine for generative AI models as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Image rendering method, image rendering model generation method and related device
CN115731336A
AI visual special effect dynamic generation system fused with multi-modal perception
CN120318379A
Adaptive rendering method and system for characters and images in AI digital human virtual and real scene fusion
CN120355827A
Scene generation and interaction method and device, electronic equipment, medium and program product
CN120726238A
3D GAN inversion method and apparatus with pose optimization
KR102625474B1
Cited By
Three-dimensional earth data fusion method and system based on multi-source space-time images
CN121350169A
Data processing method and device
CN121597447A