Text control scene rendering graph generation method and system based on foreground model
Through the foreground model-based text-controlled scene rendering generation method, text-driven controllable generation network and multimodal encoder are used to solve the problems of cumbersome design and high demand for computing resources in traditional 3D rendering technology, and efficient and flexible background rendering generation is achieved.
Patent Information
- Application Number
- CN202510226271.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
When traditional 3D rendering technology generates background renderings of complex products or scenes, there are problems such as cumbersome design processes, high demand for computing resources, and low design efficiency.
The text-controlled scene rendering image generation method based on the foreground model is adopted, and the product background rendering image that meets the requirements is automatically generated through text driver controllable generation network, combining image text multimodal encoder and shadow adaptive conditions alignment.
Improve design efficiency, reduce the demand for computing resources, enhance the flexibility and customization of image generation, and reduce design costs and time.
Smart Images

Figure CN120182483A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer graphics, and particularly to a method and system for generating a rendering diagram of a text-controlled scene based on a foreground model. Background Art
[0002] In the traditional process of generating product design images, 3D rendering is a commonly used method, but this method has significant limitations. This technology requires designers to manually construct the entire scene, and the process is extremely cumbersome. From the perspective of modeling, designers need to precisely create the three-dimensional geometric structure of each object. When dealing with complex products or scenes, it involves a large amount of point, line, and surface data processing, and designers need to possess high-level professional skills in 3D modeling, lighting control, and material settings. In terms of lighting and material settings, traditional 3D rendering also faces huge challenges. Different light source types, such as point lights, parallel lights, and ambient lights, have significantly different reflection, refraction, and shadow effects in different scenes. In interior design rendering, it is very complex to adjust the light color, intensity, and angle to create a warm or bright atmosphere. In terms of material settings, from the glossiness of metals to the texture details of wood, designers need to debug them one by one, and the interaction effects between different materials, such as the light and shadow changes when metal is adjacent to plastic, also need to be finely processed. In addition, traditional 3D rendering technology has extremely high requirements for computing resources. When rendering images with high resolution and complex scenes, it requires a powerful graphics processing unit and a large amount of memory support. In some large industrial design projects, high investment is required to purchase professional hardware equipment to meet the rendering requirements, which undoubtedly increases the design cost. Moreover, even with powerful hardware, the rendering process may still take a long time, seriously affecting the design efficiency and occupying too much of the designers' time. Therefore, how to automatically and efficiently generate a product background rendering diagram that meets the requirements only by inputting text descriptions or setting conditions, while ensuring the flexibility of the background, is a technical problem that needs to be solved. Summary of the Invention
[0003] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method and system for generating a rendering diagram of a text-controlled scene based on a foreground model, which uses a text-driven controllable generation network, guides the model to generate a suitable background through scene descriptions, and adaptively adjusts the scene image through an image-text multi-modal encoder and shadow adaptive conditional alignment.
[0004] The purpose of the present invention can be achieved by the following technical solutions:
[0005] According to one aspect of the present invention, a method for generating a rendering diagram of a text-controlled scene based on a foreground model is provided, and the specific steps include:
[0006] S1. Collect the design information of the product image to be generated. The foreground modeling module performs geometric reconstruction based on the geometric file, material texture information, and parameter constraints of the product to obtain the corresponding 3D model;
[0007] S2. Input the 3D model into the foreground rendering module for rendering processing to obtain the foreground target rendering image;
[0008] S3. The text-driven effect image generation module generates the rendering effect image of the product scene through the image-text multimodal encoder and the noise prediction network according to the scene description and the foreground target rendering image.
[0009] Furthermore, the specific steps for the text-driven effect image generation module to generate the rendering effect image of the product scene are as follows:
[0010] The image-text multimodal encoder extracts features from the text label and the query vector, and associates the text feature information and the query vector information in a shared parameter manner. The query vector performs cross-attention calculation with the foreground target rendering image, and finally performs a linear transformation of the visual-text multimodal features through the feature mapping composed of the feed-forward neural network to obtain the foreground image multimodal embedding;
[0011] The noise prediction network compresses the initial image into the latent space and gradually adds noise to obtain the noise image; receives the scene description and the foreground image multimodal embedding from the image-text multimodal encoder, associates the foreground image multimodal embedding and the noise image through the visual cross-attention mechanism, associates the scene text embedding and the noise image through the text cross-attention mechanism, and performs noise prediction; denoises the noise image according to the predicted noise to obtain the latent space generated image, and further decodes it to obtain the generated image.
[0012] Furthermore, during the training process of the text-driven effect image generation module, after obtaining the generated image, the shadow adaptive conditional alignment is used as the training target. The foreground image multimodal embedding and the background color of the initial image are made consistent through the transformation T, and then consistency constraints are performed among the initial image, the generated image, and the foreground image multimodal embedding. The expression is:
[0013]
[0014] where, L saca is the loss function of the shadow adaptive conditional alignment; z0 is the latent space image of the initial image; y is the scene text embedding; c is the foreground image multimodal embedding; t is the time step of the diffusion process; m is the mask; is the generated image; T(·) is the transformation T; x0 is the initial image.
[0015] Further, the transformation T adjusts the color distribution of the foreground image multimodal embedding c to be consistent with the initial image x0, and the expression is:
[0016]
[0017] where μ c is the mean of the foreground image multimodal embedding, and σ c is the standard deviation of the foreground image multimodal embedding; is the mean of the initial image, is the standard deviation of the initial image.
[0018] Further, the noise prediction loss function L ∈-pred of the noise prediction network is expressed as:
[0019]
[0020] where z0 is the latent space image of the initial image; y is the scene text embedding; c is the foreground image multimodal embedding; t is the time step of the diffusion process; ∈ t is the random noise added at the current time step; z t is the noise image;
[0021] The noise image z t is obtained from the latent space image z0 of the initial image, and the expression is:
[0022]
[0023] where is the preset diffusion scheduling parameter.
[0024] Further, the decoding is to decode the latent space generated image through the decoder D to obtain the generated image, and the expression is:
[0025]
[0026] where is the generated image; is the predicted noise; is the preset diffusion scheduling parameter; z t is the noise image.
[0027] According to another aspect of the present invention, a text-controlled scene rendering graph generation system based on a foreground model is provided. The system includes a foreground modeling module, a foreground rendering module, and a text-driven effect image generation module;
[0028] The text-driven effect image generation module includes an image-text multi-modal encoder unit, a noise prediction network unit, and a shadow adaptive conditional alignment unit; the noise prediction network unit includes forward noise addition and backward denoising; the forward noise addition includes an input processing sub-unit, a step-by-step noise addition sub-unit, and a time step processing sub-unit; the backward denoising includes a noise prediction and update sub-unit, a control conditional feature interaction sub-unit, and an iterative generation sub-unit; the shadow adaptive conditional alignment unit includes a transformation sub-unit.
[0029] Further, the input of the foreground modeling module is the geometric file, material texture information, and parameter constraints of the product. Through the modeling engine, geometric reconstruction is performed, and the output is a three-dimensional model including geometric shape, material texture map, and lighting parameters.
[0030] Further, the input of the foreground rendering module is the three-dimensional model output by the foreground modeling module. By parsing the geometric shape, material texture map, and lighting parameters of the three-dimensional model, the image perspective and the placement position of the product scene are adjusted. The rendering engine calculates the surface texture, reflectivity, transparency, and light and shadow effects of the model according to the set lighting parameters and perspective parameters, and combines the physical rendering material technology to obtain the foreground target rendering diagram of the product.
[0031] Further, the input of the image-text multi-modal encoder unit is the foreground target rendering diagram, query vector, and text label. Through the encoder, the corresponding embedding is obtained, specifically including: extracting features through query vector self-attention calculation to obtain query vector information; extracting features through text label self-attention calculation to obtain text feature information; sharing the parameters of query vector self-attention calculation and text label self-attention calculation, and associating query vector information and text feature information; through a feed-forward neural network, obtaining the foreground image multi-modal embedding.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] (1) Improve design efficiency: Traditional 3D rendering methods require designers to manually construct the entire scene, consuming a large amount of time and effort, and requiring designers to have high-level professional skills such as 3D modeling, lighting control, and material setting. The text-driven image generation system based on the controllable generation network of the present invention can automatically generate high-quality images that meet the specified conditions by inputting simple text descriptions or setting conditions, such as product materials, textures, background environments, etc. This reduces the workload of designers, lowers the design threshold, and enables designers to focus on the design itself rather than the detailed adjustment of rendering technology, thus significantly improving design efficiency.
[0034] (2) Enhance the flexibility and customization of image generation: Through the image-text multimodal encoder and the shadow adaptive conditional alignment module, the present invention can adaptively adjust the lighting effect and color distribution of the generated image according to different text descriptions and foreground images, ensuring that the generated background image is visually consistent with the foreground image. This enables designers to quickly generate diverse product display effects through simple text instructions, meeting the display requirements in different scenarios. In addition, the system also supports flexible adjustment of scene descriptions, object positions, and postures, further enhancing the customization of image generation.
[0035] (3) Reduce the computational resource requirements: Traditional 3D rendering technologies have extremely high requirements for computational resources, especially when rendering high-resolution and complex scenes, which require powerful graphics processing units and a large amount of memory support. However, through the diffusion model-based image generation technology, the complex 3D modeling and rendering processes are reduced, and the dependence on computational resources is decreased. The system can generate high-quality images in a short time through the controllable generation network and the noise prediction network, reducing the rendering time and hardware costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flowchart of a method for generating a text-controlled scene rendering diagram based on a foreground model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] As Figure 1 shown, it is a method for generating a text-controlled scene rendering diagram based on a foreground model, and the specific steps include:
[0039] S1. Collect the design information of the product image to be generated. The foreground modeling module performs geometric reconstruction according to the geometric file, material texture information, and parameter constraints of the product to obtain the corresponding three-dimensional model;
[0040] S2. Input the three-dimensional model into the foreground rendering module for rendering processing to obtain the foreground target rendering diagram;
[0041] S3. The text-driven effect image generation module generates the rendering effect diagram of the product scene through the image-text multimodal encoder and the noise prediction network according to the scene description and the foreground target rendering diagram.
[0042] The specific steps for generating a rendered effect image of a product scene using the text-driven effect image generation module are as follows: The image-text multimodal encoder extracts features from the text tags and query vectors, and associates the text feature information and query vector information by sharing parameters. The query vector performs cross-attention calculation with the foreground object rendering image. Finally, a linear transformation of the visual-text multimodal features is performed through a feature map composed of a feed-forward neural network to obtain the foreground image multimodal embedding; the noise prediction network compresses the initial image into the latent space and gradually adds noise to obtain a noisy image; receives the scene description and the foreground image multimodal embedding from the image-text multimodal encoder, associates the foreground image multimodal embedding and the noisy image through the visual cross-attention mechanism, associates the scene text embedding and the noisy image through the text cross-attention mechanism, and performs noise prediction; denoises the noisy image according to the predicted noise to obtain the latent space generated image, and further decodes it to obtain the generated image.
[0043] The foreground modeling module converts the appearance and structural features of the product into a 3D geometric model by importing design data such as the geometric file, material texture information, and parameter constraints of the product. In the processing stage, the module uses a modeling engine for geometric reconstruction, optimizes the polygon mesh structure, repairs surface defects, and accurately maps the material and texture information to the 3D surface. At the same time, it combines the lighting simulation technology to pre-calculate the light and shadow effects. The model also extracts important detailed features of the product, such as edges, screw holes, or complex curves, to accurately represent the design intent of the product. Finally, the module outputs a high-precision 3D model containing geometric shapes, material maps, and lighting parameters, which can be exported in various formats, such as.png,.jpg,.mp4, to provide support for subsequent rendering, simulation, and production modules, and at the same time supports model adjustment and iterative optimization based on design feedback. The foreground modeling module performs 3D modeling on various objects. The 3D modeling engine used in this embodiment has low requirements for appearance attributes such as lighting, texture, and material, and pays more attention to the color of the 3D model, that is, the RGB attribute. Therefore, the operation and modeling process are relatively simple and convenient.
[0044] The foreground rendering module sets appropriate lighting, viewing angles, and materials to make the product image more realistic and meet the actual requirements. When performing rendering, the foreground rendering module does not pay attention to the background of the object, but only focuses on the object itself. It uses ray tracing rendering and rasterization rendering, and sets and adjusts the light source and angle according to the actual effect and needs. It uses the three-point lighting method to set the light source at three random positions above the object to minimize the shadow area and obtain a foreground object rendering image of the product with uniform lighting.
[0045] The text-driven effect image generation module is based on a controllable generation network. Through the controllable generation network framework, according to text descriptions such as design requirements, colors, textures, etc., it generates product backgrounds and effect diagrams. Through text driving, the system can quickly generate effect diagrams that meet specific requirements according to different input conditions. The text-driven module enables designers to guide image generation through simple text instructions, greatly improving design efficiency. The text-driven effect image generation module based on the controllable generation network is fine-tuned using the public internal design dataset 3D-FUTURE. The fine-tuned model samples generation samples with corresponding attributes based on the given foreground image and scene description. The text-driven effect image generation module includes an image-text multimodal encoder unit, a noise prediction network unit, and a shadow adaptive conditional alignment unit. The image-text multimodal encoder unit encodes the text description and the image features of the subject into a unified multimodal embedding, providing precise conditional control information for the generation process. The noise prediction network unit takes the multimodal embedding as the conditional input, gradually predicting and removing noise during the reverse diffusion process, making the generated image gradually approach the target distribution from a random distribution. The shadow adaptive conditional alignment unit further optimizes the realism and fidelity of the generated result by adjusting the color distribution and appearance consistency of the generated image and the conditional information in the local area. The cooperation of these three ensures that the generated effect diagrams not only meet the semantic requirements of the input conditions but also have high quality and consistency visually.
[0046] During the training process of the text-driven effect image generation module, after obtaining the generated image, taking shadow adaptive conditional alignment as the training objective, making the foreground image multimodal embedding and the background color of the initial image consistent through the transformation T, and then performing consistency constraints among the initial image, the generated image, and the foreground image multimodal embedding. The expression is:
[0047]
[0048] where, L saca is the loss function of shadow adaptive conditional alignment; z0 is the latent space image of the initial image; y is the scene text embedding; c is the foreground image multimodal embedding; t is the time step of the diffusion process; m is the mask; is the generated image; T(·) is the transformation T; x0 is the initial image.
[0049] The transformation T adjusts the color distribution of the foreground image multimodal embedding c to make it consistent with the initial image x0. The expression is:
[0050]
[0051] where, μ c is the mean of the foreground image multimodal embedding, and σ c is the standard deviation of the foreground image multimodal embedding; is the mean of the initial image, and is the standard deviation of the initial image.
[0052] The noise prediction loss function L of the noise prediction network ∈-pred has the following expression:
[0053]
[0054] where z0 is the latent space image of the initial image; y is the scene text embedding; c is the foreground image multi-modal embedding; t is the time step of the diffusion process; ∈ t is the random noise added at the current time step; z t is the noise image;
[0055] The noise image z t is obtained from the latent space image z0 of the initial image, and the expression is:
[0056]
[0057] where is the preset diffusion scheduling parameter.
[0058] The generated image is obtained by decoding the latent space generated image through the decoder D, and the expression is:
[0059]
[0060] where is the generated image; is the predicted noise; is the preset diffusion scheduling parameter; z t is the noise image.
[0061] This embodiment also discloses a system for generating a text-controlled scene rendering diagram based on a foreground model. The system includes a foreground modeling module, a foreground rendering module, and a text-driven effect image generation module; the text-driven effect image generation module includes an image-text multi-modal encoder unit, a noise prediction network unit, and a shadow adaptive conditional alignment unit; the noise prediction network unit includes forward noise addition and backward denoising; the forward noise addition includes an input processing subunit, a step-by-step noise addition subunit, and a time step processing subunit; the backward denoising includes a noise prediction and update subunit, a control condition feature interaction subunit, and an iterative generation subunit; the shadow adaptive conditional alignment unit includes a transformation subunit.
[0062] The input of the foreground modeling module is the geometric file, material texture information, and parameter constraints of the product. Through the modeling engine, geometric reconstruction is performed, and the output is a three-dimensional model including geometric shape, material texture mapping, and lighting parameters.
[0063] The input of the foreground rendering module is the 3D model output by the foreground modeling module. By parsing the geometric shape, material texture mapping, and lighting parameters of the 3D model, adjusting the image perspective and the placement position of the product scene, the rendering engine calculates the surface texture, reflectivity, transparency, and light and shadow effects of the model according to the set lighting parameters and perspective parameters, combined with the physical rendering material technology, to obtain the foreground target rendering of the product.
[0064] The inputs of the image-text multi-modal encoder unit are the foreground target rendering, query vector, and text label. The corresponding embeddings are obtained through the encoder, specifically including: feature extraction through query vector self-attention calculation to obtain query vector information; feature extraction through text label self-attention calculation to obtain text feature information; sharing the parameters of query vector self-attention calculation and text label self-attention calculation to associate query vector information and text feature information; and obtaining the foreground image multi-modal embedding through a feed-forward neural network.
[0065] By adjusting the scene description and conditional image, the method and system of this embodiment can complete the generation of scene renderings of multiple control target types, including foreground objects with a single object, foreground objects with multiple objects, variable positions and poses of foreground objects, and variable scene descriptions. For a foreground object with a single object, generating a complete scene image conditional on a single object rendering is the basic task of this embodiment. The conditional object is drawn into the image, and at the same time, a shadow suitable for connection with the global environment is provided. It can work well even when solving small objects such as chairs. For example, if the foreground object with a single object is a bed, the obtained scene rendering includes the bed, bedside table, corner table, and armchair, and the scene style is comfortable, the scene is well-lit, and the details are clear. For foreground objects with multiple objects, when the provided control condition is multiple foreground objects, the background of the multiple-object foreground can be generated. For the network structure, processing the combination of multiple objects is similar to processing a single object with multiple different parts. In the scene graph rendering function of foreground objects with multiple objects, for example, the design of a bedroom reveals a bedside table, a dressing table, a bed, and a cabinet. For the scene graph rendering with variable scene descriptions, by changing the characteristics of objects, colors, styles, materials, etc. in the scene description, the scene style and the appearance, colors, and material elements of objects in the background are modified. For variable positions and poses of foreground objects, for the target item bed, different poses can be adjusted.
[0066] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for generating a text-controlled scene rendering image based on a foreground model, characterized in that: The specific steps include: S1. Collect the design information of the product image to be generated. The foreground modeling module performs geometric reconstruction according to the geometry file, material texture information and parameter constraints in the design information to obtain the corresponding three-dimensional model. S2, inputting the three-dimensional model into a foreground rendering module for rendering processing to obtain a foreground target rendering image; S3, the text-driven effect image generation module generates a rendering effect image of the product scene according to the scene description and the foreground target rendering image through the image-text multimodal encoder and the noise prediction network.
2. The method for generating a text-controlled scene rendering image based on a foreground model according to claim 1, characterized in that: The specific steps of the text-driven effect image generation module generating a rendering effect image of a product scene are as follows: The image-text multimodal encoder extracts features from text labels and query vectors, and associates text feature information with query vector information by sharing parameters. The query vector and the foreground object rendering are cross-attentionally calculated. Finally, the multimodal features of the visual text are linearly transformed through feature mapping composed of a feedforward neural network to obtain the multimodal embedding of the foreground image. The noise prediction network compresses the initial image into the latent space and gradually adds noise to obtain a noisy image; Receive the scene description and the foreground image multimodal embedding from the image-text multimodal encoder, associate the foreground image multimodal embedding with the noisy image through the visual cross-attention mechanism, associate the scene text embedding with the noisy image through the text cross-attention mechanism, and perform noise prediction; denoise the noisy image according to the predicted noise to obtain a latent space generated image, and further decode to obtain the generated image.
3. The method for generating a text-controlled scene rendering image based on a foreground model according to claim 2, characterized in that: In the training process of the text-driven effect image generation module, after obtaining the generated image, the shadow adaptive conditional alignment is used as the training target, and the background color of the foreground image multimodal embedding and the initial image are made consistent through the transformation T, and then the consistency constraint is performed between the initial image, the generated image and the foreground image multimodal embedding, and the expression is: Among them, L saca is the loss function of shadow adaptive conditional alignment; z0 is the latent space image of the initial image; y is the scene text embedding; c is the multimodal embedding of the foreground image; t is the time step of the diffusion process; m is the mask; is the generated image; T(·) is the transformation T; x0 is the initial image.
4. The method for generating a text-controlled scene rendering image based on a foreground model according to claim 3 is characterized in that: The transformation T adjusts the color distribution of the foreground image multimodal embedding c to make it consistent with the initial image x0, expressed as: Among them, μ c is the mean of the multimodal embedding of the foreground image, σ c is the standard deviation of the multimodal embedding of the foreground image; is the mean of the initial image, is the standard deviation of the initial image.
5. The method for generating a text-controlled scene rendering image based on a foreground model according to claim 2, characterized in that: The noise prediction loss function L of the noise prediction network is ∈-pred The expression is: Where z0 is the latent space image of the initial image; y is the scene text embedding; c is the multimodal embedding of the foreground image; t is the time step of the diffusion process; ∈ t is the random noise added at the current time step; z t is a noise image; The noise image z t It is obtained from the latent space image z0 of the initial image, and the expression is: in, is the preset diffusion scheduling parameter.
6. The method for generating a text-controlled scene rendering image based on a foreground model according to claim 2, characterized in that: The decoding is to decode the latent space generated image through the decoder D to obtain the generated image, and the expression is: in, To generate images; To predict noise; is the preset diffusion scheduling parameter; t is a noise image.
7. A system for the method for generating a text-controlled scene rendering image based on a foreground model according to any one of claims 1 to 6, characterized in that: The system includes a foreground modeling module, a foreground rendering module and a text-driven effect image generation module; The text-driven effect image generation module includes an image-text multimodal encoder unit, a noise prediction network unit, and a shadow adaptive conditional alignment unit; the noise prediction network unit includes forward denoising and reverse denoising; The forward denoising includes an input processing subunit, a step-by-step denoising subunit and a time step processing subunit; the reverse denoising includes a noise prediction and update subunit, a control condition feature interaction subunit, and an iterative generation subunit; the shadow adaptive condition alignment unit includes a transformation subunit.
8. The system for generating a text-controlled scene rendering image based on a foreground model according to claim 7, characterized in that: The foreground modeling module takes as input the product's geometry file, material texture information and parameter constraints, performs geometry reconstruction through the modeling engine, and outputs a three-dimensional model including geometry, material mapping and lighting parameters.
9. The system for generating a text-controlled scene rendering image based on a foreground model according to claim 7, characterized in that: The input of the foreground rendering module is the three-dimensional model output by the foreground modeling module. By analyzing the geometric shape, material mapping and lighting parameters of the three-dimensional model, adjusting the image perspective and product scene placement position, the rendering engine calculates the model surface texture, reflectivity, transparency and light and shadow effects according to the set lighting parameters and perspective parameters combined with physical rendering material technology to obtain the foreground target rendering of the product.
10. The system for generating a text-controlled scene rendering image based on a foreground model according to claim 7, characterized in that: The input of the image-text multimodal encoder unit is a foreground target rendering, a query vector and a text label, and the corresponding embedding is obtained through the encoder, specifically including: extracting features through query vector self-attention calculation to obtain query vector information; extracting features through text label self-attention calculation to obtain text feature information; sharing parameters of query vector self-attention calculation and text label self-attention calculation, associating query vector information and text feature information; and obtaining foreground image multimodal embedding through a feedforward neural network.