Conditional guided image generation method and system with physical property control capability

By encoding and linearly modulating the illumination and height parameters, the problem of insufficient physical property control in existing image generation methods is solved, achieving consistent control of the illumination distribution and imaging scale of the generated image, and improving the controllability and consistency of the generation results.

CN122634840APending Publication Date: 2026-08-25UNIV OF SCI & TECH BEIJING +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610651013.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing image generation methods lack the ability to explicitly model and control the physical properties of images, making it difficult to precisely control the generated results, especially in terms of consistency in illumination distribution and imaging scale.

Method used

By constructing a conditional coding module to encode illumination and height parameters, a conditional feature vector is generated and introduced into the diffusion model. The feature linear modulation (FiLM) method is used to fuse the feature vector with image features, and the model is optimized by combining the joint loss function to achieve explicit control of illumination and height.

Benefits of technology

It significantly improves the illumination distribution, imaging scale, and physical interpretability of the generated images, enhances the consistency between the generated results and the input physical conditions, and enables flexible adjustment and stable generation under different physical conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122634840A_ABST
    Figure CN122634840A_ABST
Patent Text Reader

Abstract

The application provides a conditional guided image generation method and system with physical property regulation capability, and belongs to the technical field of image generation and computer vision, comprising the following steps: constructing training samples comprising image information and physical property information, wherein the physical property information comprises illumination parameters and height parameters; constructing a conditional encoding module to encode the illumination parameters and the height parameters to obtain a conditional feature vector; constructing a conditional guided diffusion model with physical property regulation capability, introducing the conditional feature vector into the diffusion model, fusing the conditional feature vector with image features through a FiLM mode, constructing a conditional guided diffusion generation process, and training and optimizing the model through a joint loss function; based on an initial image and target physical parameters input by a user, using the trained diffusion model to perform conditional reasoning to generate a visible light image satisfying specified physical constraints. The application can generate a conditional guided image with physical property regulation capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image generation and computer vision technology, and specifically refers to a condition-guided image generation method and system with the ability to adjust physical properties. Background Technology

[0002] With the development of deep learning technology, generative model-based image generation methods have been widely researched and applied in the field of computer vision. Among them, diffusion models have gradually become an important technical route for high-quality image generation due to their advantages in image generation quality, diversity, and stability. However, most existing image generation methods mainly rely on data-driven statistical learning, and their generation results are more of a fit to the distribution of training data, lacking the ability to explicitly model and control the physical properties of the image. Summary of the Invention

[0003] To address the technical problems existing in the prior art, the present invention provides a condition-guided image generation method and system with physical property control capabilities, the technical solution of which is as follows: On the one hand, a condition-guided image generation method with physical property control capability is provided, the method comprising: S1. Generate illumination parameter data corresponding to the original visible light image, and construct training samples including image information and physical attribute information, wherein the physical attribute information includes illumination parameters and height parameters; S2. Construct a conditional encoding module to encode the illumination parameters and height parameters to obtain conditional feature vectors; S3. Construct a condition-guided diffusion model with physical property control capabilities, introduce the conditional feature vector into the diffusion model, fuse it with image features through feature linear modulation (FiLM) to construct a condition-guided diffusion generation process, and train and optimize the model through a joint loss function to achieve a synergistic improvement in visual quality and physical consistency of the generated image. S4. Based on the initial image and target physical parameters input by the user, use the trained diffusion model to perform conditional inference and generate a visible light image that satisfies the specified physical constraints.

[0004] On the other hand, a condition-guided image generation system with physical property control capability is provided, the system comprising: A construction module is generated to generate corresponding illumination parameter data for the original visible light image and to construct training samples that include image information and physical attribute information, wherein the physical attribute information includes illumination parameters and height parameters. The conditional coding module is used to construct the conditional coding module, which encodes the illumination parameters and height parameters to obtain the conditional feature vector; A conditional guided diffusion model is used to construct a conditional guided diffusion model with physical property control capabilities. The conditional feature vector is introduced into the diffusion model and fused with image features through feature linear modulation (FiLM) to construct a conditional guided diffusion generation process. The model is trained and optimized through a joint loss function to achieve a synergistic improvement in visual quality and physical consistency of the generated image. The inference generation module is used to generate a visible light image that satisfies specified physical constraints by performing conditional inference using a trained diffusion model based on the initial image and target physical parameters input by the user.

[0005] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described condition-guided image generation method with physical property control capability.

[0006] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement the above-described condition-guided image generation method with physical property control capability.

[0007] The beneficial effects of the technical solution provided by this invention include at least the following: (1) To address the technical problems of existing image generation methods, such as difficulty in explicitly controlling physical properties and lack of physical consistency in the generated results, this invention proposes a condition-guided diffusion generation strategy based on physical condition coding. By jointly modeling illumination and height parameters and mapping the physical parameters to conditional feature representations aligned with the feature space of the generation network, the diffusion model can explicitly perceive and respond to external physical constraints during the generation process. This avoids the problem that traditional generation models rely solely on implicit statistical distributions and are difficult to precisely control the physical properties of the generated results, significantly improving the controllability of the generated images in terms of illumination distribution, imaging scale, and physical interpretability.

[0008] (2) To address the difficulty of modeling mixed continuous and discrete physical conditions, this invention designs a unified conditional encoding and feature modulation mechanism to achieve collaborative guidance for different types of physical attributes. Specifically, for continuous physical conditions such as height parameters, continuous embedding and nonlinear mapping are used for encoding; for discrete category conditions such as illumination parameters, learnable category embedding is used for representation, and the Feature Linear Modulation (FiLM) mechanism is used to transform different physical conditions into scaling and offset operations on intermediate features of the diffusion model, enabling the model to dynamically adjust its generation behavior at different time steps and different levels, thereby improving the stability and precision of conditional guidance.

[0009] (3) To address the issue of insufficient consistency between the generated results and physical conditions in the diffusion model, a joint loss function for physical consistency is introduced to constrain the model training process. Based on the traditional denoising loss function, a physical property consistency constraint based on illumination and height parameters is further introduced. By jointly optimizing the differences between the generated image and the real image in terms of brightness distribution, illumination response, and highly correlated imaging characteristics, the model can significantly enhance the consistency between the generated results and the input physical conditions while ensuring the generation quality, effectively suppressing physical property shifts or distortions.

[0010] (4) During the inference stage, the physical properties of the generated image can be flexibly adjusted through a condition-guided inverse denoising process. By inputting different illumination and height parameters and continuously applying conditional constraints during the inverse denoising process, images that meet different physical conditions can be generated without retraining the model, thus improving the flexibility and practical value of the method in real-world applications.

[0011] (5) The method of the present invention does not require the introduction of a complex physical simulation process, and can achieve controllable modeling of physical properties in a data-driven generative framework, thus exhibiting good versatility and scalability. The method can be flexibly extended to other physical property conditions (such as atmospheric parameters, imaging angles, or sensor parameters), and is applicable to visible light image generation, remote sensing image synthesis, and image generation tasks under multiple physical constraints, possessing high engineering practical value and promising application prospects. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a condition-guided image generation method with physical property control capability provided by an embodiment of the present invention; Figure 2 This is a general block diagram of a condition-guided image generation method with physical property control capability provided by an embodiment of the present invention; Figure 3(a) , 3(b) These are schematic diagrams of the characteristic linear modulation layer and conditional connection method used in the conditional coding module provided in the embodiments of the present invention. Figure 4(a) , 4(b) This is a schematic diagram of forward noise addition and reverse noise reduction of the diffusion model provided in the embodiments of the present invention; Figures 5(a) and 5(b) show the image generation results at different altitudes under original lighting and cloudy conditions, respectively, provided by the method of the present invention. Figures 6(a) and 6(b) show the image generation results at different altitudes under foggy and nighttime conditions provided by the method of the present invention. Figures 7(a) and 7(b) show the image generation results at different altitudes under clear weather and dusk conditions, respectively, provided by the method of the present invention. Figure 8 This is a block diagram of a condition-guided image generation system with physical property control capability provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0014] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0015] This invention provides a condition-guided image generation method with physical property control capabilities. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of this method is shown below. Figure 2 The diagram shown is an overall block diagram of the method. The processing flow may include the following steps: S1. Generate illumination parameter data corresponding to the original visible light image, and construct training samples including image information and physical attribute information, wherein the physical attribute information includes illumination parameters and height parameters; Optionally, S1 specifically includes: By surveying a large number of images under different lighting conditions, a set of independently adjustable lighting parameters were extracted and defined to describe the lighting intensity, direction, or distribution characteristics of a scene, including: brightness, contrast, tone shift, local lighting and shadow, and light and fog intensity, as shown in Table 1. Table 1 Illumination parameters

[0016] Through physical modeling or external computation, generate corresponding illumination parameter data for the original visible light image; Construct training samples that include image information and physical attribute information. The physical attribute information includes illumination parameters and altitude parameters. The altitude parameters represent imaging altitude, flight altitude, or continuous physical quantities related to the imaging scale.

[0017] S2. Construct a conditional encoding module to encode the illumination parameters and height parameters to obtain conditional feature vectors; Optionally, S2 specifically includes: Five lighting conditions are constructed based on the lighting parameters, corresponding to combinations of physical parameters for typical lighting scenarios: sunny, cloudy, dusk, foggy, and night. For discrete lighting conditions, the lighting parameters are categorized as discrete categories. This is used to represent different lighting states or lighting types. These lighting categories are achieved by combining various corresponding lighting parameters, as shown in Table 2. Table 2 Examples of illumination conditions under different combinations of physical parameters

[0018] First, represent the lighting category as a discrete index. ,in Indicates the total number of lighting categories; By using a lookup table embedding method, discrete illumination categories are mapped to illumination condition embedding vectors, the expression of which is:

[0019] in, This represents the learnable lighting category embedding matrix. Embedding dimensions for lighting conditions, Embedded for lighting conditions; Through the above modeling method, different lighting categories can form distinguishable conditional representations in the feature space, thereby guiding the diffusion model to learn the brightness distribution, color response and shadow characteristics under different lighting conditions. For continuous height conditions, let the input height parameter be... First, the height parameter is normalized to eliminate the influence of different height units and value ranges on model training, expressed as:

[0020] in, and These represent the minimum and maximum values ​​of the height parameter, respectively. Normalized height parameter The input is fed into the conditional coding module, where a high-level conditional embedding vector is obtained through multi-layer nonlinear mapping. Its expression is:

[0021] in, This represents a learnable mapping function used for height modeling. For the corresponding network parameters, For high-condition embedding; By using the above method, the continuously changing height conditions are mapped to a high-dimensional feature space, enabling the diffusion model to learn the implicit relationship between height changes and image scale, detail distribution, and imaging characteristics. In obtaining height conditional embedding Embedded with lighting conditions Then, the two are fused to form a unified representation of physical conditions, and the fusion method is expressed as follows:

[0022] in, Indicates feature concatenation operation; This represents a multi-layer learnable linear mapping function used for conditional fusion, which performs nonlinear mapping on the concatenated conditional features to obtain a unified conditional feature representation. The mapping function includes at least two layers of linear transformation and corresponding nonlinear activation operations. This is the final conditional feature vector generated.

[0023] By combining the continuous height condition and the discrete illumination condition in the above-mentioned joint modeling method, the conditional encoding module of this embodiment of the invention can simultaneously characterize continuous physical changes and discrete physical states, thereby providing stable and controllable physical constraints for the subsequent condition-guided diffusion generation process.

[0024] S3. Construct a condition-guided diffusion model with physical property control capabilities, introduce the conditional feature vector into the diffusion model, fuse it with image features through feature linear modulation (FiLM) to construct a condition-guided diffusion generation process, and train and optimize the model through a joint loss function to achieve a synergistic improvement in visual quality and physical consistency of the generated image. Optionally, in step S3, the conditional feature vector is introduced into the diffusion model and fused with image features through Feature-wise Linear Modulation (FiLM), specifically including: Let the intermediate features of a certain layer in the diffusion model be represented as:

[0025] in, In the diffusion model, the first... Each network layer, C, W, H These represent the number of feature channels and the spatial resolution, respectively. The conditional feature vector is input into the FiLM generator to generate modulation parameters that match the number of feature channels in the current network layer. For the Input Conditional eigenvectors Two learnable modulation parameters are generated using the FiLM generator: scaling parameter. and offset parameters The mathematical expressions for both are:

[0026] in, This represents the conditional feature vector obtained by fusing illumination parameters and height parameters. and These represent the learnable mapping functions used to generate the scaling and offset parameters, respectively. After obtaining the modulation parameters, linear modulation is performed on the corresponding feature channels in the diffusion model. The modulation process is represented as follows:

[0027] in, Represents the feature map; Through the above FiLM modulation operation, physical conditions are applied to the intermediate features of the diffusion model in the form of scaling and offset, thereby achieving adaptive adjustment of physical properties to feature distribution without changing the network structure. The above FiLM modulation operation is repeatedly performed in multiple UNet encoding and decoding layers of the diffusion model, so that the illumination and height conditions continuously constrain the feature evolution process at different stages of diffusion and reverse denoising, thereby guiding the model to gradually form image structure and appearance characteristics that meet the corresponding physical conditions during the generation process.

[0028] The network structure of the FiLM layer and its connection with the residual module in the diffusion model are shown in Figures 3(a) and 3(b). By using the FiLM method to construct conditional inputs and fuse them with image features, the stability of the generative model can be guaranteed while effectively controlling continuous height conditions and discrete illumination conditions.

[0029] Optionally, as shown in Figures 4(a) and 4(b), the diffusion model employs a generation mechanism of progressive noise addition and progressive noise reduction, including: During the training phase, the real visible light images are first subjected to forward noise addition: Let the real visible light image be In time step The following is a forward noise-adding process to obtain the corresponding noisy image. The calculation method is as follows:

[0030] in, For random noise sampled from a standard Gaussian distribution, The time-step related noise coefficient is determined according to the preset noise scheduling strategy and is used to control the ratio of original image information to noise components in different diffusion stages; In the reverse denoising process, the diffusion model uses the noisy image... Corresponding time step and the conditional feature vector generated by the conditional coding module As a joint input, the noise corresponding to the current time step is predicted, and the prediction process is expressed as follows:

[0031] in, Represented by neural network parameters The noise prediction network characterized by the conditional feature vector This is used to apply physical properties to guide the generation process of diffusion models, enabling the models to perceive and respond to external physical constraints during the denoising process.

[0032] Optionally, a denoising loss function is constructed based on the noise prediction results. This is used to measure the error between the model's predicted noise and the actual noise, and it is defined as follows:

[0033] in, This is used to constrain the denoising accuracy and training stability of the diffusion model in the reverse denoising process, ensuring that the model can accurately recover image structural information at different time steps; Based on the above denoising loss, a physical condition consistency constraint loss function is introduced to enhance the consistency between the generated image and the input physical conditions: Given conditional eigenvectors In this case, the image generated by the reverse denoising process is denoted as Based on the lighting parameters and height parameters, a physical property mapping operator is constructed. Implemented by a learnable lightweight convolutional neural network, it is used to extract physical attribute features related to illumination distribution and imaging scale from real and generated images, respectively. The physical condition consistency constraint loss is defined as:

[0034] in, This represents a mapping function that maps the input real image to the physical property space; This indicates that, given the eigenvectors... Below, the physical properties of the target are represented by the corresponding real images; The diffusion model is trained and optimized using a joint loss function, the expression of which is:

[0035] in, and These are the weighting coefficients used to balance the denoising loss from the physical condition consistency constraint loss.

[0036] S4. Based on the initial image and target physical parameters input by the user, use the trained diffusion model to perform conditional inference and generate a visible light image that satisfies the specified physical constraints.

[0037] Optionally, S4 specifically includes: During the inference phase, the system first receives the user's input of an initial image and target physical parameters, including target illumination parameters and imaging height parameters. The initial image and the target physical parameters are input into the conditional coding module to generate a conditional feature vector, which is then fused with the image features using Feature Linear Modulation (FiLM) to guide the generation behavior of the diffusion model during subsequent inference. Subsequently, a random noise image is obtained by adding noise to the initial image. The latent space representation of the image is obtained through a variational automatic inference encoder and serves as the initial state for the back inference of the diffusion model. The random noise image follows a preset Gaussian distribution. The random noise image, the corresponding time step information, and the conditional feature tensor are input into the trained diffusion model. The back denoising process is executed step by step according to a preset denoising scheduling strategy to obtain the predicted latent space image representation. After decoding, a generated image controlled by height and illumination parameters is finally obtained. The generated image is consistent with the target physical parameters input by the user in terms of overall brightness distribution, illumination direction consistency, shadow morphology, and geometric or radiometric characteristics related to imaging height, thereby realizing condition-guided image generation with physical property control capabilities. In each time step of the reverse denoising process, while the diffusion model performs noise prediction and denoising updates, it continuously modulates the intermediate features within the model using the conditional feature tensor. The modulation method is FiLM modulation, which enables the model to continuously perceive and respond to the constraint information of the illumination parameters and height parameters during the gradual removal of noise, thereby guiding the generated result to gradually approach the target state that meets the physical conditions at the structural and appearance levels.

[0038] To verify the effectiveness of the condition-guided image generation method with physical property control capabilities provided in this embodiment of the invention, its performance was tested using the following evaluation metrics: (1) CLIP Score: This measure of semantic consistency is achieved by calculating the similarity between the generated image and the target image in the shared semantic space of the multimodal foundation model CLIP. Its value range is [value range missing]. The higher the value, the better; (2) Root Mean Square Error (MSE): This measures the reconstruction error by calculating the average squared difference between the generated image and the reference image at the pixel level. Its value range is... The smaller the value, the better; (3) Peak Signal-to-Noise Ratio (PENR): Based on the ratio between pixel error and the maximum pixel value of the image, it evaluates the image reconstruction quality in decibels (dB), and its value range is... The higher the value, the better; (4) Structural Similarity Index (SSIM): This index assesses the structural similarity between the generated image and the reference image in terms of brightness, contrast, and structure. Its value range is [range missing]. The higher the value, the better; (5) Normalized Global Error (EGARS): This is a metric used to evaluate the quality of remote sensing images. It is commonly used to assess the performance of image processing or compression algorithms, as well as the quality of remote sensing images. A lower value is better.

[0039] In this embodiment, the ISPRS Potsdam and Toronto City datasets are used for training, validation, and testing. This dataset contains 22,852 aerial RGB images, of which 18,304 are used as the training set, 6 as the test set, and the remainder as the validation set. The ISPRS Potsdam dataset, collected in the Potsdam region of Germany, offers images with a resolution as high as 5 cm / pixel. Each image is taken from a top-down perspective, providing richly detailed RGB aerial images of urban areas. The Toronto City dataset consists of large-scale aerial images of the Toronto area in Canada, with image resolutions ranging from 5 to 10 cm / pixel. The RGB images in this dataset come from different times and regions, showcasing a complete and diverse urban environment, including dense urban areas, suburbs, highways, and green spaces.

[0040] Both datasets contain urban information, which allows the model to learn more detailed and complex scene information during training.

[0041] During training, the weight coefficients of the loss function are set to... The initial learning rate of the diffusion model is set to The training run consisted of 500 epochs, with a batch size of 12.

[0042] Table 3 shows the image quantization results at different heights generated under the original lighting conditions in this embodiment: Table 3. Quantization results of images generated at different heights under original illumination.

[0043] Table 4 shows the image quantization results at different altitudes under cloudy conditions generated in this embodiment: Table 4. Quantization results of images generated at different altitudes under cloudy conditions.

[0044] Table 5 shows the image quantization results at different altitudes under foggy conditions generated in this embodiment: Table 5. Quantization results of images generated at different altitudes under foggy conditions.

[0045] Table 6 shows the image quantization results at different altitudes under nighttime conditions generated in this embodiment: Table 6. Quantization results of images generated at different altitudes under nighttime conditions.

[0046] Table 7 shows the image quantization results at different altitudes under clear weather conditions generated in this embodiment: Table 7. Quantization results of images generated at different altitudes under clear weather conditions.

[0047] Table 8 shows the image quantization results at different altitudes under the dusk conditions generated in this embodiment: Table 8. Quantization results of images generated at different altitudes under twilight conditions.

[0048] The image results generated in this embodiment at different altitudes under original lighting and cloudy conditions are shown in Figure 5(a) and Figure 5(b); the image results at different altitudes under foggy and nighttime conditions are shown in Figure 6(a) and Figure 6(b); and the image results at different altitudes under sunny and twilight conditions are shown in Figure 7(a) and Figure 7(b).

[0049] The quantitative evaluation results and visual comparison results of this embodiment show that the condition-guided image generation method with physical property control capability described in this embodiment can stably generate visible light images corresponding to the physical conditions given user-input target illumination and imaging height conditions. The generated results maintain a high degree of consistency with the input conditions in terms of overall brightness distribution, shadow representation, and structural and detail features related to imaging height. Furthermore, under different combinations of physical conditions, the generated images exhibit good generation quality in terms of sharpness, structural integrity, and visual coherence, indicating that the method of this embodiment has good technical effects in terms of both physical condition controllability and generation effect stability.

[0050] like Figure 8 As shown, this embodiment of the invention also provides a condition-guided image generation system with physical property control capabilities, the system comprising: A construction module 810 is used to generate illumination parameter data corresponding to the original visible light image and construct training samples including image information and physical attribute information, wherein the physical attribute information includes illumination parameters and height parameters. Conditional coding module 820 is used to construct a conditional coding module to encode illumination parameters and height parameters to obtain conditional feature vectors; Conditional guided diffusion model 830 is used to construct a conditional guided diffusion model with physical property control capability. The conditional feature vector is introduced into the diffusion model and fused with image features through feature linear modulation (FiLM) to construct a conditional guided diffusion generation process. The model is trained and optimized through a joint loss function to achieve a synergistic improvement in visual quality and physical consistency of the generated image. The inference generation module 840 is used to generate a visible light image that satisfies specified physical constraints by performing conditional inference using a trained diffusion model based on the initial image and target physical parameters input by the user.

[0051] The present invention provides a condition-guided image generation system with physical property control capability, the functional structure of which corresponds to the condition-guided image generation method with physical property control capability provided in the present invention, and will not be described again here.

[0052] Figure 9 This is a schematic diagram of the structure of an electronic device 900 provided in an embodiment of the present invention. The electronic device 900 may vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 901 and one or more memories 902. The memory 902 stores at least one instruction, which is loaded and executed by the processor 901 to implement the steps of the condition-guided image generation method with physical attribute control capability described above.

[0053] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the conditionally guided image generation method with physical attribute control capabilities. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0054] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0055] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A condition-guided image generation method with physical property control capability, characterized in that, The method includes: S1. Generate illumination parameter data corresponding to the original visible light image, and construct training samples including image information and physical attribute information, wherein the physical attribute information includes illumination parameters and height parameters; S2. Construct a conditional encoding module to encode the illumination parameters and height parameters to obtain conditional feature vectors; S3. Construct a condition-guided diffusion model with physical property control capabilities, introduce the conditional feature vector into the diffusion model, fuse it with image features through feature linear modulation (FiLM) to construct a condition-guided diffusion generation process, and train and optimize the model through a joint loss function to achieve a synergistic improvement in visual quality and physical consistency of the generated image. S4. Based on the initial image and target physical parameters input by the user, use the trained diffusion model to perform conditional inference and generate a visible light image that satisfies the specified physical constraints.

2. The method according to claim 1, characterized in that, S1 specifically includes: By surveying a large number of images under different lighting conditions, a set of independently adjustable lighting parameters were extracted and defined to describe the lighting intensity, direction or distribution characteristics of the scene, including: brightness, contrast, tone shift, local lighting and shadow, and light and fog intensity. Through physical modeling or external computation, generate corresponding illumination parameter data for the original visible light image; Construct training samples that include image information and physical attribute information. The physical attribute information includes illumination parameters and altitude parameters. The altitude parameters represent imaging altitude, flight altitude, or continuous physical quantities related to the imaging scale.

3. The method according to claim 1, characterized in that, S2 specifically includes: Five lighting conditions are constructed based on the lighting parameters, corresponding to combinations of physical parameters for typical lighting scenarios: sunny, cloudy, dusk, foggy, and night. For discrete lighting conditions, the lighting parameters are categorized as discrete categories. , used to represent different lighting states or lighting types, these lighting categories are achieved by a combination of corresponding lighting parameters; First, represent the lighting category as a discrete index. ,in Indicates the total number of lighting categories; By using a lookup table embedding method, discrete illumination categories are mapped to illumination condition embedding vectors, the expression of which is: in, This represents the learnable lighting category embedding matrix. Embedding dimensions for lighting conditions, Embedded for lighting conditions; Through the above modeling method, different lighting categories can form distinguishable conditional representations in the feature space, thereby guiding the diffusion model to learn the brightness distribution, color response and shadow characteristics under different lighting conditions. For continuous height conditions, let the input height parameter be... First, the height parameter is normalized to eliminate the influence of different height units and value ranges on model training, expressed as: in, and These represent the minimum and maximum values ​​of the height parameter, respectively. Normalized height parameter The input is fed into the conditional coding module, where a high-level conditional embedding vector is obtained through multi-layer nonlinear mapping. Its expression is: in, This represents a learnable mapping function used for height modeling. For the corresponding network parameters, For high-condition embedding; By using the above method, the continuously changing height conditions are mapped to a high-dimensional feature space, enabling the diffusion model to learn the implicit relationship between height changes and image scale, detail distribution, and imaging characteristics. In obtaining height conditional embedding Embedded with lighting conditions Then, the two are fused to form a unified representation of physical conditions, and the fusion method is expressed as follows: in, Indicates feature concatenation operation; This represents a multi-layer learnable linear mapping function used for conditional fusion, which performs nonlinear mapping on the concatenated conditional features to obtain a unified representation of the conditional features. The mapping function includes at least two layers of linear transformation and corresponding nonlinear activation operations. This is the final generated conditional feature vector.

4. The method according to claim 1, characterized in that, In step S3, the conditional feature vector is introduced into the diffusion model and fused with image features through Feature Linear Modulation (FiLM). Specifically, this includes: Let the intermediate features of a certain layer in the diffusion model be represented as: in, In the diffusion model, the first... Each network layer, C, W, H These represent the number of feature channels and the spatial resolution, respectively. The conditional feature vector is input into the FiLM generator to generate modulation parameters that match the number of feature channels in the current network layer. For the Input Conditional eigenvectors Two learnable modulation parameters are generated using the FiLM generator: scaling parameter. and offset parameters The mathematical expressions for both are: in, This represents the conditional feature vector obtained by fusing illumination parameters and height parameters. and These represent the learnable mapping functions used to generate the scaling and offset parameters, respectively. After obtaining the modulation parameters, linear modulation is performed on the corresponding feature channels in the diffusion model. The modulation process is represented as follows: in, Represents the feature map; Through the above FiLM modulation operation, physical conditions are applied to the intermediate features of the diffusion model in the form of scaling and offset, thereby achieving adaptive adjustment of physical properties to feature distribution without changing the network structure. The above FiLM modulation operation is repeatedly performed in multiple UNet encoding and decoding layers of the diffusion model, so that the illumination and height conditions continuously constrain the feature evolution process at different stages of diffusion and reverse denoising, thereby guiding the model to gradually form image structure and appearance characteristics that meet the corresponding physical conditions during the generation process.

5. The method according to claim 1, characterized in that, The diffusion model employs a generation mechanism of progressive noise addition and progressive noise removal, including: During the training phase, the real visible light images are first subjected to forward noise addition: Let the real visible light image be In time step The following is a forward noise-adding process to obtain the corresponding noisy image. The calculation method is as follows: in, For random noise sampled from a standard Gaussian distribution, The time-step related noise coefficient is determined according to the preset noise scheduling strategy and is used to control the ratio of original image information to noise components in different diffusion stages; In the reverse denoising process, the diffusion model uses the noisy image... Corresponding time step and the conditional feature vector generated by the conditional coding module As a joint input, the noise corresponding to the current time step is predicted, and the prediction process is expressed as follows: in, Represented by neural network parameters The noise prediction network characterized by the conditional feature vector This is used to apply physical properties to guide the generation process of diffusion models, enabling the models to perceive and respond to external physical constraints during the denoising process.

6. The method according to claim 5, characterized in that, Based on the noise prediction results, a denoising loss function is constructed. This is used to measure the error between the model's predicted noise and the actual noise, and it is defined as follows: in, This is used to constrain the denoising accuracy and training stability of the diffusion model in the reverse denoising process, ensuring that the model can accurately recover image structural information at different time steps; Based on the above denoising loss, a physical condition consistency constraint loss function is introduced to enhance the consistency between the generated image and the input physical conditions: Given conditional eigenvectors In this case, the image generated by the reverse denoising process is denoted as Based on the lighting parameters and height parameters, a physical property mapping operator is constructed. Implemented by a learnable lightweight convolutional neural network, it is used to extract physical attribute features related to illumination distribution and imaging scale from real and generated images, respectively. The physical condition consistency constraint loss is defined as: in, This represents a mapping function that maps the input real image to the physical property space; This indicates that, given the eigenvectors... Below, the physical properties of the target are represented by the corresponding real images; The diffusion model is trained and optimized using a joint loss function, the expression of which is: in, and These are the weighting coefficients used to balance the denoising loss from the physical condition consistency constraint loss.

7. The method according to claim 1, characterized in that, S4 specifically includes: During the inference phase, the system first receives the user's input of an initial image and target physical parameters, including target illumination parameters and imaging height parameters. The initial image and the target physical parameters are input into the conditional coding module to generate a conditional feature vector, which is then fused with the image features using Feature Linear Modulation (FiLM) to guide the generation behavior of the diffusion model during subsequent inference. Subsequently, a random noise image is obtained by adding noise to the initial image. The latent space representation of the image is obtained through a variational automatic inference encoder and serves as the initial state for the back inference of the diffusion model. The random noise image follows a preset Gaussian distribution. The random noise image, the corresponding time step information, and the conditional feature tensor are input into the trained diffusion model. The back denoising process is executed step by step according to a preset denoising scheduling strategy to obtain the predicted latent space image representation. After decoding, a generated image controlled by height and illumination parameters is finally obtained. The generated image is consistent with the target physical parameters input by the user in terms of overall brightness distribution, illumination direction consistency, shadow morphology, and geometric or radiometric characteristics related to imaging height, thereby realizing condition-guided image generation with physical property control capabilities. In each time step of the reverse denoising process, while the diffusion model performs noise prediction and denoising updates, it continuously modulates the intermediate features within the model using the conditional feature tensor. The modulation method is FiLM modulation, which enables the model to continuously perceive and respond to the constraint information of the illumination parameters and height parameters during the gradual removal of noise, thereby guiding the generated result to gradually approach the target state that meets the physical conditions at the structural and appearance levels.

8. A condition-guided image generation system with physical property control capabilities, characterized in that, The system includes: A construction module is generated to generate corresponding illumination parameter data for the original visible light image and to construct training samples that include image information and physical attribute information, wherein the physical attribute information includes illumination parameters and height parameters. The conditional coding module is used to construct the conditional coding module, which encodes the illumination parameters and height parameters to obtain conditional feature vectors; A conditional guided diffusion model is used to construct a conditional guided diffusion model with physical property control capabilities. The conditional feature vector is introduced into the diffusion model and fused with image features through feature linear modulation (FiLM) to construct a conditional guided diffusion generation process. The model is trained and optimized through a joint loss function to achieve a synergistic improvement in visual quality and physical consistency of the generated image. The inference generation module is used to generate visible light images that satisfy specified physical constraints by performing conditional inference using a trained diffusion model based on the initial image and target physical parameters input by the user.

9. An electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, characterized in that, The processor loads and executes at least one instruction to implement the condition-guided image generation method with physical property control capability as described in any one of claims 1-7.

10. A computer-readable storage medium storing at least one instruction, characterized in that, The at least one instruction is loaded and executed by the processor to implement the condition-guided image generation method with physical property control capability as described in any one of claims 1-7.