Remote sensing image generation method, device and equipment based on multi-level description
By employing a multi-level description-based remote sensing image generation method, a trained noise prediction network is used to recover the structure of remote sensing images under the control of multi-level description. This solves the problems of target omission and unreasonable spatial layout caused by single-level text prompts, and achieves efficient and refined remote sensing image generation.
Patent Information
- Application Number
- CN202610892492.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-19
- Publication Date
- 2026-08-25
AI Technical Summary
In existing remote sensing image generation technologies, single-level text prompts lead to missing targets, unreasonable spatial layout, insufficient local details, and semantic deviations, making it difficult to generate high-quality remote sensing images.
A multi-level description method is adopted to obtain preprocessed remote sensing image samples and their corresponding multi-level descriptions, including scene concepts, spatial relationships and ground feature descriptions. Remote sensing images are generated by training a noise prediction network, and the image structure is recovered from the noise latent variables using the information of the multi-level description. The generation result is controlled by the target multi-level description.
It improves the structural rationality, semantic consistency, and detail fidelity of the generated remote sensing image results, reduces the semantic ambiguity of single-level text prompts, and enhances generation efficiency and fine-grained control capabilities.
Smart Images

Figure CN122636784A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of remote sensing image processing and computer vision, and in particular to a method, apparatus and device for generating remote sensing images based on multi-level description. Background Technology
[0002] Remote sensing imagery can reflect spatial information such as land cover, urban morphology, transportation infrastructure, natural topography, and human activities. Remote sensing image generation technologies can be used for data augmentation, scene simulation, change analysis, urban planning, and emergency drills. With the development of text-generated image diffusion models, generating remote sensing imagery using textual conditions has become an important direction for intelligent production of remote sensing data.
[0003] However, remote sensing imagery is characterized by its vertical observation perspective, large differences in ground feature scale, dense spatial relationships, and highly specialized geographic semantics. In existing technologies, single-level text prompts typically only express a scene category or a general description, leading to issues such as missing targets, illogical spatial layout, insufficient local detail, and semantic bias in the generated remote sensing imagery. Summary of the Invention
[0004] In view of this, the present invention provides a method, apparatus and device for generating remote sensing images based on multi-level description, which improves the quality of remote sensing image generation results.
[0005] According to one aspect of the present invention, a method for generating remote sensing images based on multi-level description is provided, the method comprising: Obtain preprocessed remote sensing image samples, and obtain multi-level descriptions corresponding to the remote sensing image samples. The multi-level descriptions include at least one of scene concept description, spatial relationship description, and ground feature description. The multi-level descriptions are connected into a unified condition text according to a preset connection order, and the unified condition text is encoded to obtain multi-level description condition features. The remote sensing image samples are compressed into sample latent variables in the latent space. Random noise is added to the sample latent variables to obtain noise latent variables. The noise latent variables, the denoising time step, and the multi-level descriptive condition features are input into the initial noise prediction network to obtain the predicted noise. The initial noise prediction network is updated according to the loss value between the predicted noise and the random noise until a preset stopping condition is reached to obtain the trained noise prediction network. A multi-level description of the target is obtained, which includes at least one level of the target scene concept description, target spatial relationship description, and target feature description. The multi-level descriptions are connected into a unified target conditional text according to a preset connection order. The unified target conditional text is encoded to obtain the multi-level description conditional features. Random latent variables are obtained. The random latent variables, the denoising time step, and the multi-level description conditional features are input into the noise prediction network to obtain target prediction noise. According to the denoising time step, the random latent variables are denoised based on the target prediction noise to output target latent spatial variables. The target latent spatial variables are decoded to generate a remote sensing image generation result.
[0006] According to another aspect of the present invention, a remote sensing image generation apparatus based on multi-level description is provided, the apparatus comprising: The description acquisition module is used to acquire preprocessed remote sensing image samples and acquire multi-level descriptions corresponding to the remote sensing image samples. The multi-level descriptions include at least one of scene concept description, spatial relationship description, and ground feature description. The text connection and encoding module is used to connect the multi-level descriptions into a unified condition text according to a preset connection order, and encode the unified condition text to obtain multi-level description condition features; The training module is used to compress the remote sensing image samples into sample latent variables in the latent space, add random noise to the sample latent variables to obtain noise latent variables, input the noise latent variables, the denoising time step and the multi-level descriptive condition features into the initial noise prediction network to obtain the predicted noise, update the initial noise prediction network according to the loss value between the predicted noise and the random noise, until a preset stopping condition is reached to obtain the trained noise prediction network. The image generation module is used to acquire a multi-level description of the target, which includes at least one of the following: a target scene concept description, a target spatial relationship description, and a target feature description. The multi-level descriptions are connected into a unified target conditional text according to a preset connection order. The unified target conditional text is encoded to obtain multi-level target description conditional features. Random latent variables are acquired. The random latent variables, the denoising time step, and the multi-level target description conditional features are input into a noise prediction network to obtain target prediction noise. Based on the target prediction noise and the denoising time step, the random latent variables are denoised to output target latent spatial variables. The target latent spatial variables are decoded to generate a remote sensing image generation result.
[0007] According to another aspect of the present invention, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described remote sensing image generation method based on multi-level description.
[0008] According to another aspect of the present invention, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor, when executing the program, implements the above-described remote sensing image generation method based on multi-level description.
[0009] By means of the above technical solution, the present invention provides a remote sensing image generation method, apparatus, and device based on multi-level description. First, a preprocessed remote sensing image sample is acquired, and a multi-level description corresponding to the remote sensing image sample is obtained. The multi-level description includes at least one of scene concept description, spatial relationship description, and ground feature description. Second, the multi-level descriptions are connected into a unified conditional text according to a preset connection order, and the unified conditional text is encoded to obtain multi-level description conditional features. Then, the remote sensing image sample is compressed into sample latent variables in a latent space, random noise is added to the sample latent variables to obtain noise latent variables, and the noise latent variables, denoising time steps, and multi-level description conditional features are input into an initial noise prediction network to obtain predicted noise. The predicted noise is then further determined based on the loss value between the predicted noise and the random noise. The initial noise prediction network is trained until a preset stopping condition is met, resulting in a fully trained noise prediction network. Finally, a multi-level target description is obtained, comprising at least one level of target scene concept description, target spatial relationship description, and target feature description. This multi-level description is then connected into a unified target conditional text according to a preset connection order. The unified target conditional text is encoded to obtain multi-level target description conditional features. Random latent variables are obtained, and the random latent variables, the denoising time step, and the multi-level target description conditional features are input into the noise prediction network to obtain target prediction noise. Based on the target prediction noise, the random latent variables are denoised according to the denoising time step, and target latent spatial variables are output. The target latent spatial variables are then decoded to generate a remote sensing image. Through the technical solution of this invention, on the one hand, since the noise prediction network is trained, it learns to recover the structure of remote sensing image samples from noise latent variables under the control of multi-level description, that is, it learns the mapping relationship between multi-level description and remote sensing image samples. Therefore, the encoded target multi-level description conditional features are input into the noise prediction network, enabling the noise prediction network to utilize the information of the target multi-level description and obtain target prediction noise under the constraints of the target multi-level description. The target prediction noise is then gradually removed from the random latent variables to recover the remote sensing image structure and generate the remote sensing image generation result. On the other hand, the target multi-level description is directly connected through the target unified conditional text, and the encoded... The target multi-level description conditional features are then input into the noise prediction network, meaning the conditional control input is provided all at once, rather than in stages, thus improving generation efficiency. Furthermore, the existing single-level text prompts are extended to multi-level target descriptions, and the number of levels of the multi-level target description can be selected according to the required level of refinement of the remote sensing image generation results. This not only controls the quality of the remote sensing image generation results according to the required level of refinement, improving generation efficiency, but also reduces the target missing and spatial layout distortion caused by the semantic ambiguity of single-level text prompts, improving the structural rationality, semantic consistency, and detail fidelity of the remote sensing image generation results.
[0010] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of this application. In the drawings: Figure 1 The diagram illustrates a flowchart of a remote sensing image generation method based on multi-level description provided by an embodiment of the present invention. Figure 2 A flowchart illustrating another remote sensing image generation method based on multi-level description provided by an embodiment of the present invention is shown. Figure 3 This diagram illustrates the structure of a remote sensing image generation device based on multi-level description provided in an embodiment of the present invention. Figure 4 This invention provides a schematic diagram of another remote sensing image generation device based on multi-level description. Figure 5 This diagram illustrates a process of training an initial noise prediction network using remote sensing image samples and multi-level descriptions, as provided in an embodiment of the present invention. Figure 6 This illustration shows a process for generating remote sensing images from multi-level target descriptions, as provided in an embodiment of the present invention. Detailed Implementation
[0012] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0013] This embodiment provides a remote sensing image generation method based on multi-level description, such as Figure 1 As shown, the method includes: 101. Obtain the preprocessed remote sensing image sample, and obtain the multi-level description corresponding to the remote sensing image sample. The multi-level description includes at least one of scene concept description, spatial relationship description, and ground feature description.
[0014] In this embodiment, an initial remote sensing image sample is acquired and preprocessed to obtain a preprocessed remote sensing image sample. The preprocessing includes, but is not limited to: for multi-channel initial remote sensing image samples, channel selection and format conversion are performed to form an RGB three-channel image; for initial remote sensing image samples of different sizes, slicing, cropping and padding are performed to meet the size requirements of the image encoder, so that in step 103 of the embodiment, the image encoder is used to compress the remote sensing image sample into a latent variable in the latent space.
[0015]
[0016] in, This is the initial remote sensing image sample. For preprocessing functions, and These represent the height and width of the image, respectively. This is a remote sensing image sample.
[0017] It should be noted that the more levels included in the multi-level description, the richer the semantic information of the unified conditional text is across multiple levels, rather than simply being richer at a single level. Therefore, in step 103 of the embodiment, it can provide stronger constraints for the initial noise prediction network, and the mapping relationship between the learned multi-level description and the remote sensing image sample is more accurate. In this way, when the noise prediction model is applied in step 104 of the embodiment, the generated remote sensing image will have stronger structural rationality, semantic consistency with the target multi-level description, and detail fidelity, that is, the higher the refinement of the remote sensing image generation result.
[0018] However, if the number of levels of multi-level description is always set to three regardless of the required level of detail in the generated remote sensing image results, it will greatly increase the computational complexity and reduce the efficiency of the initial noise prediction network during training. It will also greatly increase the computational complexity and reduce the efficiency of the noise prediction network during inference applications. Therefore, in this embodiment, the number of levels of multi-level description can be selected according to the required level of detail in the generated remote sensing image results. Similarly, step 104 of this embodiment can select the target level of multi-level description according to the required level of detail in the generated remote sensing image results, thereby balancing the level of detail, computational complexity, and efficiency of the generated remote sensing image results.
[0019] 102. Connect the multi-level descriptions into a unified condition text according to a preset connection order, and encode the unified condition text to obtain multi-level description condition features.
[0020] In this embodiment, the preset connection order is determined according to the purpose of use. For example, the multi-level description includes a scene concept description. Spatial Relationship Description Description of ground features Multi-level description For example, if the preset connection order is scene concept description, spatial relationship description, and feature description (i.e., including but not limited to connecting scene concept description first, then spatial relationship description, and finally feature description), then the text will be directly connected according to the preset connection order to obtain a unified conditional text. for:
[0021] in, This indicates a join operation.
[0022] The unified conditional text is encoded using a text encoder to obtain multi-level descriptive conditional features.
[0023]
[0024] in, For text encoders, For multi-level description of conditional features, when the text encoder is the CLIP text encoder or other pre-trained language models, This is a vector sequence representation for unified conditional text. This vector sequence representation simultaneously incorporates three levels of semantics: scene concept, spatial relationship, and geographic features.
[0025] 103. Compress the remote sensing image samples into sample latent variables in the latent space, add random noise to the sample latent variables to obtain noise latent variables, input the noise latent variables, the denoising time step and the multi-level descriptive condition features into the initial noise prediction network to obtain predicted noise, update the initial noise prediction network according to the loss value between the predicted noise and the random noise, until a preset stopping condition is reached to obtain the trained noise prediction network.
[0026] For this embodiment, as Figure 5 The diagram shows a process of training an initial noise prediction network using remote sensing image samples and multi-level descriptions, where the multi-level descriptions include scene concept descriptions, spatial relationship descriptions, and ground feature descriptions.
[0027] Adding random noise to the latent variables of the sample includes sampling the denoising time steps according to a preset noise schedule, and adding random noise to the latent variables of the sample based on the cumulative noise control coefficient corresponding to the denoising time steps.
[0028] By performing diffusion modeling in the latent space rather than the original image space, computational complexity and memory consumption can be reduced, while preserving the main structural information of the remote sensing image samples.
[0029] Preferably, the initial noise prediction network further includes a cross-attention layer so that after the noise latent variables, denoising time steps, and multi-level descriptive conditional features are input into the initial noise prediction network, the cross-attention mechanism of the cross-attention layer is used to constrain the noise prediction, thereby obtaining the predicted noise. Specifically, the noise latent variables provide a query vector. Multi-level descriptive conditional features provide key vectors Sum value vector The noise prediction network calculates cross-attention using query vectors, key vectors, and value vectors. This cross-attention is used to predict noise. In this process, the initial noise prediction network establishes a correspondence between latent noise variables and multi-level descriptive features. Under the constraints of these features, predicted noise is obtained. If the multi-level description includes scene concept description, spatial relationship description, and feature description, then the predicted noise is obtained under the combined constraints of these three features. Furthermore, this cross-attention layer in the noise prediction network also constrains the network during denoising by the target multi-level description. For example, if the target multi-level description includes target scene concept description, target spatial relationship description, and target feature description, then the noise prediction network is constrained by these three descriptions during denoising.
[0030] Preferably, updating the initial noise prediction network includes updating the parameters of the initial noise prediction network, including but not limited to full parameter fine-tuning, low-rank adaptation, and text inversion. To improve training efficiency, low-rank adaptation can be used to fine-tune the cross-attention layers of the initial noise prediction network. For the weights to be fine-tuned... Introducing a low-rank matrix and The updated weights Represented as:
[0031] in, It is a low-rank dimension. is the scaling factor. Using low-rank adaptation reduces the number of training parameters and allows the model to focus on learning the correspondence between the unified conditional text and the spatial structure of remote sensing images.
[0032] 104. Obtain a multi-level description of the target, which includes at least one of the following: target scene concept description, target spatial relationship description, and target feature description. Connect the multi-level descriptions of the target into a unified conditional text according to the preset connection order. Encode the unified conditional text of the target to obtain the multi-level description conditional features of the target. Obtain random latent variables. Input the random latent variables, the denoising time step, and the multi-level description conditional features of the target into the noise prediction network to obtain target prediction noise. According to the denoising time step, denoise the random latent variables based on the target prediction noise to output target latent spatial variables. Decode the target latent spatial variables to generate a remote sensing image generation result.
[0033] In this embodiment, the reasoning application step involves a multi-level description of the target, which corresponds to the generated result of the remote sensing image to be generated. For example... Figure 6 The diagram shows a process for generating remote sensing images from multi-level target descriptions, where the multi-level target descriptions include target scene concept descriptions, target spatial relationship descriptions, and target feature descriptions.
[0034] This invention provides a remote sensing image generation method, apparatus, and device based on multi-level description. First, preprocessed remote sensing image samples are acquired, and multi-level descriptions corresponding to the remote sensing image samples are obtained. These multi-level descriptions include at least one of scene concept description, spatial relationship description, and ground feature description. Second, the multi-level descriptions are connected into a unified conditional text according to a preset connection order, and the unified conditional text is encoded to obtain multi-level description conditional features. Then, the remote sensing image samples are compressed into sample latent variables in a latent space, and random noise is added to these sample latent variables to obtain noise latent variables. The noise latent variables, denoising time steps, and the multi-level description conditional features are input into an initial noise prediction network to obtain predicted noise. The initial noise prediction network is updated based on the loss value between the predicted noise and the random noise. The noise prediction network is trained until a preset stopping condition is met, resulting in a fully trained noise prediction network. Finally, a multi-level target description is obtained, comprising at least one level of target scene concept description, target spatial relationship description, and target feature description. This multi-level description is connected into a unified target conditional text according to a preset connection order. The unified target conditional text is encoded to obtain multi-level target description conditional features. Random latent variables are obtained, and the random latent variables, the denoising time step, and the multi-level target description conditional features are input into the noise prediction network to obtain target prediction noise. Based on the target prediction noise, the random latent variables are denoised according to the denoising time step, and target latent spatial variables are output. The target latent spatial variables are decoded to generate the remote sensing image generation result. Through the technical solution of this invention, on the one hand, since the noise prediction network is trained, it learns to recover the structure of remote sensing image samples from noise latent variables under the control of multi-level description, that is, it learns the mapping relationship between multi-level description and remote sensing image samples. Therefore, the encoded target multi-level description conditional features are input into the noise prediction network, enabling the noise prediction network to utilize the information of the target multi-level description and obtain target prediction noise under the constraints of the target multi-level description. The target prediction noise is then gradually removed from the random latent variables to recover the remote sensing image structure and generate the remote sensing image generation result. On the other hand, the target multi-level description is directly connected through the target unified conditional text, and the encoded... The target multi-level description conditional features are then input into the noise prediction network, meaning the conditional control input is provided all at once, rather than in stages, thus improving generation efficiency. Furthermore, the existing single-level text prompts are extended to multi-level target descriptions, and the number of levels of the multi-level target description can be selected according to the required level of refinement of the remote sensing image generation results. This not only controls the quality of the remote sensing image generation results according to the required level of refinement, improving generation efficiency, but also reduces the target missing and spatial layout distortion caused by the semantic ambiguity of single-level text prompts, improving the structural rationality, semantic consistency, and detail fidelity of the remote sensing image generation results.
[0035] Furthermore, as a refinement and extension of the specific implementation methods described above, and to fully illustrate the specific implementation process in this embodiment, another remote sensing image generation method based on multi-level description is provided, such as... Figure 2 As shown, the method includes: 201. Obtain the preprocessed remote sensing image samples.
[0036] The specific implementation steps in this embodiment are the same as those in step 101 of the embodiment, and will not be repeated here.
[0037] 202. Determine the preset description refinement level corresponding to the remote sensing image sample, and determine the multi-level description corresponding to the remote sensing image sample based on the preset description refinement level. The multi-level description includes at least one of scene concept description, spatial relationship description, and ground feature description.
[0038] In this embodiment, the scene concept description includes: scene category; the spatial relationship description includes at least one of the relationships between ground features, such as adjacent, surrounding, crossing, containing, and located; and the ground feature element description includes at least one of the elements of buildings, roads, vegetation, water bodies, vehicles, bare land, and site elements.
[0039] As one implementation method, determining the multi-level description corresponding to the remote sensing image sample based on the preset description refinement level includes: if the preset description refinement level is high, and the level of the multi-level description corresponding to the remote sensing image sample is determined to be three, then the multi-level description corresponding to the remote sensing image sample includes scene concept description, spatial relationship description, and ground feature description; if the preset description refinement level is medium, and the level of the multi-level description corresponding to the remote sensing image sample is determined to be two, then the multi-level description corresponding to the remote sensing image sample includes any two levels among scene concept description, spatial relationship description, and ground feature description; if the preset description refinement level is low, and the level of the multi-level description corresponding to the remote sensing image sample is determined to be one, then the multi-level description corresponding to the remote sensing image sample includes any one level among scene concept description, spatial relationship description, and ground feature description.
[0040] For example, multi-level descriptions include scene concept descriptions, spatial relationship descriptions, and feature descriptions. The scene concept description is: park; the spatial relationship description is: a remote sensing image sample of a park with a lake in the center, surrounded by trees and roads; and the feature descriptions are: buildings, city, roads, trees, rooftops, plazas, and paths.
[0041] It should be noted that the higher the preset description refinement level, the higher the required level of refinement.
[0042] As another implementation, determining the multi-level description corresponding to the remote sensing image sample based on the preset description refinement level includes: if the preset description refinement level is high, determining the number of levels of the multi-level description corresponding to the remote sensing image sample to be three, and determining the first preset description length and the number of first preset description fields for each level of description, then determining that the multi-level description corresponding to the remote sensing image sample includes scene concept description, spatial relationship description, and ground feature description, and determining the first preset description length and the number of first preset description fields for scene concept description, spatial relationship description, and ground feature description corresponding to the high preset description refinement level; if the preset description refinement level is medium, determining the number of levels of the multi-level description corresponding to the remote sensing image sample to be two, and determining the second preset description length and the number of second preset description fields for each level of description, then determining that the multi-level description corresponding to the remote sensing image sample includes scene concept description, spatial relationship description, and ground feature description. The multi-level description corresponding to the remote sensing image sample is determined to be any two levels among concept description, spatial relationship description, and feature description, and the second preset description length and the number of second preset description fields corresponding to any two levels of the preset description refinement level are determined. If the preset description refinement level is low, the level of the multi-level description corresponding to the remote sensing image sample is determined to be one, and the third preset description length and the number of third preset description fields for each level of description are determined. Then, the multi-level description corresponding to the remote sensing image sample is determined to include any one level among scene concept description, spatial relationship description, and feature description, and the third preset description length and the number of third preset description fields corresponding to any one level of the preset description refinement level are determined. The first description refinement corresponding to the first preset description length and the first preset description field number is greater than the second description refinement corresponding to the second preset description length and the second preset description field number, and the second description refinement is greater than the third description refinement corresponding to the third preset description length and the third preset description field number.
[0043] It should be noted that the preset description refinement level corresponds not only to the number of levels in the multi-level description, but also to the description refinement within each level of the multi-level description: For the preset description refinement level of high: for scene concept description, the first preset description length and the number of first preset description fields are used; for spatial relationship description, the first preset description length and the number of first preset description fields are used; for feature description, the first preset description length and the number of first preset description fields are used. For a preset description refinement level of medium (e.g., any second level included in a multi-level description is a scene concept description and a spatial relationship description): for a scene concept description, the second preset description length and the number of second preset description fields are used; for a spatial relationship description, the second preset description length and the number of second preset description fields are used. For a preset description refinement level of low (e.g., any level included in a multi-level description is a scene concept description): For scene concept descriptions, use the third preset description length and the number of third preset description fields of the scene concept description.
[0044] 203. Connect the multi-level descriptions into a unified condition text according to a preset connection order, and encode the unified condition text to obtain multi-level description condition features.
[0045] In this embodiment, the step of connecting the multi-level descriptions into a unified conditional text according to a preset connection order and encoding the unified conditional text to obtain multi-level description conditional features includes: obtaining the level identifier corresponding to each level description, adding the level identifier to the header of the corresponding level description, connecting the level descriptions with the added level identifier into a unified conditional text according to a preset connection order, and encoding the unified conditional text using a text encoder to obtain multi-level description conditional features.
[0046] Each level of description is pre-defined, with a corresponding level identifier. For example, the level identifier for the scene concept description is: The level identifier corresponding to the spatial relationship description is: The level identifiers corresponding to the descriptions of geographic features are: Level identifiers are used to preserve the semantic source of the three-level description, enabling the unified conditional text to be processed by the text encoder and to express the multi-level structure of the remote sensing scene.
[0047] The default connection order is determined based on the intended use; for example, multi-level descriptions may include scene concept descriptions. Spatial Relationship Description Description of ground features Multi-level description For example, if the preset connection order is scene concept description, spatial relationship description, and feature description (i.e., including but not limited to connecting scene concept description first, then spatial relationship description, and finally feature description), then the level identifier is added to the header of the corresponding level description. According to the preset connection order, the level descriptions with the added level identifier are directly connected into a unified conditional text to obtain the unified conditional text. for:
[0048] in, This indicates a join operation.
[0049] The specific implementation method for encoding the unified conditional text using a text encoder to obtain multi-level descriptive conditional features is the same as step 102 in the embodiment, and will not be repeated here.
[0050] 204. Compress the remote sensing image samples into sample latent variables in the latent space, add random noise to the sample latent variables to obtain noise latent variables, input the noise latent variables, the denoising time step and the multi-level descriptive condition features into the initial noise prediction network to obtain predicted noise, update the initial noise prediction network according to the loss value between the predicted noise and the random noise, until a preset stopping condition is reached to obtain the trained noise prediction network.
[0051] In this embodiment, an image encoder is used. remote sensing image samples Sample latent variables compressed into the latent space :
[0052] For latent variables in the sample Add random noise The noise latent variables are obtained, specifically: set up For the noise reduction time step, From step 1 to step 2 The cumulative noise control factor of the step. For random noise sampled from a standard Gaussian distribution, the sample latent variable is... In the Step noise latent variables for:
[0053] Wherein, this formula represents the time step in any denoising process. The noise latent variables corresponding to the noise intensity are directly obtained and used to train the initial noise prediction network.
[0054] The noise latent variable Noise reduction time step and the multi-level descriptive condition features Input Initial Noise Prediction Network The predicted noise is obtained as follows:
[0055] With random noise For real noise, use real noise The mean square error between the error and the predicted noise is used as the loss function to calculate the loss value. :
[0056] Determine if the loss value is less than a preset threshold. If not, adjust the model parameters of the initial noise prediction network to update the initial noise prediction network until the loss value is less than the preset threshold, reaching the preset stopping condition, and obtain the trained noise prediction network.
[0057] It should be noted that this initial noise prediction network learns to recover the remote sensing image structure from the noise latent variables under a multi-level description. Specifically, based on the predicted noise, the noise latent variables are denoised, and the initial noise prediction network outputs latent spatial variables, which are then processed by the image decoder. The latent space variables are decoded to recover the corresponding remote sensing image samples.
[0058] 205. Obtain a multi-level description of the target, which includes at least one of the following: target scene concept description, target spatial relationship description, and target feature description. Connect the multi-level descriptions of the target into a unified conditional text according to the preset connection order. Encode the unified conditional text of the target to obtain the multi-level description conditional features of the target. Obtain random latent variables. Input the random latent variables, the denoising time step, and the multi-level description conditional features of the target into the noise prediction network to obtain target prediction noise. According to the denoising time step, denoise the random latent variables based on the target prediction noise to output target latent spatial variables. Decode the target latent spatial variables to generate a remote sensing image generation result.
[0059] In this embodiment, obtaining the target multi-level description includes: determining the target preset description refinement level corresponding to the remote sensing image generation result to be generated; and determining the target multi-level description corresponding to the remote sensing image generation result to be generated based on the target preset description refinement level. Specifically, the methods for determining the target preset description refinement level corresponding to the remote sensing image generation result to be generated and determining the target multi-level description corresponding to the remote sensing image generation result to be generated based on the target preset description refinement level are the same as steps 202 in the embodiment, and will not be repeated here.
[0060] For example, a multi-level description of a target includes a description of the target scene concept. Description of target spatial relationships Description of target features Then the target is described in multiple levels. for:
[0061] Target Unification Condition Text for:
[0062] Target multi-level description condition features for:
[0063] From random latent variables Initially, based on a noise prediction network Stepwise denoising yields the target latent space variables for:
[0064] The remote sensing image generation result is obtained from the image decoder. for:
[0065] This invention provides a remote sensing image generation method, apparatus, and device based on multi-level description. First, preprocessed remote sensing image samples are acquired, and multi-level descriptions corresponding to the remote sensing image samples are obtained. These multi-level descriptions include at least one of scene concept description, spatial relationship description, and ground feature description. Second, the multi-level descriptions are connected into a unified conditional text according to a preset connection order, and the unified conditional text is encoded to obtain multi-level description conditional features. Then, the remote sensing image samples are compressed into sample latent variables in a latent space, and random noise is added to these sample latent variables to obtain noise latent variables. The noise latent variables, denoising time steps, and the multi-level description conditional features are input into an initial noise prediction network to obtain predicted noise. The initial noise prediction network is updated based on the loss value between the predicted noise and the random noise. The noise prediction network is trained until a preset stopping condition is met, resulting in a fully trained noise prediction network. Finally, a multi-level target description is obtained, comprising at least one level of target scene concept description, target spatial relationship description, and target feature description. This multi-level description is connected into a unified target conditional text according to a preset connection order. The unified target conditional text is encoded to obtain multi-level target description conditional features. Random latent variables are obtained, and the random latent variables, the denoising time step, and the multi-level target description conditional features are input into the noise prediction network to obtain target prediction noise. Based on the target prediction noise, the random latent variables are denoised according to the denoising time step, and target latent spatial variables are output. The target latent spatial variables are decoded to generate the remote sensing image generation result. Through the technical solution of this invention, on the one hand, since the noise prediction network is trained, it learns to recover the structure of remote sensing image samples from noise latent variables under the control of multi-level description, that is, it learns the mapping relationship between multi-level description and remote sensing image samples. Therefore, the encoded target multi-level description conditional features are input into the noise prediction network, enabling the noise prediction network to utilize the information of the target multi-level description and obtain target prediction noise under the constraints of the target multi-level description. The target prediction noise is then gradually removed from the random latent variables to recover the remote sensing image structure and generate the remote sensing image generation result. On the other hand, the target multi-level description is directly connected through the target unified conditional text, and the encoded... The target multi-level description conditional features are then input into the noise prediction network, meaning the conditional control input is provided all at once, rather than in stages, thus improving generation efficiency. Furthermore, the existing single-level text prompts are extended to multi-level target descriptions, and the number of levels of the multi-level target description can be selected according to the required level of refinement of the remote sensing image generation results. This not only controls the quality of the remote sensing image generation results according to the required level of refinement, improving generation efficiency, but also reduces the target missing and spatial layout distortion caused by the semantic ambiguity of single-level text prompts, improving the structural rationality, semantic consistency, and detail fidelity of the remote sensing image generation results.
[0066] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this invention provides a remote sensing image generation device based on multi-level description, such as... Figure 3 As shown, the device includes: a description acquisition module 31, a text connection and encoding module 32, a training module 33, and an image generation module 34; The description acquisition module 31 is used to acquire preprocessed remote sensing image samples and acquire multi-level descriptions corresponding to the remote sensing image samples. The multi-level descriptions include at least one of scene concept description, spatial relationship description, and ground feature description. The text connection and encoding module 32 is used to connect the multi-level descriptions into a unified condition text according to a preset connection order, and encode the unified condition text to obtain multi-level description condition features; Training module 33 is used to compress the remote sensing image samples into sample latent variables in the latent space, add random noise to the sample latent variables to obtain noise latent variables, input the noise latent variables, the denoising time step and the multi-level descriptive condition features into the initial noise prediction network to obtain predicted noise, update the initial noise prediction network according to the loss value between the predicted noise and the random noise, until a preset stopping condition is reached to obtain the trained noise prediction network. The image generation module 34 is used to acquire a multi-level description of the target, which includes at least one of the following: a target scene concept description, a target spatial relationship description, and a target feature description. The multi-level description is connected into a unified target conditional text according to a preset connection order. The unified target conditional text is encoded to obtain multi-level target description conditional features. Random latent variables are acquired. The random latent variables, the denoising time step, and the multi-level target description conditional features are input into the noise prediction network to obtain target prediction noise. Based on the target prediction noise and the denoising time step, the random latent variables are denoised to output target latent spatial variables. The target latent spatial variables are decoded to generate a remote sensing image generation result.
[0067] Accordingly, in order to obtain a multi-level description corresponding to the remote sensing image sample, such as Figure 4 As shown, the description acquisition module 31 includes: a first determining unit 311 and a second determining unit 312; The first determining unit 311 is used to determine the preset description refinement level corresponding to the remote sensing image sample; The second determining unit 312 is used to determine a multi-level description corresponding to the remote sensing image sample based on the preset description refinement level.
[0068] Accordingly, in order to determine the multi-level description corresponding to the remote sensing image sample based on the preset description refinement level, the second determining unit 312 is configured to: if the preset description refinement level is high, determine that the level of the multi-level description corresponding to the remote sensing image sample is three, then determine that the multi-level description corresponding to the remote sensing image sample includes scene concept description, spatial relationship description, and ground feature description; if the preset description refinement level is medium, determine that the level of the multi-level description corresponding to the remote sensing image sample is two, then determine that the multi-level description corresponding to the remote sensing image sample includes any two levels among scene concept description, spatial relationship description, and ground feature description; if the preset description refinement level is low, determine that the level of the multi-level description corresponding to the remote sensing image sample is one, then determine that the multi-level description corresponding to the remote sensing image sample includes any one level among scene concept description, spatial relationship description, and ground feature description.
[0069] Accordingly, in order to determine the multi-level description corresponding to the remote sensing image sample based on the preset description refinement level, the second determining unit 312 is configured to: if the preset description refinement level is high, determine that the number of levels of the multi-level description corresponding to the remote sensing image sample is three, and determine the first preset description length and the number of first preset description fields for each level of description; then determine that the multi-level description corresponding to the remote sensing image sample includes scene concept description, spatial relationship description, and ground feature description; and determine the first preset description length and the number of first preset description fields for the scene concept description, the first preset description length and the number of first preset description fields for the spatial relationship description, and the first preset description length and the number of first preset description fields for the ground feature description corresponding to the high preset description refinement level; if the preset description refinement level is medium, determine that the number of levels of the multi-level description corresponding to the remote sensing image sample is two, and determine the second preset description length and the number of second preset description fields for each level of description; then determine that the multi-level description package corresponding to the remote sensing image sample includes... The multi-level description corresponding to the remote sensing image sample includes any two levels of scene concept description, spatial relationship description, and feature description, and determines the second preset description length and the number of second preset description fields corresponding to any two levels of the preset description refinement level. If the preset description refinement level is low, the level of the multi-level description corresponding to the remote sensing image sample is determined to be one, and the third preset description length and the number of third preset description fields for each level of description are determined. Then, the multi-level description corresponding to the remote sensing image sample includes any one level of scene concept description, spatial relationship description, and feature description, and determines the third preset description length and the number of third preset description fields corresponding to any one level of the preset description refinement level. The first description refinement corresponding to the first preset description length and the first preset description field number is greater than the second description refinement corresponding to the second preset description length and the second preset description field number, and the second description refinement is greater than the third description refinement corresponding to the third preset description length and the third preset description field number.
[0070] Accordingly, in order to connect the multi-level descriptions into a unified conditional text according to a preset connection order, and to encode the unified conditional text to obtain multi-level description conditional features, the text connection and encoding module 32 is used to obtain the level identifier corresponding to each level description, add the level identifier to the header of the corresponding level description, connect the level descriptions with the added level identifier into a unified conditional text according to a preset connection order, and encode the unified conditional text using a text encoder to obtain multi-level description conditional features.
[0071] Accordingly, the description acquisition module 31, the scene concept description includes: scene category, the spatial relationship description includes at least one of the relationships between ground features such as adjacency, surrounding, crossing, containing, and located, and the ground feature element description includes at least one of the elements such as buildings, roads, vegetation, water bodies, vehicles, bare land, and site elements.
[0072] Accordingly, in order to obtain a multi-level description of the target, the image generation module 34 is used to determine the target preset description refinement level corresponding to the remote sensing image generation result to be generated; and to determine the target multi-level description corresponding to the remote sensing image generation result to be generated based on the target preset description refinement level.
[0073] It should be noted that other corresponding descriptions of the functional units involved in the multi-level description-based remote sensing image generation device provided in this embodiment can be found in [reference]. Figures 1 to 2 The corresponding description will not be repeated here.
[0074] Based on the above, Figures 1 to 2 Accordingly, this embodiment also provides a storage medium, which may be volatile or non-volatile, storing a computer program that, when executed by a processor, implements the above-described method. Figures 1 to 2 The remote sensing image generation method based on multi-level description is shown.
[0075] Based on this understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of the present invention.
[0076] Based on the above, Figures 1 to 2 The method shown and Figure 3 , Figure 4 To achieve the above objectives, the present application also provides a computer device, specifically a personal computer, server, network device, etc., as shown in the illustrated embodiment. This computer device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 2 The remote sensing image generation method based on multi-level description is shown.
[0077] Optionally, the computer device may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0078] Those skilled in the art will understand that the computer device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0079] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned computer device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the non-volatile storage medium, as well as communication with other hardware and software in the information processing entity device.
[0080] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.
[0081] This invention provides a remote sensing image generation method, apparatus, and device based on multi-level description. First, preprocessed remote sensing image samples are acquired, and multi-level descriptions corresponding to the remote sensing image samples are obtained. These multi-level descriptions include at least one of scene concept description, spatial relationship description, and ground feature description. Second, the multi-level descriptions are connected into a unified conditional text according to a preset connection order, and the unified conditional text is encoded to obtain multi-level description conditional features. Then, the remote sensing image samples are compressed into sample latent variables in a latent space, and random noise is added to these sample latent variables to obtain noise latent variables. The noise latent variables, denoising time steps, and the multi-level description conditional features are input into an initial noise prediction network to obtain predicted noise. The initial noise prediction network is updated based on the loss value between the predicted noise and the random noise. The noise prediction network is trained until a preset stopping condition is met, resulting in a fully trained noise prediction network. Finally, a multi-level target description is obtained, comprising at least one level of target scene concept description, target spatial relationship description, and target feature description. This multi-level description is connected into a unified target conditional text according to a preset connection order. The unified target conditional text is encoded to obtain multi-level target description conditional features. Random latent variables are obtained, and the random latent variables, the denoising time step, and the multi-level target description conditional features are input into the noise prediction network to obtain target prediction noise. Based on the target prediction noise, the random latent variables are denoised according to the denoising time step, and target latent spatial variables are output. The target latent spatial variables are decoded to generate the remote sensing image generation result. Through the technical solution of this invention, on the one hand, since the noise prediction network is trained, it learns to recover the structure of remote sensing image samples from noise latent variables under the control of multi-level description, that is, it learns the mapping relationship between multi-level description and remote sensing image samples. Therefore, the encoded target multi-level description conditional features are input into the noise prediction network, enabling the noise prediction network to utilize the information of the target multi-level description and obtain target prediction noise under the constraints of the target multi-level description. The target prediction noise is then gradually removed from the random latent variables to recover the remote sensing image structure and generate the remote sensing image generation result. On the other hand, the target multi-level description is directly connected through the target unified conditional text, and the encoded... The target multi-level description conditional features are then input into the noise prediction network, meaning the conditional control input is provided all at once, rather than in stages, thus improving generation efficiency. Furthermore, the existing single-level text prompts are extended to multi-level target descriptions, and the number of levels of the multi-level target description can be selected according to the required level of refinement of the remote sensing image generation results. This not only controls the quality of the remote sensing image generation results according to the required level of refinement, improving generation efficiency, but also reduces the target missing and spatial layout distortion caused by the semantic ambiguity of single-level text prompts, improving the structural rationality, semantic consistency, and detail fidelity of the remote sensing image generation results.
[0082] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or they can be located in one or more apparatuses different from this embodiment, with corresponding changes. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0083] The serial numbers used above are for descriptive purposes only and do not represent the superiority or inferiority of the implementation scenarios. The above disclosures are merely a few specific implementation scenarios of the present invention; however, the present invention is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A remote sensing image generation method based on multi-level description, characterized in that, The method includes: Obtain preprocessed remote sensing image samples, and obtain multi-level descriptions corresponding to the remote sensing image samples. The multi-level descriptions include at least one of scene concept description, spatial relationship description, and ground feature description. The multi-level descriptions are connected into a unified condition text according to a preset connection order, and the unified condition text is encoded to obtain multi-level description condition features. The remote sensing image samples are compressed into sample latent variables in the latent space. Random noise is added to the sample latent variables to obtain noise latent variables. The noise latent variables, the denoising time step, and the multi-level descriptive condition features are input into the initial noise prediction network to obtain the predicted noise. The initial noise prediction network is updated according to the loss value between the predicted noise and the random noise until a preset stopping condition is reached to obtain the trained noise prediction network. A multi-level description of the target is obtained, which includes at least one level of the target scene concept description, target spatial relationship description, and target feature description. The multi-level descriptions are connected into a unified target conditional text according to a preset connection order. The unified target conditional text is encoded to obtain the multi-level description conditional features. Random latent variables are obtained. The random latent variables, the denoising time step, and the multi-level description conditional features are input into the noise prediction network to obtain target prediction noise. According to the denoising time step, the random latent variables are denoised based on the target prediction noise to output target latent spatial variables. The target latent spatial variables are decoded to generate a remote sensing image generation result.
2. The method according to claim 1, characterized in that, The acquisition of the multi-level description corresponding to the remote sensing image sample includes: Determine the preset description refinement level corresponding to the remote sensing image sample; The multi-level description corresponding to the remote sensing image sample is determined based on the preset description refinement level.
3. The method according to claim 2, characterized in that, The step of determining the multi-level description corresponding to the remote sensing image sample based on the preset description refinement level includes: If the preset description refinement level is high, and the level of the multi-level description corresponding to the remote sensing image sample is determined to be three, then the multi-level description corresponding to the remote sensing image sample is determined to include scene concept description, spatial relationship description and ground feature description. If the preset description refinement level is medium, and the level of the multi-level description corresponding to the remote sensing image sample is determined to be two, then the multi-level description corresponding to the remote sensing image sample is determined to include any two levels among scene concept description, spatial relationship description, and ground feature description. If the preset description refinement level is low, and the level of the multi-level description corresponding to the remote sensing image sample is determined to be one, then the multi-level description corresponding to the remote sensing image sample is determined to include any one of the following levels: scene concept description, spatial relationship description, and ground feature description.
4. The method according to claim 2, characterized in that, The step of determining the multi-level description corresponding to the remote sensing image sample based on the preset description refinement level includes: If the preset description refinement level is high, the number of levels of the multi-level description corresponding to the remote sensing image sample is determined to be three, and the first preset description length and the number of first preset description fields for each level of description are determined. Then, the multi-level description corresponding to the remote sensing image sample is determined to include scene concept description, spatial relationship description, and ground feature description. The first preset description length and the number of first preset description fields for scene concept description, spatial relationship description, and ground feature description are determined to correspond to the high preset description refinement level. If the preset description refinement level is medium, the level of the multi-level description corresponding to the remote sensing image sample is determined to be two, and the second preset description length and the number of second preset description fields of each level description are determined. Then, the multi-level description corresponding to the remote sensing image sample is determined to include any second level among scene concept description, spatial relationship description and ground feature description, and the second preset description length and the number of second preset description fields of any second level description corresponding to the preset description refinement level of medium are determined. If the preset description refinement level is low, the level of the multi-level description corresponding to the remote sensing image sample is determined to be one, and the third preset description length and the number of third preset description fields for each level of description are determined. Then, the multi-level description corresponding to the remote sensing image sample is determined to include any one of scene concept description, spatial relationship description, and ground feature description. And the third preset description length and the number of third preset description fields for any one level of description corresponding to the preset description refinement level are determined. Wherein, the first description refinement corresponding to the first preset description length and the first preset description field number is greater than the second description refinement corresponding to the second preset description length and the second preset description field number, and the second description refinement is greater than the third description refinement corresponding to the third preset description length and the third preset description field number.
5. The method according to claim 1, characterized in that, The step of connecting the multi-level descriptions into a unified conditional text according to a preset connection order, and encoding the unified conditional text to obtain multi-level description conditional features includes: Obtain the level identifier corresponding to each level description, add the level identifier to the header of the corresponding level description, and connect the level descriptions with the added level identifier into a unified conditional text according to a preset connection order; The unified conditional text is encoded using a text encoder to obtain multi-level descriptive conditional features.
6. The method according to claim 1, characterized in that, The scene concept description includes: scene category; the spatial relationship description includes at least one of the relationships between ground features, such as adjacency, surrounding, crossing, containing, or located; and the ground feature element description includes at least one of the following elements: building, road, vegetation, water body, vehicle, bare land, or site element.
7. The method according to claim 1, characterized in that, The acquisition of the target multi-level description includes: Determine the target preset description refinement level corresponding to the generated remote sensing image; The target multi-level description is determined based on the target preset description refinement level and the target multi-level description corresponding to the remote sensing image generation result to be generated.
8. A remote sensing image generation device based on multi-level description, characterized in that, The device includes: The description acquisition module is used to acquire preprocessed remote sensing image samples and acquire multi-level descriptions corresponding to the remote sensing image samples. The multi-level descriptions include at least one of scene concept description, spatial relationship description, and ground feature description. The text connection and encoding module is used to connect the multi-level descriptions into a unified condition text according to a preset connection order, and encode the unified condition text to obtain multi-level description condition features; The training module is used to compress the remote sensing image samples into sample latent variables in the latent space, add random noise to the sample latent variables to obtain noise latent variables, input the noise latent variables, the denoising time step and the multi-level descriptive condition features into the initial noise prediction network to obtain the predicted noise, update the initial noise prediction network according to the loss value between the predicted noise and the random noise, until a preset stopping condition is reached to obtain the trained noise prediction network. The image generation module is used to acquire a multi-level description of the target, which includes at least one of the following: a target scene concept description, a target spatial relationship description, and a target feature description. The multi-level descriptions are connected into a unified target conditional text according to a preset connection order. The unified target conditional text is encoded to obtain multi-level target description conditional features. Random latent variables are acquired. The random latent variables, the denoising time step, and the multi-level target description conditional features are input into a noise prediction network to obtain target prediction noise. Based on the target prediction noise and the denoising time step, the random latent variables are denoised to output target latent spatial variables. The target latent spatial variables are decoded to generate a remote sensing image generation result.
9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the remote sensing image generation method based on multi-level description as described in any one of claims 1 to 7.
10. A computer device comprising a memory, a processor, and a computer program stored on a storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the remote sensing image generation method based on multi-level description as described in any one of claims 1 to 7.