Small target image generation method and device based on diffusion model, equipment and medium

By using a diffusion model to generate small target images, and leveraging feature extraction and cross-modal attention mechanisms, we can generate remote sensing images with enhanced details. This solves the problem of low clarity for small targets in remote sensing images and achieves efficient recognition and generation.

CN122067072BActive Publication Date: 2026-08-25XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610534391.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-22
Publication Date
2026-08-25
Estimated Expiration
2046-04-22

AI Technical Summary

Technical Problem

Currently, small targets in remote sensing images have low clarity and recognizability, making it difficult to generate images with enhanced details.

Method used

A small target image generation method using a diffusion model extracts features through a first image encoder, a second image encoder, and a text encoder, and combines a cross-modal attention mechanism and a student model to perform feature fusion, generating images with enhanced details.

Benefits of technology

It improves the recognition capability and generation efficiency of small target images, reduces manual operation, and increases the generation speed of image detail enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122067072B_ABST
    Figure CN122067072B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of remote sensing analysis and the technical field of artificial intelligence, and discloses a small target image generation method and device based on a diffusion model, equipment and a medium, the method comprising the following steps: performing feature extraction on a current remote sensing image to generate local features of the current remote sensing image, global features of the current remote sensing image and text features of the current remote sensing image; fusing the local features of the current remote sensing image, the global features of the current remote sensing image, the text features of the current remote sensing image, local features of a small target image corresponding to the current remote sensing image, global features of the small target image corresponding to the current remote sensing image and text features of the small target image corresponding to the current remote sensing image to generate fusion features of the current remote sensing image; and inputting the fusion features into a trained student model to generate the current remote sensing image after a small target is enhanced through the trained student model. The application can improve the generation efficiency of the current remote sensing image after a small target is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of remote sensing analysis technology and artificial intelligence technology, and in particular to methods, apparatus, devices and media for generating small target images based on diffusion models. Background Technology

[0002] With the widespread application of remote sensing technology, current remote sensing imagery serves as an important data source for Earth observation, providing rich information about the Earth's surface and covering many fields such as land use, environmental monitoring, and disaster assessment.

[0003] However, due to limitations imposed by sensor performance and atmospheric conditions during the acquisition of current remote sensing images, and the fact that small targets themselves occupy few pixels and have weak features, their clarity and recognizability in current remote sensing images are low. Therefore, how to generate detailed-enhanced images of small targets based on current remote sensing images is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] This application provides a method, apparatus, device, and medium for generating small target images based on a diffusion model, in order to solve the aforementioned technical problem of how to generate small target images with enhanced details based on current remote sensing images.

[0005] In a first aspect, embodiments of this application provide a method for generating small target images based on a diffusion model, applied to an electronic device deployed with a student model, wherein the student model is a diffusion model, and the method for generating small target images includes: Acquire current remote sensing images from the remote sensing system; The first image encoder, the second image encoder, and the text encoder are used to extract features from the current remote sensing image, respectively, and generate local features, global features, and text features of the current remote sensing image. The first image encoder, the second image encoder, and the text encoder are used to extract features from the small target image corresponding to the current remote sensing image, respectively, and generate local features, global features, and text features of the small target image corresponding to the current remote sensing image. Using a cross-modal attention mechanism, the local features, global features, and text features of the current remote sensing image are fused together to generate the fused features of the current remote sensing image. The fusion features of the current remote sensing image are input into the trained student model. The trained student model reconstructs the fusion features of the current remote sensing image to generate reconstructed information. The image details of small targets in the current remote sensing image are enhanced using the reconstructed information to generate the current remote sensing image with enhanced small targets.

[0006] In one possible implementation of the first aspect, the small target image generation method includes, before acquiring the current remote sensing image from the remote sensing system: Obtain the training set, which consists of multiple samples. Each sample includes a preset remote sensing image and a small target image corresponding to the preset remote sensing image. The first image encoder, the second image encoder, and the text encoder are used to extract features from the preset remote sensing image, generating local features, global features, and text features of the preset remote sensing image, respectively. The first image encoder, the second image encoder, and the text encoder are used to extract features from the small target image corresponding to the preset remote sensing image, generating local features, global features, and text features of the small target image corresponding to the preset remote sensing image, respectively. The cross-modal attention mechanism is used to fuse the local features, global features, text features, local features of the small target image corresponding to the preset remote sensing image, global features of the small target image corresponding to the preset remote sensing image, and text features of the small target image corresponding to the preset remote sensing image to generate the fused features of the preset remote sensing image. The fusion features of the preset remote sensing images are input into the student model, and the predicted noise of the student model is generated through the diffusion module of the student model. The fusion features of the preset remote sensing images are input into the teacher model, and the predicted noise of the teacher model is generated through the diffusion module of the teacher model. Based on the predicted noise of the student model, the real noise of the preset remote sensing image, and the first loss model, the diffusion loss of the student model is generated. Based on the predicted noise of the student model, the predicted noise of the teacher model, and the second loss model, the output layer distillation loss of the student model is generated. Based on the output features of each layer of the student model and each layer of the teacher model, and the third loss model, the feature layer distillation loss of the student model is generated. The total loss of the student model is generated based on the diffusion loss of the student model, the output layer distillation loss of the student model, the feature layer distillation loss of the student model, and the total loss model. The model parameters of the student model are trained using backpropagation until the total loss is less than the preset loss value. Only then is the training of the student model stopped and the trained student model saved.

[0007] In one possible implementation of the first aspect, the small target image corresponding to the preset remote sensing image is the vehicle image in the preset remote sensing image, and the small target image corresponding to the current remote sensing image is the vehicle image in the current remote sensing image. Alternatively, the small target image corresponding to the preset remote sensing image is the pedestrian image in the preset remote sensing image, and the small target image corresponding to the current remote sensing image is the pedestrian image in the current remote sensing image.

[0008] In one possible implementation of the first aspect, the first loss model is defined as follows: ; ; It is the diffusion loss of the student model; To introduce noise latent variables into remote sensing images, For noise scheduling functions, Represents a random variable; Indicates a normal distribution. This indicates that the random variable follows a normal distribution. t is the time step number. As expected, For the velocity field, These are the prediction parameters for the student model, used to predict the direction of noise. These are the initial latent variables for the preset remote sensing images; This is a conditional vector for preset remote sensing images; Through A function that predicts noise at the t-th time step under feature guidance; It is the square of the Euclidean distance.

[0009] In one possible implementation of the first aspect, the second loss model is defined as follows: ; in, For the output layer distillation loss of the student model, For the velocity field of the student model, For the prediction parameters of the student model, For the velocity field of the teacher model, Predict parameters for the teacher model. To introduce noise latent variables into remote sensing images, The image is a pre-defined image of a small target with a noise latent variable corresponding to a remote sensing image, where t is the time step number. As expected, The square of the Euclidean distance. , These are the conditional vectors of the small target image corresponding to the preset remote sensing image and the conditional vectors of the preset remote sensing image, respectively.

[0010] In one possible implementation of the first aspect, the third loss model is defined as follows: ; in, For the feature layer distillation loss of the student model, The first teacher model Features of the layer; The student model represents the first Features of the layer A set of feature layers for small target images corresponding to preset remote sensing images; As a feature projector, it only requires two convolutional layers to map student features to the dimension of teacher features. The remote sensing image contains latent noise variables; Add noise latent variables to the small target images corresponding to the preset remote sensing images; , These are the conditional vectors of the small target image corresponding to the preset remote sensing image and the conditional vectors of the preset remote sensing image, respectively. The total loss model is defined as follows: ; ; ; ; This represents the total loss of the student model; This represents the task loss of the student model; , , These are the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient; These are weighting coefficients; express The value at the t-th time step; express The value at the t-th time step; express The value at the t-th time step; It is the diffusion loss of the student model; For the feature layer distillation loss of the student model; The output layer distillation loss for the student model; It is the output layer distillation loss of the trained student model; It is the feature layer distillation loss of the trained student model.

[0011] Secondly, embodiments of this application provide a small target image generation device based on a diffusion model, applied to an electronic device with a student model deployed thereon, the student model being a diffusion model, comprising: The save module is used to acquire the current remote sensing image from the remote sensing system; The first extraction module is used to extract features from the current remote sensing image using a first image encoder, a second image encoder, and a text encoder, respectively, and generate local features, global features, and text features of the current remote sensing image. The second extraction module is used to extract features from the small target image corresponding to the current remote sensing image using the first image encoder, the second image encoder, and the text encoder, respectively, and generate the local features, global features, and text features of the small target image corresponding to the current remote sensing image. The first generation module is used to use a cross-modal attention mechanism to fuse the local features of the current remote sensing image, the global features of the current remote sensing image, the text features of the current remote sensing image, the local features of the small target image corresponding to the current remote sensing image, the global features of the small target image corresponding to the current remote sensing image, and the text features of the small target image corresponding to the current remote sensing image to generate the fused features of the current remote sensing image. The second generation module is used to input the fusion features of the current remote sensing image into the trained student model, reconstruct the fusion features of the current remote sensing image through the trained student model, generate reconstruction information, and perform image detail enhancement processing on small targets in the current remote sensing image through the reconstruction information to generate the current remote sensing image with enhanced small targets.

[0012] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the small target image generation method described in the first aspect above.

[0013] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the small target image generation method described in the first aspect above.

[0014] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the small target image generation method described in the first aspect.

[0015] The beneficial effects of the embodiments of this application are as follows: Firstly, the trained student model uses the teacher model to perform knowledge distillation, thus enabling the trained student model to effectively learn the parameters of the teacher model and significantly improve its own recognition ability. Therefore, the trained student model can generate current remote sensing images with enhanced small targets. Secondly, the fusion features of the current remote sensing image are input into the trained student model. The trained student model reconstructs the fusion features of the current remote sensing image to generate reconstructed information. The image details of small targets in the current remote sensing image are enhanced using the reconstructed information to generate the current remote sensing image with enhanced small targets. No manual operation is required, thus reducing the generation time of the current remote sensing image with enhanced small targets and improving the generation efficiency. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is an application scenario diagram of the small target image generation method provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the small target image generation method provided in an embodiment of this application; Figure 3 A schematic block diagram of a small target image generation apparatus provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0019] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0020] It should be understood that in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. The terms "comprising," "including," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0021] Furthermore, the technical solutions of the various embodiments can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0022] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0023] The small target image generation method provided in this application embodiment can be applied to electronic devices with student models deployed. The student model is a diffusion model. Electronic devices include, but are not limited to, servers, mobile phones, tablets, wearable devices, vehicle-mounted devices, and laptops. This application embodiment does not impose any restrictions on the specific type of electronic device.

[0024] Please see Figure 1 , Figure 1 The application scenario diagram of the small target image generation method provided in the embodiments of this application is described in detail below: The electronic device stores the trained student model, connects to the remote sensing system, and uses the first image encoder, the second image encoder, and the text encoder in the remote sensing system to extract features from the current remote sensing image, generating local features, global features, and text features of the current remote sensing image, respectively.

[0025] In the embodiments of this application, acquiring the current remote sensing image from the remote sensing system can quickly and accurately capture the current remote sensing image, reduce the acquisition time of the current remote sensing image, and help improve the acquisition efficiency of the current remote sensing image.

[0026] Please see Figure 2 , Figure 2 This is a flowchart illustrating a small target image generation method provided in an embodiment of this application, which can be applied to electronic devices.

[0027] like Figure 2 As shown, the small target image generation method provided in this application includes the following steps, which are detailed below: S201, Acquire current remote sensing imagery from the remote sensing system; The methods for generating small target images before acquiring current remote sensing images from the remote sensing system include: Obtain the training set, which consists of multiple samples. Each sample includes a preset remote sensing image and a small target image corresponding to the preset remote sensing image. The first image encoder, the second image encoder, and the text encoder are used to extract features from the preset remote sensing image, generating local features, global features, and text features of the preset remote sensing image, respectively. The first image encoder, the second image encoder, and the text encoder are used to extract features from the small target image corresponding to the preset remote sensing image, generating local features, global features, and text features of the small target image corresponding to the preset remote sensing image, respectively. The cross-modal attention mechanism is used to fuse the local features, global features, text features, local features of the small target image corresponding to the preset remote sensing image, global features of the small target image corresponding to the preset remote sensing image, and text features of the small target image corresponding to the preset remote sensing image to generate the fused features of the preset remote sensing image. The fusion features of the preset remote sensing images are input into the student model, and the predicted noise of the student model is generated through the diffusion module of the student model. The fusion features of the preset remote sensing images are input into the teacher model, and the predicted noise of the teacher model is generated through the diffusion module of the teacher model. Based on the predicted noise of the student model, the real noise of the preset remote sensing image, and the first loss model, the diffusion loss of the student model is generated. Based on the predicted noise of the student model, the predicted noise of the teacher model, and the second loss model, the output layer distillation loss of the student model is generated. Based on the output features of each layer of the student model and each layer of the teacher model, and the third loss model, the feature layer distillation loss of the student model is generated. The total loss of the student model is generated based on the diffusion loss of the student model, the output layer distillation loss of the student model, the feature layer distillation loss of the student model, and the total loss model. The model parameters of the student model are trained using backpropagation until the total loss is less than the preset loss value. Only then is the training of the student model stopped and the trained student model saved.

[0028] The student model's parameters are trained using backpropagation until the total loss is less than a preset loss value. Training then stops, and the trained student model is saved, including: The model parameters of the student model are trained using backpropagation until the total loss is less than the preset loss value. Training of the student model parameters is then stopped. The generated images of the trained student model are obtained, and all small target regions in the generated images are located using an object detector. All small target regions in the generated images are cropped and aligned with the corresponding small target images of the preset remote sensing images. The structural similarity index of all target regions is calculated. When the structural similarity index is greater than the preset index, the trained student model is saved.

[0029] Training the teacher model is a prerequisite for training the student model. Only when the teacher model has been fully trained and reaches its ideal performance state can it provide reliable guidance and knowledge transfer for training the student model.

[0030] For example, the fusion features of preset remote sensing images are input into the student model, and the student model's diffusion module generates the predicted noise of the student model. The fusion features of preset remote sensing images are input into the teacher model, and the teacher model's diffusion module generates the predicted noise of the teacher model. After training the teacher model based on the predicted noise of the teacher model and the teacher model's loss model, the diffusion loss of the student model is generated based on the predicted noise of the student model, the real noise of the preset remote sensing images, and the first loss model. The output layer distillation loss of the student model is generated based on the predicted noise of the student model, the predicted noise of the teacher model, and the second loss model. The feature layer distillation loss of the student model is generated based on the output features of each layer of the student model, the output features of each layer of the teacher model, and the third loss model.

[0031] The loss model for the teacher model is defined as follows: ; It is the diffusion loss of the teacher model; To introduce noise latent variables into remote sensing images, For noise scheduling functions, Represents a random variable; Indicates a normal distribution. This indicates that the random variable follows a normal distribution. t is the time step number. As expected, Velocity field These are the prediction parameters for the teacher model, used to predict the direction of noise. These are the initial latent variables for the preset remote sensing images; This is a conditional vector for preset remote sensing images; Through A function that predicts noise at the t-th time step under feature guidance; The square of the Euclidean distance is given. The first loss model is defined as follows: ; ; It is the diffusion loss of the student model; To introduce noise latent variables into remote sensing images, For noise scheduling functions, Represents a random variable; Indicates a normal distribution. This indicates that the random variable follows a normal distribution. t is the time step number. As expected, For the velocity field, These are the prediction parameters for the student model, used to predict the direction of noise. These are the initial latent variables for the preset remote sensing images; This is a conditional vector for preset remote sensing images; Through A function that predicts noise at the t-th time step under feature guidance; It is the square of the Euclidean distance.

[0032] Among them, the noisy latent variable of remote sensing image is a noisy latent variable generated by the preset remote sensing image at the t-th time step iteration.

[0033] The second loss model is defined as follows: ; in, For the output layer distillation loss of the student model, For the velocity field of the student model, For the prediction parameters of the student model, For the velocity field of the teacher model, Predict parameters for the teacher model. To introduce noise latent variables into remote sensing images, The image is a pre-defined image of a small target with a noise latent variable corresponding to a remote sensing image, where t is the time step number. As expected, The square of the Euclidean distance. , These are the conditional vectors of the small target image corresponding to the preset remote sensing image and the conditional vectors of the preset remote sensing image, respectively.

[0034] The third loss model is defined as follows: ; ; in, For the feature layer distillation loss of the student model, The first teacher model Features of the layer; The student model represents the first Features of the layer A set of feature layers for small target images corresponding to preset remote sensing images; As a feature projector, it only requires two convolutional layers to map student features to the dimension of teacher features. The remote sensing image contains latent noise variables; Add noise latent variables to the small target images corresponding to the preset remote sensing images; , These are the conditional vectors of the small target image corresponding to the preset remote sensing image and the conditional vectors of the preset remote sensing image, respectively. The total loss model is defined as follows: ; ; ; ; This represents the total loss of the student model; This represents the task loss of the student model; , , These are the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient; These are weighting coefficients; express The value at the t-th time step; express The value at the t-th time step; express The value at the t-th time step; It is the diffusion loss of the student model; For the feature layer distillation loss of the student model; The output layer distillation loss for the student model; It is the output layer distillation loss of the trained student model; It is the feature layer distillation loss of the trained student model.

[0035] S202, the first image encoder, the second image encoder, and the text encoder are used to extract features from the current remote sensing image, respectively, and generate local features, global features, and text features of the current remote sensing image. S203, the first image encoder, the second image encoder, and the text encoder are used to extract features from the small target image corresponding to the current remote sensing image, respectively, and generate local features, global features, and text features of the small target image corresponding to the current remote sensing image. The first image encoder and the second image encoder are different image encoders.

[0036] For ease of explanation, the following example is provided: For example, the first image encoder is a LIP-G encoder, and the second image encoder is a CLIP-L encoder; The LIP-G encoder is built on logarithmic image processing theory and focuses on the extraction and optimization of low-level visual features of images. It can complete basic visual processing such as image enhancement and edge detection. The CLIP-L encoder is a long text optimized version of the CLIP series image encoder. Its core function is to achieve cross-modal alignment of visual and text features. It is good at parsing complex visual information and adapting to tasks such as cross-modal matching and image classification.

[0037] Optionally, the text encoder adopts the Gemma-2-2b model. The Gemma-2-2b model can complete core tasks such as text understanding, feature vectorization, and semantic representation, and is a lightweight model for text encoding tasks.

[0038] S204 uses a cross-modal attention mechanism to fuse the local features, global features, text features, local features of the small target image corresponding to the current remote sensing image, global features of the small target image corresponding to the current remote sensing image, and text features of the small target image corresponding to the current remote sensing image to generate the fused features of the current remote sensing image. Among them, cross-modal attention mechanism is a core technology for integrating multimodal data in deep learning. Its core logic is to achieve the focus and fusion of key information between modalities through dynamic weight allocation. It can adaptively focus on important information in different modalities according to task requirements, while suppressing redundant parts.

[0039] S205: Input the fusion features of the current remote sensing image into the trained student model. The trained student model reconstructs the fusion features of the current remote sensing image to generate reconstructed information. The reconstructed information is used to enhance the image details of small targets in the current remote sensing image to generate the current remote sensing image with enhanced small targets.

[0040] In this embodiment, image detail enhancement processing is performed on small targets in the current remote sensing image by reconstructing information to generate a small target image with enhanced details. On the one hand, the small target image with enhanced details can clearly restore minute features such as texture, edges, and contours, weakening the problems of detail blurring and noise masking caused by shooting equipment, environmental interference, and transmission loss, making the image more layered and the information expression more complete, and improving the visual recognition efficiency of the human eye in the small target image in the current remote sensing image. On the other hand, the small target image with enhanced details can provide richer and more accurate feature dimensions for the recognition algorithm, effectively reducing the recognition error and matching deviation caused by feature loss and blurring, and improving the accuracy and stability of the recognition algorithm.

[0041] Among them, the current remote sensing image after enhancing small targets is the remote sensing image after enhancing the image details of small targets, including texture, edges, and contours.

[0042] Among them, the small target image corresponding to the preset remote sensing image is the vehicle image in the preset remote sensing image, and the small target image corresponding to the current remote sensing image is the vehicle image in the current remote sensing image; Alternatively, the small target image corresponding to the preset remote sensing image is the pedestrian image in the preset remote sensing image, and the small target image corresponding to the current remote sensing image is the pedestrian image in the current remote sensing image.

[0043] For ease of explanation, the following example is provided: For example, the small target image corresponding to the current remote sensing image is the vehicle image in the current remote sensing image. The fusion features of the current remote sensing image are input into the trained student model. Through the trained student model, the vehicle image with enhanced details is generated. After enhancing the vehicles in the current remote sensing image, the distinction between vehicle outlines, body boundaries and background will be significantly improved. Vehicles that were originally blurred and blended into the road surface and green belts can be clearly seen. When controlling traffic on urban main roads, individual vehicles in a single lane can be clearly distinguished, and adjacent vehicles will no longer be mixed up due to image blur. The traffic density of each lane can be accurately judged, providing clear image data for traffic police departments to adjust traffic light timing in real time and plan temporary diversion routes, avoiding control delays caused by inaccurate traffic flow judgment.

[0044] For example, the small target image corresponding to the current remote sensing image is the pedestrian image in the current remote sensing image. The fusion features of the current remote sensing image are input into the trained student model, and the trained student model generates pedestrian images with enhanced details. In urban public space planning, it can clearly identify the pedestrian distribution in public areas such as urban parks, riverside walkways, and urban squares, accurately see the pedestrian gathering density and traffic flow in each area, and even distinguish the pedestrian activity patterns at different times. This provides planners with real and accurate pedestrian data support for optimizing the layout of public space walkways, adding rest facilities, and planning pedestrian flow guidance channels, making public space planning more in line with the actual needs of citizens.

[0045] The beneficial effects of the embodiments of this application are as follows: Firstly, the trained student model uses the teacher model to perform knowledge distillation, thus enabling the trained student model to effectively learn the parameters of the teacher model and significantly improve its own recognition ability. Therefore, the trained student model can generate current remote sensing images with enhanced small targets. Secondly, the fusion features of the current remote sensing image are input into the trained student model. The trained student model reconstructs the fusion features of the current remote sensing image to generate reconstructed information. The image details of small targets in the current remote sensing image are enhanced using the reconstructed information to generate the current remote sensing image with enhanced small targets. No manual operation is required, thus reducing the generation time of the current remote sensing image with enhanced small targets and improving the generation efficiency.

[0046] For the small target image generation method described in the above embodiments, please refer to [link / reference]. Figure 3 , Figure 3 This is a schematic block diagram of a small target image generation apparatus provided in an embodiment of this application. Figure 3 The small target image generation device 400 shown can be applied to, for example... Figure 1 The application scenario diagram shows electronic devices. The following section uses electronic devices as an example to illustrate this. Figure 3 The small target image generation device 400 shown will be described in detail. The small target image generation device 400 may include a storage module 401, a first extraction module 402, a second extraction module 403, a first generation module 404, and a second generation module 405.

[0047] The storage module 401 is used to acquire the current remote sensing image from the remote sensing system; The first extraction module 402 is used to extract features from the current remote sensing image using a first image encoder, a second image encoder, and a text encoder, respectively, and generate local features, global features, and text features of the current remote sensing image. The second extraction module 403 is used to extract features from the small target image corresponding to the current remote sensing image using the first image encoder, the second image encoder, and the text encoder, respectively, and generate the local features, global features, and text features of the small target image corresponding to the current remote sensing image. The first generation module 404 is used to use a cross-modal attention mechanism to fuse the local features of the current remote sensing image, the global features of the current remote sensing image, the text features of the current remote sensing image, the local features of the small target image corresponding to the current remote sensing image, the global features of the small target image corresponding to the current remote sensing image, and the text features of the small target image corresponding to the current remote sensing image to generate the fused features of the current remote sensing image. The second generation module 405 is used to input the fusion features of the current remote sensing image into the trained student model, reconstruct the fusion features of the current remote sensing image through the trained student model, generate reconstruction information, and perform image detail enhancement processing on small targets in the current remote sensing image through the reconstruction information to generate the current remote sensing image with enhanced small targets.

[0048] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0049] The beneficial effects of the embodiments of this application are as follows: Firstly, the trained student model uses the teacher model to perform knowledge distillation, thus enabling the trained student model to effectively learn the parameters of the teacher model and significantly improve its own recognition ability. Therefore, the trained student model can generate current remote sensing images with enhanced small targets. Secondly, the fusion features of the current remote sensing image are input into the trained student model. The trained student model reconstructs the fusion features of the current remote sensing image to generate reconstructed information. The image details of small targets in the current remote sensing image are enhanced using the reconstructed information to generate the current remote sensing image with enhanced small targets. No manual operation is required, thus reducing the generation time of the current remote sensing image with enhanced small targets and improving the generation efficiency.

[0050] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0051] like Figure 4 As shown, Figure 4 The electronic device 2 includes: at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20, wherein the processor 20 executes the computer program 22 to implement the steps in any of the above method embodiments.

[0052] The electronic device 2 may include, but is not limited to, a processor 20 and a memory 21. Those skilled in the art will understand that... Figure 4 This is merely an example of electronic device 2 and does not constitute a limitation on electronic device 2. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0053] The processor 20 is used to run a computer program 22 stored in the memory 21, and performs the following steps when executing the computer program 22: Acquire current remote sensing images from the remote sensing system; The first image encoder, the second image encoder, and the text encoder are used to extract features from the current remote sensing image, respectively, and generate local features, global features, and text features of the current remote sensing image. The first image encoder, the second image encoder, and the text encoder are used to extract features from the small target image corresponding to the current remote sensing image, respectively, and generate local features, global features, and text features of the small target image corresponding to the current remote sensing image. Using a cross-modal attention mechanism, the local features, global features, and text features of the current remote sensing image are fused together to generate the fused features of the current remote sensing image. The fusion features of the current remote sensing image are input into the trained student model. The trained student model reconstructs the fusion features of the current remote sensing image to generate reconstructed information. The image details of small targets in the current remote sensing image are enhanced using the reconstructed information to generate the current remote sensing image with enhanced small targets.

[0054] The processor 20 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors, field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0055] In some embodiments, the memory 21 may be an internal storage unit of the electronic device 2, such as a hard disk or memory of the electronic device 2. In other embodiments, the memory 21 may be an external storage device of the electronic device 2, such as a plug-in hard disk, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard equipped on the electronic device 2.

[0056] Furthermore, the memory 21 may include both internal storage units and external storage devices of the electronic device 2. The memory 21 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 21 can also be used to temporarily store data that has been output or will be output.

[0057] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0058] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0059] The computer-readable storage medium may also be an external storage device of the small target image generating device or electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, non-transitory computer-readable storage medium, etc., equipped on the small target image generating device or electronic device.

[0060] Since the computer program stored in the computer-readable storage medium can execute any of the small target image generation methods based on the diffusion model provided in the embodiments of this application, the computer-readable storage medium can achieve the beneficial effects that any of the small target image generation methods based on the diffusion model provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0061] This application provides a computer program product that, when run on an electronic device, causes the electronic device to execute the above-described small target image generation method.

[0062] When a computer program is loaded into an electronic device, it can perform the following steps: Acquire current remote sensing images from the remote sensing system; The first image encoder, the second image encoder, and the text encoder are used to extract features from the current remote sensing image, respectively, and generate local features, global features, and text features of the current remote sensing image. The first image encoder, the second image encoder, and the text encoder are used to extract features from the small target image corresponding to the current remote sensing image, respectively, and generate local features, global features, and text features of the small target image corresponding to the current remote sensing image. Using a cross-modal attention mechanism, the local features, global features, and text features of the current remote sensing image are fused together to generate the fused features of the current remote sensing image. The fusion features of the current remote sensing image are input into the trained student model. The trained student model reconstructs the fusion features of the current remote sensing image to generate reconstructed information. The image details of small targets in the current remote sensing image are enhanced using the reconstructed information to generate the current remote sensing image with enhanced small targets.

[0063] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0064] Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium includes: an entity or device for carrying computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium.

[0065] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0066] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for generating small target images based on a diffusion model, characterized in that, The small target image generation method, applicable to electronic devices deployed with a student model (a diffusion model), includes: Obtain the training set, which consists of multiple samples. Each sample includes a preset remote sensing image and a small target image corresponding to the preset remote sensing image. The first image encoder, the second image encoder, and the text encoder are used to extract features from the preset remote sensing image, generating local features, global features, and text features of the preset remote sensing image, respectively. The first image encoder, the second image encoder, and the text encoder are used to extract features from the small target image corresponding to the preset remote sensing image, generating local features, global features, and text features of the small target image corresponding to the preset remote sensing image, respectively. The cross-modal attention mechanism is used to fuse the local features, global features, text features, local features of the small target image corresponding to the preset remote sensing image, global features of the small target image corresponding to the preset remote sensing image, and text features of the small target image corresponding to the preset remote sensing image to generate the fused features of the preset remote sensing image. The fusion features of the preset remote sensing images are input into the student model, and the predicted noise of the student model is generated through the diffusion module of the student model. The fusion features of the preset remote sensing images are input into the teacher model, and the predicted noise of the teacher model is generated through the diffusion module of the teacher model. Based on the predicted noise of the student model, the real noise of the preset remote sensing image, and the first loss model, the diffusion loss of the student model is generated. Based on the predicted noise of the student model, the predicted noise of the teacher model, and the second loss model, the output layer distillation loss of the student model is generated. Based on the output features of each layer of the student model and each layer of the teacher model, and the third loss model, the feature layer distillation loss of the student model is generated. The total loss of the student model is generated based on the diffusion loss of the student model, the output layer distillation loss of the student model, the feature layer distillation loss of the student model, and the total loss model. The model parameters of the student model are trained using backpropagation until the total loss is less than the preset loss value. Then the training of the student model parameters is stopped and the trained student model is saved. Acquire current remote sensing images from the remote sensing system; The first image encoder, the second image encoder, and the text encoder are used to extract features from the current remote sensing image, respectively, and generate local features, global features, and text features of the current remote sensing image. The first image encoder, the second image encoder, and the text encoder are used to extract features from the small target image corresponding to the current remote sensing image, respectively, and generate local features, global features, and text features of the small target image corresponding to the current remote sensing image. Using a cross-modal attention mechanism, the local features, global features, and text features of the current remote sensing image are fused together to generate the fused features of the current remote sensing image. The fusion features of the current remote sensing image are input into the trained student model. The trained student model reconstructs the fusion features of the current remote sensing image to generate reconstructed information. The image details of small targets in the current remote sensing image are enhanced using the reconstructed information to generate the current remote sensing image with enhanced small targets.

2. The method for generating small target images according to claim 1, characterized in that, The small target image corresponding to the preset remote sensing image is the vehicle image in the preset remote sensing image, and the small target image corresponding to the current remote sensing image is the vehicle image in the current remote sensing image; Alternatively, the small target image corresponding to the preset remote sensing image is the pedestrian image in the preset remote sensing image, and the small target image corresponding to the current remote sensing image is the pedestrian image in the current remote sensing image.

3. The method for generating small target images according to claim 1, characterized in that, The first loss model is defined as follows: ; ; It is the diffusion loss of the student model; To introduce noise latent variables into remote sensing images, For noise scheduling functions, Represents a random variable; Indicates a normal distribution. This indicates that the random variable follows a normal distribution. t is the time step number. As expected, For the velocity field, These are the prediction parameters for the student model, used to predict the direction of noise. These are the initial latent variables for the preset remote sensing images; This is a conditional vector for preset remote sensing images; Through A function that predicts noise at the t-th time step under feature guidance; It is the square of the Euclidean distance.

4. The method for generating small target images according to claim 2, characterized in that, The second loss model is defined as follows: ; in, For the output layer distillation loss of the student model, For the velocity field of the student model, These are the prediction parameters for the student model. For the velocity field of the teacher model, Predict parameters for the teacher model. To introduce noise latent variables into remote sensing images, The image is a pre-defined image of a small target with a noise latent variable corresponding to a remote sensing image, where t is the time step number. As expected, The square of the Euclidean distance. , These are the conditional vectors of the small target image corresponding to the preset remote sensing image and the conditional vectors of the preset remote sensing image, respectively.

5. The method for generating small target images according to claim 1, characterized in that... ; The third loss model is defined as follows: ; in, For the feature layer distillation loss of the student model, The first teacher model Characteristics of the layer; The student model represents the first Features of the layer A set of feature layers for small target images corresponding to preset remote sensing images; As a feature projector, it only requires two convolutional layers to map student features to the dimension of teacher features. The remote sensing image contains latent noise variables; Add noise latent variables to the small target images corresponding to the preset remote sensing images; , These are the conditional vectors of the small target image corresponding to the preset remote sensing image and the conditional vectors of the preset remote sensing image, respectively. The total loss model is defined as follows: ; ; ; ; This represents the total loss of the student model; This represents the task loss of the student model; , , These are the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient; These are weighting coefficients; express The value at the t-th time step; express The value at the t-th time step; express The value at the t-th time step; It is the diffusion loss of the student model; For the feature layer distillation loss of the student model; The output layer distillation loss for the student model; It is the output layer distillation loss of the trained student model; It is the feature layer distillation loss of the trained student model.

6. A small target image generation device based on a diffusion model, characterized in that, Applied to electronic devices that deploy student models, where the student model is a diffusion model, including: The storage module is used to acquire the training set, which consists of multiple samples. Each sample includes a preset remote sensing image and a small target image corresponding to the preset remote sensing image. A first image encoder, a second image encoder, and a text encoder are used to extract features from the preset remote sensing image, generating local features, global features, and text features of the preset remote sensing image, respectively. Similarly, the same first image encoder, second image encoder, and text encoder are used to extract features from the small target image corresponding to the preset remote sensing image, generating local features, global features, and text features of the small target image corresponding to the preset remote sensing image, respectively. A cross-modal attention mechanism is used to fuse these features, combining them to generate the preset remote sensing image. The image fusion features are input into the student model, and the student model's diffusion module generates predicted noise. The fusion features of the preset remote sensing images are input into the teacher model, and the teacher model's diffusion module generates predicted noise. Based on the student model's predicted noise, the preset remote sensing images' actual noise, and the first loss model, the student model's diffusion loss is generated. Based on the student model's predicted noise, the teacher model's predicted noise, and the second loss model, the student model's output layer distillation loss is generated. Based on the output features of each layer of the student model and the teacher model, and the third loss model, the student model's feature layer distillation loss is generated. Based on the student model's diffusion loss, output layer distillation loss, feature layer distillation loss, and total loss model, the student model's total loss is generated. The student model's parameters are trained using backpropagation until the total loss is less than the preset loss value. Training of the student model's parameters stops, the trained student model is saved, and the current remote sensing image is retrieved from the remote sensing system. The first extraction module is used to extract features from the current remote sensing image using a first image encoder, a second image encoder, and a text encoder, respectively, and generate local features, global features, and text features of the current remote sensing image. The second extraction module is used to extract features from the small target image corresponding to the current remote sensing image using the first image encoder, the second image encoder, and the text encoder, respectively, and generate the local features, global features, and text features of the small target image corresponding to the current remote sensing image. The first generation module is used to use a cross-modal attention mechanism to fuse the local features of the current remote sensing image, the global features of the current remote sensing image, the text features of the current remote sensing image, the local features of the small target image corresponding to the current remote sensing image, the global features of the small target image corresponding to the current remote sensing image, and the text features of the small target image corresponding to the current remote sensing image to generate the fused features of the current remote sensing image. The second generation module is used to input the fusion features of the current remote sensing image into the trained student model, reconstruct the fusion features of the current remote sensing image through the trained student model, generate reconstruction information, and perform image detail enhancement processing on small targets in the current remote sensing image through the reconstruction information to generate the current remote sensing image with enhanced small targets.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the small target image generation method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the small target image generation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Remote sensing image super-resolution method and product based on diffusion model and multi-modal large language model

    CN119722462A

  • Small target detection method and device, equipment and storage medium

    CN119851149A