Extreme environment degraded image simulation generation method and device
By obtaining degraded images and text prompt information in extreme environments, and using the diffusion variational autoencoder and dynamic semantic relationship aggregation model to generate context-decoupled feature vectors, the problem of poor image quality in extreme environments is solved and higher quality degraded image generation is achieved.
Patent Information
- Application Number
- CN202510843055.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In extreme environmental degradation scenarios, the existing L2I diffusion model fails to effectively handle the visual coupling and occlusion problems between foreground instances and backgrounds, resulting in poor quality of generated degraded images.
A method for simulating and generating images degraded in extreme environments is adopted. By acquiring degraded images, text prompt information, and object geometric layout information, a diffusion variational autoencoder, a two-end prototype resampling model, and a dynamic semantic relationship aggregation model are used to generate context-decoupled feature vectors and perform weighted fusion to improve image quality.
It improves the image quality in extremely degraded scenes, solves the visual coupling and occlusion problems between foreground instances and background, and generates higher quality degraded images.
Smart Images

Figure CN120707682A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method and apparatus for simulating and generating images degraded in extreme environments. Background Art
[0002] Currently, images captured in degraded scenes in extreme environments (e.g., low light, remote sensing, underwater, and extreme weather conditions like fog and rain) are of low quality and limited in data volume. This results in poor auxiliary training data sources for downstream visual models. Given the scarcity of visual data resources, generating images in extremely degraded scenes has become a growing concern. A common approach for simulating degraded images in extreme environments is to use an L2I (Layout-to-Image) diffusion model to encode the geometric layout and label information corresponding to the degraded image into position-aware and category-aware tags. These tags are then input into a latent diffusion space. Multi-instance mask generation is then performed on the geometric layout information to generate degraded images in extreme environments.
[0003] However, it has been found in practice that when the above method is used to simulate and generate extreme environment degradation images, the following technical problems often occur: In existing complex extreme environment degradation scenarios, there is a high degree of visual coupling between foreground instances and backgrounds due to their similarity in appearance, and there is frequent spatial occlusion between foreground instances. The L2I diffusion model does not take into account the occlusion problems between foreground instances and between foreground instances and backgrounds, resulting in poor quality of the generated degraded images.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention
[0005] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure provide a method and apparatus for generating simulated images of extreme environmental degradation to solve one or more of the technical problems mentioned in the above background technology section.
[0007] In a first aspect, some embodiments of the present disclosure provide a method for simulating and generating degraded images in extreme environments, including: obtaining degraded images, text prompt information, and object geometric layout information in extreme environments; inputting the degraded images into a diffusion variational autoencoder to obtain a potential image feature vector; generating a foreground prior visual feature vector set and a background prior visual feature vector based on the text prompt information and the object geometric layout information; inputting the foreground prior visual feature vector set and the background prior visual feature vector into a double-end prototype resampling model to obtain a foreground visual perception mark feature vector set and a background visual perception mark feature vector; and generating a foreground visual perception mark feature vector based on the potential image feature vector. , the above-mentioned object geometric layout information, the above-mentioned foreground visual perception mark feature vector set, the above-mentioned text prompt information, and the above-mentioned background visual perception mark feature vector are used to generate a context decoupling feature vector set; the above-mentioned context decoupling feature vector set is input into the context topological relationship coordination model to obtain a context topological relationship correction feature vector set; the above-mentioned context topological relationship correction feature vector set is input into the dynamic semantic relationship aggregation model to obtain a context semantic feature vector; the above-mentioned context semantic feature vector and the above-mentioned context decoupling feature vector set are weightedly fused to obtain a degraded visual feature vector; the above-mentioned degraded visual feature vector is decoded to obtain a target degraded image.
[0008] In a second aspect, some embodiments of the present disclosure provide a device for simulating and generating degraded images in extreme environments, including: an acquisition unit configured to acquire degraded images, text prompt information, and object geometric layout information in extreme environments; a first input unit configured to input the degraded image into a diffusion variational autoencoder to obtain a potential image feature vector; a first generation unit configured to generate a foreground visual prior feature vector set and a background prior visual feature vector based on the text prompt information and the object geometric layout information; a second input unit configured to input the foreground prior visual feature vector set and the background prior visual feature vector into a double-end prototype resampling model to obtain a foreground visual perception label feature vector set and a background visual perception label feature vector; a second generation unit configured to generate a foreground visual prior feature vector set and a background visual perception label feature vector based on the potential The image feature vector, the above-mentioned object geometric layout information, the above-mentioned foreground visual perception mark feature vector set, the above-mentioned text prompt information, and the above-mentioned background visual perception mark feature vector are used to generate a context decoupling feature vector set; the third input unit is configured to input the above-mentioned context decoupling feature vector set into the context topological relationship coordination model to obtain a context topological relationship correction feature vector set; the fourth input unit is configured to input the above-mentioned context topological relationship correction feature vector set into the dynamic semantic relationship aggregation model to obtain a context semantic feature vector; the fusion unit is configured to perform weighted fusion processing on the above-mentioned context semantic feature vector and the above-mentioned context decoupling feature vector set to obtain a degraded visual feature vector; the decoding unit is configured to perform decoding processing on the above-mentioned degraded visual feature vector to obtain a target degraded image.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0011] The above-described various embodiments of the present disclosure have the following beneficial effects: the simulation generation method for degraded images in extreme environments according to some embodiments of the present disclosure can improve the quality of images in extremely degraded scenes. Specifically, the poor quality of the generated degraded images is caused by the fact that in existing complex extreme degraded scenes, there is a high degree of visual coupling between foreground instances and backgrounds due to their appearance similarity, as well as frequent spatial occlusion between foreground instances. The L2I diffusion model does not account for occlusion between foreground instances or between foreground instances and background, resulting in poor quality of the generated degraded images. Based on this, the simulation generation method for degraded images in some embodiments of the present disclosure can first obtain a degraded image, textual prompt information, and object geometric layout information in an extreme environment. Here, the degraded image provides realistic visual information of the degraded scene, the textual prompt information describes the global semantic features of the scene, and the object geometric layout information accurately locates the position and category of foreground instances through bounding boxes, providing spatial constraints and semantic guidance for the subsequent generation process. Secondly, the degraded image is input into a diffusion variational autoencoder to obtain a latent image feature vector. Here, a diffuse variational autoencoder compresses the degraded image into a low-dimensional latent feature vector to capture essential global visual information from the degraded image, providing foundational information for the subsequent generation of context-decoupled feature vectors. Furthermore, based on the textual cue information and the object geometric layout information, a set of foreground and background prior visual feature vectors are generated. This method uses the object geometric layout information to obtain visual prior information of the same category as the foreground instance, while the textual cue information is used to obtain background prior information. This supplements detailed information from extreme environments and resolves blurring issues in extremely degraded scenes. Next, the foreground and background prior visual feature vectors are input into a two-end prototype resampling model to generate a set of foreground and background visual perception marker feature vectors. The two-end prototype resampling model includes a prototype resampling model for the foreground instance and a background resampling model to extract rich visual structural information from the foreground instance and background, enhancing the perceived visual difference between the two. Then, a context-decoupled feature vector set is generated based on the latent image feature vector, the object geometric layout information, the foreground visual perception marker feature vector set, the text prompt information, and the background visual perception marker feature vector. Here, the context-decoupled feature vector set includes a foreground decoupling feature vector set and a background decoupling feature vector. The foreground instances and background are decoupled by interacting with the latent image feature vector through parallel decoupling of the foreground instances and background. Subsequently, the context-decoupled feature vector set is input into a context topology coordination model to obtain a context topology correction feature vector set.Here, the topological relationship coordination model can correct the mutual topological associations of the contexts by constructing a topological graph including at least one foreground instance and background, then extracting and strengthening occlusion connectivity relationships within the context to suppress irrelevant and redundant connections. The context topological relationship correction feature vector set is then input into a dynamic semantic relationship aggregation model to obtain a context semantic feature vector. The dynamic semantic relationship aggregation model aggregates common and differential semantic information of contexts corresponding to different categories with different weights to improve the comprehensiveness and visual detail of the context semantic information. The context semantic feature vector and the context decoupling feature vector set are then weightedly fused to obtain a degraded visual feature vector. By assigning different attention weights to the foreground instance and background, more visual detail can be extracted, increasing the visual detail included in the degraded visual feature vector, enabling subsequent decoding to generate a higher-quality degraded image. Finally, the degraded visual feature vector is decoded to obtain a target degraded image. Consequently, this method for simulating and generating degraded images in extreme environments can improve the quality of images in extremely degraded scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0013] Figure 1 is a flow chart of some embodiments of the extreme environment degradation image simulation generation method according to the present disclosure;
[0014] Figure 2 is a schematic diagram of the conversion of the topological occlusion relationship between foreground instances in some embodiments of the extreme environment degradation image simulation generation method disclosed herein;
[0015] Figure 3 is a comparative schematic diagram of generated environmental degradation images in some embodiments of the extreme environmental degradation image simulation generation method disclosed herein;
[0016] Figure 4 is a schematic structural diagram of some embodiments of the extreme environment degradation image simulation generation device according to the present disclosure;
[0017] Figure 5 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0024] Figure 1 A process 100 of some embodiments of the degraded image generation method according to the present disclosure is shown. The degraded image generation method includes the following steps:
[0025] Step 101: Obtain a degraded image, text prompt information, and object geometric layout information under an extreme environment.
[0026] In some embodiments, the execution subject (e.g., an electronic device) of the above-mentioned extreme environment degradation image simulation generation method can obtain degraded images, text prompt information, and object geometric layout information through a wired connection or a wireless connection. The above-mentioned degraded image can be an image in an extreme environment, where the detailed information and structural information of each object in the image are lost or destroyed. For example, the extremely degraded scene can be, but is not limited to, at least one of the following: underwater environment, foggy and rainy weather, remote sensing image, artifacts, low light, and blur. The above-mentioned text prompt information can be a natural language statement describing the overall visual appearance of the degraded scene. The above-mentioned object geometric layout information can be the position information and semantic information of each object in the above-mentioned degraded image. For example, the above-mentioned degraded image can be an image of a turtle and a jellyfish in an underwater low-light environment. The above-mentioned position information can be the coordinates of the upper left corner and lower right corner of the bounding box of the turtle and jellyfish's respective locations, and the above-mentioned semantic information can be information about the semantic category of the turtle and jellyfish.
[0027] Step 102: Input the degraded image into a diffusion variational autoencoder to obtain a latent image feature vector.
[0028] In some embodiments, the execution entity may input the degraded image into a diffuse variational autoencoder to obtain a latent image feature vector. The latent image feature vector may represent information such as the texture, color, and geometric shape distribution of the degraded image. For example, if the degraded image is a turtle or jellyfish in a dimly lit underwater environment, the latent image feature vector may include information about the turtle's shell outline and the transparency of the jellyfish's tentacles. If the degraded image is a vehicle in foggy or smoggy weather, the latent image feature vector may include information about the vehicle's outline and the fog concentration gradient.
[0029] Step 103: Generate a foreground priori visual feature vector set and a background priori visual feature vector according to the text prompt information and the object geometric layout information.
[0030] In some embodiments, the execution entity may generate a foreground prior visual feature vector set and a background prior visual feature vector based on the text prompt information and the object geometric layout information. The foreground prior visual feature vector in the foreground prior visual feature vector set may represent the geometric visual information of existing objects in the degraded image in the foreground object candidate dictionary that belong to the same category as a foreground object in the degraded image. The foreground object candidate dictionary may be a database for storing visual feature information of various foreground objects. The background prior visual feature vector may represent the overall scene information of the environment and background information surrounding each foreground instance in the degraded image. For example, the foreground object may be a turtle, and the foreground prior visual feature vector may include but is not limited to at least one of the following: turtle edge information, texture information, lighting information, and geometric layout position information.
[0031] As an example, the above-mentioned execution entity can first retrieve at least one image of the same category from the RUOD dataset (Rethinking general Underwater Object Detection) based on the semantic information of the geometric layout information of the above-mentioned object. Secondly, the foreground information feature vector of the above-mentioned image is extracted through ResNet-50 (ResidualNetwork50) as a foreground prior visual feature vector set. Finally, the above-mentioned text prompt information is input into the SDXL (Stable Diffusion XL) model to obtain a background image. Afterwards, the background prior visual feature vector of the background image is extracted through DINOv2 (DIstillation with NO labels v2, unlabeled distillation model).
[0032] In some optional implementations of some embodiments, generating a foreground priori visual feature vector set and a background priori visual feature vector set based on the text prompt information and the object geometric layout information may include the following steps:
[0033] In the first step, a candidate degenerate feature instance group that matches the geometric layout information space and semantics of the object is retrieved from a preset candidate instance dictionary as a target candidate instance group, thereby obtaining a target candidate instance group set. The preset candidate instance dictionary may be a pre-constructed offline database that stores images of foreground instance objects by semantic category. The target candidate instances in the target candidate instance group set may be instance images in the preset candidate instance dictionary that match the semantic information. For example, if the semantic information is a turtle, the target candidate instance groups in the target candidate instance group set may include, but are not limited to, at least one of the following: a low-light turtle image, a turtle shell image, and an underwater turtle belly image.
[0034] In the second step, based on the above target candidate instance set, the following average pooling steps are performed:
[0035] Sub-step 1, performing image processing on each target candidate instance in the above target candidate instance set to generate a processed candidate instance and obtain a processed candidate instance set. The processed candidate instance in the above processed candidate instance set can be an instance image that matches the above semantic information and undergoes different image processing operations, and the processed results are spliced with the original image to generate an image. For example, the instance image that matches the above semantic information can be a tortoise shell image, and the above different image processing operations can include but are not limited to at least one of the following: image processing of extracting a binary image of the tortoise shell contour by the Canny operator, image processing of extracting a concave-convex texture map of the tortoise shell by Gabor filtering, image processing of extracting a dark area brightness distribution map by grayscale transformation, and image splicing of the tortoise shell image. The above splicing of the tortoise shell image can be a process of splicing the tortoise shell image, the binary image of the tortoise shell contour, the concave-convex texture map of the tortoise shell, and the brightness distribution map of the dark area.
[0036] Sub-step 2, inputting the above-mentioned processed candidate instance set into the image convolution feature extraction model to obtain a foreground instance visual feature vector set. The foreground instance visual feature vectors in the above-mentioned foreground instance visual feature vector set can represent the visual information of the foreground instance. The above-mentioned image convolution feature extraction model can be a model that takes the above-mentioned processed candidate instance set as input, selectively extracts significant visual features, and outputs a foreground instance visual feature vector set. For example, the above-mentioned image convolution feature extraction model can be but is not limited to at least one of the following: a deformable convolutional neural network, a dynamic area perception convolution model. For example, if the above-mentioned processed candidate instance is a turtle image, the foreground instance visual feature vectors in the above-mentioned foreground instance visual feature vector set can include but is not limited to at least one of the following: turtle shell texture feature information, turtle shell edge contour information.
[0037] Sub-step 3: performing average pooling processing on the foreground instance visual feature vector set to obtain a foreground prior visual feature vector set.
[0038] The third step is to retrieve a training background image from a preset training background image set that matches the text prompt corresponding to the degraded image, and use it as the target training background image. The preset training background image set can be a pre-trained set of background images for various extreme environments. In practice, the execution entity can utilize a CLIP (Contrastive Language-Image Pre-training, a multimodal pre-training model) model to perform matching screening from the preset training background image set based on the text prompt information to obtain the target training background image.
[0039] In the fourth step, the target training background image is determined as a target candidate instance group set, and the average pooling step is performed again, and the obtained foreground prior visual feature vector set is used as the background prior visual feature vector.
[0040] Step 104 : Input the foreground prior visual feature vector set and the background prior visual feature vector into a two-terminal prototype resampling model to obtain a foreground visual perception mark feature vector set and a background visual perception mark feature vector.
[0041] In some embodiments, the execution entity may input the foreground prior visual feature vector set and the background prior visual feature vector into a dual-end prototype resampling model to obtain a foreground visual perception labeling feature vector set and a background visual perception labeling feature vector. The dual-end prototype resampling model may perform feature projection, resampling processing, and dynamic adaptation operations on the input foreground prior visual feature vector set and background prior visual feature vector to output a model of the foreground visual perception labeling feature vector set and the background visual perception labeling feature vector. For example, the dual-end prototype resampling model may be a Transformer decoder model. The foreground visual perception labeling feature vector in the foreground visual perception labeling feature vector set may represent the visual information of the foreground instance in the degraded image, and the information obtained through interactive training of the dual-end prototype resampling model. The background visual perception labeling feature vector may represent the visual information of the background in the environmental image, and the information obtained through interactive training of the dual-end prototype resampling model. For example, the degraded image may be an image of a turtle and jellyfish in an underwater, dimly lit coral reef environment. In this scenario, the foreground visual perception marker feature vector may include, but is not limited to, at least one of the following: texture information of the turtle's back markings in dim light, information about correction of the turtle's color distortion caused by water scattering, and information about the underwater light refraction pattern of the jellyfish's tentacles. The background visual perception marker feature vector may also include, but is not limited to, at least one of the following: distribution information about the blurred outline of the coral reef community and information about the scattering path of the light source in the water.
[0042] In some optional implementations of some embodiments, the above-mentioned two-end prototype resampling model includes: a foreground prototype resampling model and a background prototype resampling model. Among them, the input of the above-mentioned foreground prototype resampling model can be the above-mentioned foreground prior visual feature vector set and the above-mentioned foreground learnable query label, and the output is a model of the above-mentioned foreground visual perception label feature vector set. The input of the above-mentioned background prototype resampling model is the above-mentioned background prior visual feature vector and the above-mentioned background learnable query label, and the output is a model of the above-mentioned contextual visual perception label feature vector. The above-mentioned foreground prototype resampling model can be a model including 4 network layers, each network layer including a cross-attention mechanism and a feedforward neural network stacked in series. The above-mentioned background prototype resampling model can be a model with the same network structure as the above-mentioned foreground prototype resampling model, but with different inputs and outputs and different model parameters.
[0043] Optionally, the step of inputting the foreground prior visual feature vector set and the background prior visual feature vector into a two-terminal prototype resampling model to obtain a foreground visual perception mark feature vector set and a background visual perception mark feature vector may include the following steps:
[0044] The first step is to obtain a foreground learnable query tag corresponding to the foreground prior visual feature vector set and a background learnable query tag corresponding to the background prior visual feature vector. The foreground learnable query tag may be a randomly initialized and trainable parameter matrix corresponding to the foreground prototype resampling model. The background learnable query tag information may be a randomly initialized and trainable parameter matrix corresponding to the background prototype resampling model.
[0045] In the second step, a linear transformation is performed on the foreground visual feature vector set to obtain a foreground prior key vector set and a foreground prior value vector set. The foreground prior key vector in the foreground prior key vector set can be characterized as a retrievable identification feature vector containing abstract semantic feature information of the foreground category and feature information of key visual attributes. The foreground prior value vector in the foreground prior value vector set can be characterized as containing richer and more specific visual feature information of the foreground category. For example, in an image of a puppy on the grass, the foreground category is puppy. The foreground prior key vector can include but is not limited to at least one of the following: semantic label information representing "dog", visual information representing "standing posture", and the foreground prior value vector can include but is not limited to at least one of the following: morphological local feature information of erect ears, and fur color texture information of black hair.
[0046] In the third step, the foreground learnable query tag, the foreground prior key vector set, and the foreground prior value vector set are input into the foreground prototype resampling model to obtain the foreground visual perception tag feature vector set.
[0047] The fourth step is to perform a linear transformation on the above-mentioned background visual feature vector to obtain a background prior key vector and a background prior value vector. Among them, the background prior key vector in the above-mentioned background prior key vector set can be characterized as a retrievable identification feature vector containing abstract semantic feature information of the background category and feature information of key visual attributes. The background prior value vector in the above-mentioned background prior value vector set can be characterized as containing visual feature information with richer and more specific background information. For example, in an image of a puppy on grass, the background is grass. The above-mentioned background prior key vector may include but is not limited to at least one of the following: scene semantic information representing "grassland", structural feature visual information representing "horizontal extension", and the above-mentioned background prior value vector may include but is not limited to at least one of the following: visual information representing short and dense green grass and scattered wildflowers, and spatial regularity information representing that the texture is clear near and blurred far away.
[0048] In the fifth step, the background learnable query tag, the background prior key vector, and the background prior value vector are input into the background prototype resampling model to obtain a background visual perception tag feature vector.
[0049] Step 105 : Generate a context-decoupled feature vector set based on the latent image feature vector, the object geometric layout information, the foreground visual perception mark feature vector set, the text prompt information, and the background visual perception mark feature vector.
[0050] In some embodiments, the execution entity may generate a context-decoupled feature vector set based on the latent image feature vector, the object geometric layout information, the foreground visual perception marker feature vector set, the text prompt information, and the background visual perception marker feature vector. The context-decoupled feature vectors in the context-decoupled feature vector set may represent the association between the foreground instance and the background.
[0051] In the process of adopting technical solutions to solve the above-mentioned technical problem one, the following technical problem two is often accompanied: only considering the decoupling between the foreground instance and the background, resulting in unclear extraction of the basic visual information of the degraded image when identifying the foreground instance and the background, which in turn leads to the problem of low quality of the generated degraded image. In response to the above-mentioned technical problem two, the conventional solution is generally to interact the foreground instance with the latent image features through a single-path decoupling attention to extract richer basic visual information of the degraded image. However, the above-mentioned conventional solution still has the following problem: the degradation information of the background area is not extracted independently, resulting in unclear background visual information of the degraded image, which in turn leads to low quality of the generated degraded image. The inventors took into account the shortcomings of the above-mentioned conventional solutions, and combined with the advantages / technical status of the feature interaction decoupling technology of the foreground feature vector, background feature vector and latent feature vector owned by the inventor's company, we decided to adopt the following solution:
[0052] In some optional implementations of some embodiments, the object geometric layout information includes: a foreground instance bounding box position information set and a foreground instance semantic category information set. The foreground instance bounding box position information in the foreground instance bounding box position information set may include coordinate information of the upper left and lower right corners of the foreground instance bounding box. The foreground instance semantic category information in the foreground instance semantic category information set may include category name information of the foreground instance.
[0053] Optionally, generating a context-decoupled feature vector set based on the latent image feature vector, the object geometric layout information, the foreground visual perception marker feature vector set, the text prompt information, and the background visual perception marker feature vector, and generating a target degraded image based on the context-decoupled feature vector set may include the following steps:
[0054] The first step is to perform text encoding on the foreground instance semantic category information set to obtain a foreground category semantic feature vector set. The foreground category semantic feature vectors in the foreground category semantic feature vector set can be represented as semantic vector representations of the category information to which the object name belongs. For example, the semantic category can be a car, and the scene category semantic feature vectors in the foreground category semantic feature vector set can be vectors representing the basic semantic information of a car having windows, tires, and a metal shell. In practice, the execution entity can use the CLIP model to perform text encoding on the foreground instance semantic category information set to obtain a foreground category semantic feature vector set.
[0055] The second step is to perform bounding box position encoding on the foreground instance bounding box position information set to obtain a bounding box position feature vector set. The bounding box position feature vectors in the bounding box position feature vector set can be represented as numerical encoding representations of object position coordinates. In practice, the execution entity can use Fourier embedding to perform position coordinate conversion on the foreground instance bounding box position information set to obtain a Fourier transformed bounding box position embedding vector, and then use an MLP (Multilayer Perceptron) model feature extraction operation to obtain the bounding box position feature vector set.
[0056] The third step is to perform feature concatenation on the foreground visual perception mark feature vector set, the foreground category semantic feature vector set, and the bounding box position feature vector set in the feature dimension to obtain a foreground concatenated feature vector set. The foreground concatenated feature vectors in the foreground concatenated feature vector set can be represented as a comprehensive vector representation combining object visual features, category semantics, and position information.
[0057] The fourth step is to perform a linear transformation on the latent image feature vector to obtain an image visual transformation feature vector. The image visual transformation feature vector can be a query vector generated by a linear projection layer of the latent image feature vector, which is used as a feature vector for retrieving key information in the attention mechanism.
[0058] In the fifth step, the image visual transformation feature vector and the foreground splicing feature vector set are input into a multi-head cross-attention mechanism layer to obtain a foreground visual attention feature vector set. The foreground visual attention feature vector in the foreground visual attention feature vector set can represent the visual semantic information of the foreground instance integrated with the global environment features.
[0059] The sixth step is to determine the bounding box mask vector set of the foreground instance bounding box position information set. The bounding box mask vectors in the bounding box mask vector set can be represented as position identification information of the foreground object spatial region expressed in binary form. For example, the foreground instance can be the car bounding box coordinates [120, 80, 300, 200], and the bounding box mask vectors in the bounding box mask vector set can be bounding box coordinates x∈[120, 300], y∈[80, 200] is 1, and other areas are 0. In practice, the execution entity can use the attention mask mechanism to determine the bounding box mask vector set of the foreground instance bounding box position information set.
[0060] The seventh step is to determine the element-by-element product of each foreground visual attention feature vector in the foreground visual attention feature vector set and the corresponding bounding box mask vector in the bounding box mask vector set to obtain the foreground instance layout interaction feature vector set.
[0061] Step 8: Text-encode the text prompt information to obtain a text prompt feature vector. The text prompt feature vector can represent semantic information describing the global scene. For example, the text prompt information can be "City streets in dense fog, visibility less than 100 meters," and a vector representing dense fog, thick fog, and low visibility is output. In practice, the execution entity can use the CLIP model to perform a text-to-numeric vector operation on the text prompt information to obtain the text prompt feature vector.
[0062] In the ninth step, the background visual perception mark feature vector and the text prompt feature vector are concatenated in the feature dimension to obtain a background concatenated feature vector. The background concatenated feature vector set can be represented as a fusion of degraded visual information and scene semantic information.
[0063] In the tenth step, the image visual transformation feature vector and the background splicing feature vector are input into a multi-head cross-attention mechanism layer to obtain a background visual attention feature vector. The background visual attention feature vector can be represented as the semantic information of the text prompt information and the image semantic information after the attention weighted fusion.
[0064] In the eleventh step, a background region mask vector is generated based on the bounding box mask vector set, wherein the background region mask vector can represent the region position information of the background region in the degraded image except for the bounding box of the foreground instance.
[0065] As an example, the execution subject may use a foreground instance separation decoupling function to generate a background region mask vector based on the bounding box mask vector set. The foreground instance separation decoupling function may be:
[0066]
[0067] Where N represents the total number of bounding boxes of the foreground instances. i represents the binary attention mask of the i-th bounding box in the above foreground instance.
[0068] In step 12, the element-wise product of the background visual attention feature vector and the background region mask vector is determined to obtain a background text interaction feature vector. The background text interaction feature vector can be characterized as background description information indicating that the background region conforms to the degraded environmental characteristics described in the text. For example, if the text prompt information is about hazy weather, the background description information may include, but is not limited to, at least one of the following: road fog concentration gradient information, grayish-white sky tonal characteristics, and area restriction information.
[0069] Step 13: Combine the foreground instance context layout interaction feature vector set and the background text interaction feature vector set to obtain a context-decoupled feature vector set. In practice, the execution entity may store the foreground instance context layout interaction feature vector set and the background text interaction feature vector set in a preset set to obtain the context-decoupled feature vector set. The preset set may be a pre-designed set for storing feature vectors.
[0070] In step 14, a target degraded image is generated based on the context decoupled feature vector set. As an example, the implementation of this step can refer to the implementation of steps 106-109, which will not be described in detail again.
[0071] Steps 1 through 14, and their related content, serve as an inventive feature of an embodiment of the present disclosure. They address the second technical issue mentioned in the background art: "Only considering the decoupling between foreground instances and background leads to unclear extraction of basic visual information from the degraded image when identifying foreground instances and background, resulting in low quality of the generated degraded image." Factors contributing to this unclear extraction of basic visual information from the degraded image when identifying foreground instances and background, and thus low quality of the generated environmentally degraded image, are often as follows: Only performing feature interaction on the foreground instances in a single path without independently processing basic background visual information, resulting in unclear background visual information in the degraded image and, consequently, low quality of the generated degraded image. Addressing these factors can extract more basic image visual information, improve the realism and detail restoration of the degraded image, and produce high-quality degraded images. To achieve this, the present disclosure first utilizes the CLIP model to encode foreground instance semantic category information into a foreground category semantic feature vector containing semantic information, providing guidance for object detail extraction in the foreground instance decoupling path. Foreground instance bounding box position information is transformed via Fourier embedding, and then generated via an MLP to lock the foreground instance bounding box position feature vector. The foreground visual perceptual marker feature vector, foreground category semantic feature vector, and bounding box position feature vector are integrated to form a foreground splicing feature vector set, which serves as the decoupled basic information and prevents the incorporation of environmental information. Next, information is extracted independently through the foreground instance path. The interaction between the image latent image feature vector after feature projection and the decoupled basic information set is used to extract foreground instance details. A binary bounding box mask is generated based on the foreground instance bounding box position information. The feature modification range is strictly limited through element-wise multiplication to obtain a foreground instance layout interaction feature vector set, ensuring that foreground instance details do not affect the background area. Next, information is extracted independently through the background path. CLIP is used to encode textual cue information to extract the environmental visual basic information and obtain a text cue feature vector. The background visual marker and text cue feature vector are then spliced together to construct a background splicing feature vector. Multi-head cross-attention is used to calculate the interaction between the image latent image feature vector after feature projection and the background splicing feature vector set to extract background visual basic details. A back-masking operation is performed to obtain a background-text interaction feature vector, ensuring that background effects are limited to the background area and avoid affecting the foreground instance area. Afterwards, the foreground instance layout interaction feature vector set and the background text interaction feature vector are spliced together to obtain the context-decoupled feature vector set. By combining the foreground instance path-independent extraction information and the background path-independent extraction information, interaction decoupling is performed with the latent feature vector from two aspects, which can improve the accuracy of interaction decoupling, remove redundant information of the context-decoupled feature vector set, and improve the quality of the context-decoupled feature vector set.Finally, the decoder is used to reconstruct the context-decoupled feature vector set to obtain a degraded image with richer visual information, greater realism, and higher detail restoration.
[0072] Step 106 : Input the context decoupling feature vector set into the context topology relationship coordination model to obtain the context topology relationship correction feature vector set.
[0073] In some embodiments, the execution subject inputs the context decoupling feature vector set into the context topology relationship coordination model to obtain a context topology relationship correction feature vector set. The context topology relationship coordination model can be a model that takes the context decoupling feature vector set as input, constructs a feature adjacency matrix and performs topology feature reprojection operations, and outputs a context topology relationship correction feature vector set. The context topology relationship coordination model includes: a relationship coordination model and a graph convolutional neural network model. The context topology relationship correction feature vectors in the context topology relationship correction feature vector set can represent the occlusion association information and hierarchical information between foreground instances and backgrounds, and between foreground instances. Figure 2 As shown, the information of the occlusion association between instances represented by the above-mentioned context topological relationship correction feature vector set is displayed.
[0074] In some optional implementations of some embodiments, inputting the context decoupling feature vector set into the context topology relationship coordination model to obtain the context topology relationship correction feature vector set may include the following steps:
[0075] The first step is to determine the context adjacency matrix of the context-decoupled feature vector set through the above-mentioned relationship coordination model. Among them, the above-mentioned relationship coordination model can be a model that takes the above-mentioned context-decoupled feature vector set as input, performs similarity calculation operations between the context-decoupled feature vectors in the above-mentioned context-decoupled feature vector set, and outputs the context adjacency matrix of the context-decoupled feature vector set. The above-mentioned context adjacency matrix can be characterized as the spatial occlusion and semantic association strength information between each foreground instance and between each foreground instance and the background. For example, the above-mentioned degraded image can be an image of a turtle and a jellyfish in an underwater dark light environment, where the jellyfish occludes the turtle. The elements in the above-mentioned instance adjacency matrix can be the similarity of 0.7 between the turtle and the jellyfish, indicating a strong spatial correlation between the turtle and the jellyfish. Among them, the formula corresponding to the above-mentioned relationship coordination model is:
[0076]
[0077] Where A represents the above context adjacency matrix. j represents the jth context-decoupled feature vector. represents the transposed feature vector of the i-th context-decoupled feature vector. |||| represents the modulus of the context-decoupled feature vector.
[0078] In the second step, the above-mentioned graph convolutional neural network model is used to perform graph symmetric normalization on the above-mentioned context adjacency relationship matrix to obtain a set of context topological relationship correction feature vectors.
[0079] Among them, the formula corresponding to the above graph convolutional neural network model is:
[0080]
[0081] in, Represents the context topology correction feature vector set. σ() represents the ReLU (Rectified Linear Unit, linear rectification function) activation function. represents the adjacency relationship between nodes in a graph with self-loops, The calculation formula can be:
[0082] Υ represents the identity matrix, express The matrix has dimensions |N+1|×|N+1|. express The degree matrix of Express Sum each row of . express The element in row i and column j.
[0083] Step 107 : Input the context topology relationship correction feature vector set into the dynamic semantic relationship aggregation model to obtain the context instance semantic feature vector.
[0084] In some embodiments, the execution subject inputs the context topology relationship correction feature vector set into the dynamic semantic relationship aggregation model to obtain a context semantic feature vector. The dynamic semantic relationship aggregation model may be a model that inputs a context topology relationship correction feature vector set, performs gated routing weighting, expert feature extraction and semantic aggregation operations, and outputs a context semantic feature vector. The dynamic semantic relationship aggregation model includes: a gating network, multiple expert networks and an aggregation operation network. The aggregation network may be a network that performs a weighted summation of the outputs of at least two expert networks activated by the output of the gating network. The context semantic feature vector may represent information about the common and different semantic attributes of objects of the same category as each foreground instance and background, which are aggregated according to different aggregation weights.
[0085] In some optional implementations of some embodiments, inputting the context topology relationship correction feature vector set into the dynamic semantic relationship aggregation model to obtain the context semantic feature vector may include the following steps:
[0086] The first step is to obtain a normally distributed noise vector. The normally distributed noise vector may represent noise information that conforms to a normal distribution. For example, the normally distributed noise vector may include, but is not limited to, at least one of the following: a feature vector representing simulated water flow information, or a feature vector representing light spot jitter information.
[0087] In the second step, the normally distributed noise vector and the context topology relationship correction feature vector set are input into the gating network to obtain the expert activation probability weight set for the multiple expert networks. The gating network may be a model that takes the normally distributed noise vector and the context topology relationship correction feature vector set as input, performs noise weighted score calculation and probability mapping operations, and outputs the expert activation probability weight set for the multiple expert networks. For example, the gating network may be a sparse gating model with noise. The expert activation probability weight of the expert network in the expert activation probability weight set of the multiple expert networks may be the probability weight value of the expert being selected.
[0088] Step 3: Select at least two expert activation probability weights from the expert activation probability weight set that meet a preset activation probability condition. The preset activation probability condition may be a condition where a preset number of experts with the highest activation probability weights are selected. The preset number may be a pre-set value. For example, the preset number may be 2.
[0089] In the fourth step, at least two expert networks corresponding to the activation probability weights of the at least two experts are determined from the multiple expert networks.
[0090] The fifth step is to input the above-mentioned context topological relationship correction feature vector set into the above-mentioned at least two expert networks to obtain a context semantic aggregation feature vector set. The above-mentioned expert network can be a model that takes the above-mentioned context topological relationship correction feature vector set as input, performs semantic feature enhancement processing operations, and outputs a context semantic aggregation feature vector set. For example, the above-mentioned expert network can be a multi-layer perceptron model with two linear layers and a nonlinear ReLU activation function. The context semantic aggregation feature vector in the above-mentioned context semantic aggregation feature vector set can be represented as the semantic information of foreground instances belonging to the same category or the information of common attributes and difference attributes of the background extracted for an expert network.
[0091] Step 6: Input the contextual semantic aggregated feature vector set and the at least two expert activation probability weights into the aggregation operation network to obtain a contextual semantic feature vector. The aggregation operation network may be a network that inputs the contextual semantic aggregated feature vector set and the at least two expert activation probability weights, performs a weighted feature fusion operation, and outputs the contextual semantic feature vector.
[0092] Step 108 : Perform weighted fusion processing on the context semantic feature vector and the context decoupling feature vector set to obtain a degraded visual feature vector.
[0093] In some embodiments, the execution entity performs weighted fusion processing on the context semantic feature vector and the context decoupling feature vector set to obtain a degraded visual feature vector. The degraded visual feature vector can represent the occlusion information and more detailed information of the foreground instance and background in the same category, and has clearer image quality and more detailed texture information than the latent image feature vector. The degraded visual feature vector can represent, but is not limited to, at least one of the following: object attribute information, environmental parameter information, occlusion logic between objects, and semantic relevance.
[0094] In some optional implementations of some embodiments, performing weighted fusion processing on the context semantic feature vector and the context decoupling feature vector set to obtain a degraded visual feature vector may include the following steps:
[0095] The first step is to normalize the contextual semantic feature vector to obtain the contextual semantic relationship aggregation weight. The contextual semantic relationship aggregation weight can be represented as the semantic importance weight information of the context. In practice, the execution entity can use a softmax (normalized exponential) function to normalize the contextual semantic feature vector to obtain the contextual semantic relationship aggregation weight.
[0096] The second step is to perform element-wise multiplication of the context semantic relationship aggregation weight and each context decoupling feature vector in the context decoupling feature vector set to obtain a context weight feature vector set. The context weight feature vector set can be represented as context feature information weighted by semantic importance.
[0097] The third step is to accumulate the above context weight feature vector set to obtain the degraded visual feature vector.
[0098] Step 109: Decode the degraded visual feature vector to obtain a target degraded image.
[0099] In some embodiments, the execution subject decodes the degraded visual feature vector to obtain a target degraded image. The target degraded image may be a high-fidelity degraded image that is well aligned with a given layout in terms of space and semantics. In practice, the execution subject may utilize the decoder in the diffusion variational autoencoder to decode the degraded visual feature vector to obtain a target degraded image. The diffusion variational autoencoder model, the two-end prototype resampling model, the contextual topological relationship coordination model, and the dynamic semantic relationship aggregation model may be the various component network models of the degradation network generation model. The degradation network generation model may be obtained by training through a loss function. The loss function may be:
[0100]
[0101] Among them, L VPEC represents the loss function of the degenerate network generative model. θ′ represents the trainable parameters. θ represents the frozen parameters of the pre-trained latent diffusion model, and ξ represents the mathematical expectation. z0 represents the latent image feature vector. ε~N(0,1) represents standard Gaussian noise (mean 0, variance 1). t represents the diffusion time step. Γ represents the text prompt information. ε represents the randomly generated Gaussian noise. Ω(·) represents the noise prediction network of the degenerate network generative model. z t represents the latent image feature vector at time step t with noise. B represents the set of position information in the geometric layout information of the object. Q represents the set of foreground visual perception marker feature vectors and background visual perception marker feature vectors. Represents the square of the L2 norm, that is, the mean square error.
[0102] like Figure 3 As shown, a schematic diagram of the comparison of degraded images generated by the above-mentioned degraded network generation model, Layout-Diffusion model, MIGC (Multi-instance Generation Controller) model, and CC-Diff (Enhancing Contextual Coherence in Remote Sensing Image Synthesis) model according to each degraded image in multiple degraded images and corresponding object geometric layout information is shown, wherein VPEC can be the degraded network generation model.
[0103] The above-described various embodiments of the present disclosure have the following beneficial effects: the simulation generation method for degraded images in extreme environments according to some embodiments of the present disclosure can improve the quality of images in extremely degraded scenes. Specifically, the poor quality of the generated degraded images is caused by the fact that in existing complex extreme degraded scenes, there is a high degree of visual coupling between foreground instances and backgrounds due to their appearance similarity, as well as frequent spatial occlusion between foreground instances. The L2I diffusion model does not account for occlusion between foreground instances or between foreground instances and background, resulting in poor quality of the generated degraded images. Based on this, the simulation generation method for degraded images in some embodiments of the present disclosure can first obtain a degraded image, textual prompt information, and object geometric layout information in an extreme environment. Here, the degraded image provides realistic visual information of the degraded scene, the textual prompt information describes the global semantic features of the scene, and the object geometric layout information accurately locates the position and category of foreground instances through bounding boxes, providing spatial constraints and semantic guidance for the subsequent generation process. Secondly, the degraded image is input into a diffusion variational autoencoder to obtain a latent image feature vector. Here, a diffuse variational autoencoder compresses the degraded image into a low-dimensional latent feature vector to capture essential global visual information from the degraded image, providing foundational information for the subsequent generation of context-decoupled feature vectors. Furthermore, based on the textual cue information and the object geometric layout information, a set of foreground and background prior visual feature vectors are generated. This method uses the object geometric layout information to obtain visual prior information of the same category as the foreground instance, while the textual cue information is used to obtain background prior information. This supplements detailed information from extreme environments and resolves blurring issues in extremely degraded scenes. Next, the foreground and background prior visual feature vectors are input into a two-end prototype resampling model to generate a set of foreground and background visual perception marker feature vectors. The two-end prototype resampling model includes a prototype resampling model for the foreground instance and a background resampling model to extract rich visual structural information from the foreground instance and background, enhancing the perceived visual difference between the two. Then, a context-decoupled feature vector set is generated based on the latent image feature vector, the object geometric layout information, the foreground visual perception marker feature vector set, the text prompt information, and the background visual perception marker feature vector. Here, the context-decoupled feature vector set includes a foreground decoupling feature vector set and a background decoupling feature vector. The foreground instances and background are decoupled by interacting with the latent image feature vector through parallel decoupling of the foreground instances and background. Subsequently, the context-decoupled feature vector set is input into a context topology coordination model to obtain a context topology correction feature vector set.Here, the topological relationship coordination model can correct the mutual topological associations of the contexts by constructing a topological graph including at least one foreground instance and background, then extracting and strengthening occlusion connectivity relationships within the context to suppress irrelevant and redundant connections. The context topological relationship correction feature vector set is then input into a dynamic semantic relationship aggregation model to obtain a context semantic feature vector. The dynamic semantic relationship aggregation model aggregates common and differential semantic information of contexts corresponding to different categories with different weights to improve the comprehensiveness and visual detail of the context semantic information. The context semantic feature vector and the context decoupling feature vector set are then weightedly fused to obtain a degraded visual feature vector. By assigning different attention weights to the foreground instance and background, more visual detail can be extracted, increasing the visual detail included in the degraded visual feature vector, enabling subsequent decoding to generate a higher-quality degraded image. Finally, the degraded visual feature vector is decoded to obtain a target degraded image. Consequently, this method for simulating and generating degraded images in extreme environments can improve the quality of images in extremely degraded scenarios.
[0104] Further references Figure 4 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an extreme environment degradation image simulation generation device. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the extreme environment degradation image simulation generation device can be specifically applied to various electronic devices.
[0105] like Figure 4As shown, a device 400 for simulating and generating degraded images in extreme environments includes: an acquisition unit 401, a first input unit 402, a first generation unit 403, a second input unit 404, a second generation unit 405, a fourth input unit 407, a fusion unit 408, and a decoding unit 409. The acquisition unit 401 is configured to acquire a degraded image, text prompt information, and object geometric layout information in an extreme environment. The first input unit 402 is configured to input the degraded image into a diffuse variational autoencoder to obtain a latent image feature vector. The first generation unit 403 is configured to generate a foreground prior visual feature vector set and a background prior visual feature vector based on the text prompt information and the object geometric layout information. The second input unit 404 is configured to input the foreground prior visual feature vector set and the background prior visual feature vector into a two-end prototype resampling model to obtain a foreground visual perception label feature vector set and a background visual perception label feature vector. The second generation unit 405 is configured to generate a context decoupling feature vector set based on the above-mentioned latent image feature vector, the above-mentioned object geometric layout information, the above-mentioned foreground visual perception mark feature vector set, the above-mentioned text prompt information, and the above-mentioned background visual perception mark feature vector. The third input unit 406 is configured to input the above-mentioned context decoupling feature vector set into the context topological relationship coordination model to obtain a context topological relationship correction feature vector set. The fourth input unit 407 is configured to input the above-mentioned context topological relationship correction feature vector set into the dynamic semantic relationship aggregation model to obtain a context semantic feature vector. The fusion unit 408 is configured to perform weighted fusion processing on the above-mentioned context semantic feature vector and the above-mentioned context decoupling feature vector set to obtain a degraded visual feature vector. The decoding unit 409 is configured to perform decoding processing on the above-mentioned degraded visual feature vector to obtain a target degraded image.
[0106] It is understandable that the units described in the extreme environment degradation image simulation generation device 400 are similar to those in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the extreme environment degradation image simulation generation device 400 and the units included therein, and will not be repeated here.
[0107] Reference below Figure 5 , which shows a structural schematic diagram of an electronic device (eg, an electronic device) 500 suitable for implementing some embodiments of the present disclosure. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0108] like Figure 5As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0109] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 5 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0110] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0111] It should be noted that in some embodiments of the present disclosure, the computer-readable medium mentioned above may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0112] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0113] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains degraded images, text prompt information and object geometric layout information under extreme environments; inputs the above-mentioned degraded images into the diffusion variational autoencoder to obtain a potential image feature vector; generates a foreground prior visual feature vector set and a background prior visual feature vector based on the above-mentioned text prompt information and the above-mentioned object geometric layout information; inputs the above-mentioned foreground prior visual feature vector set and the above-mentioned background prior visual feature vector into a two-end prototype resampling model to obtain a foreground visual perception label feature vector set and a background visual perception label feature vector; according to the above-mentioned potential The image feature vector, the geometric layout information of the above-mentioned object, the foreground visual perception mark feature vector set, the text prompt information, and the background visual perception mark feature vector are used to generate a context decoupling feature vector set; the context decoupling feature vector set is input into a context topological relationship coordination model to obtain a context topological relationship correction feature vector set; the context topological relationship correction feature vector set is input into a dynamic semantic relationship aggregation model to obtain a context semantic feature vector; the context semantic feature vector and the context decoupling feature vector set are weightedly fused to obtain a degraded visual feature vector; the degraded visual feature vector is decoded to obtain a target degraded image.
[0114] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0116] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor. For example, they may be described as follows: a processor comprising an acquisition unit, a first input unit, a first generation unit, a second input unit, a second generation unit, a third input unit, a fourth input unit, a fusion unit, and a decoding unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the acquisition unit may also be described as a "unit for acquiring degraded images, text prompt information, and object geometric layout information in extreme environments."
[0117] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0118] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for simulating and generating an image degradation under extreme conditions, comprising: Obtain degraded images, text prompt information, and object geometric layout information in extreme environments; Inputting the degraded image into a diffusion variational autoencoder to obtain a latent image feature vector; generating a foreground priori visual feature vector set and a background priori visual feature vector according to the text prompt information and the object geometric layout information; Inputting the foreground prior visual feature vector set and the background prior visual feature vector into a double-ended prototype resampling model to obtain a foreground visual perception mark feature vector set and a background visual perception mark feature vector; generating a context-decoupled feature vector set according to the potential image feature vector, the object geometric layout information, the foreground visual perception mark feature vector set, the text prompt information, and the background visual perception mark feature vector; Inputting the context decoupling feature vector set into a context topology relationship coordination model to obtain a context topology relationship correction feature vector set; Inputting the context topology relationship correction feature vector set into a dynamic semantic relationship aggregation model to obtain a context semantic feature vector; Performing weighted fusion processing on the context semantic feature vector and the context decoupling feature vector set to obtain a degraded visual feature vector; The degraded visual feature vector is decoded to obtain a target degraded image.
2. The method according to claim 1, wherein The step of generating a foreground priori visual feature vector set and a background priori visual feature vector set according to the text prompt information and the object geometric layout information includes: Retrieving a candidate instance group of the same category as each semantic information included in the geometric layout information of the object from a preset candidate instance dictionary as a target candidate instance group, thereby obtaining a target candidate instance group set; Based on the target candidate instance set, perform the following average pooling steps: performing image processing on each target candidate instance in the target candidate instance set to generate a processed candidate instance, thereby obtaining a processed candidate instance set; Inputting the processed candidate instance set into an image convolutional feature extraction model to obtain a foreground instance visual feature vector set; performing average pooling processing on the foreground instance visual feature vector set to obtain a foreground prior visual feature vector set; Retrieving a training background image that matches the text prompt corresponding to the degraded image from a preset training background image set as a target training background image; The target training background image is determined as a target candidate instance set to perform the average pooling step again, and the obtained foreground prior visual feature vector set is used as the background prior visual feature vector.
3. The method according to claim 1, wherein: The dual-end prototype resampling model includes: a foreground prototype resampling model and a background prototype resampling model; and The step of inputting the foreground prior visual feature vector set and the background prior visual feature vector into a two-terminal prototype resampling model to obtain a foreground visual perception mark feature vector set and a background visual perception mark feature vector comprises: Obtaining a foreground learnable query tag corresponding to the foreground prior visual feature vector set and a background learnable query tag corresponding to the background prior visual feature vector; Performing linear transformation processing on the foreground prior visual feature vector set to obtain a foreground prior key vector set and a foreground prior value vector set; Inputting the foreground learnable query tag, the foreground prior key vector set, and the foreground prior value vector set into the foreground prototype resampling model to obtain a foreground visual perception tag feature vector set, wherein the foreground prototype resampling model includes: multiple cross attention layers and multiple feedforward neural networks; Performing linear transformation on the background prior visual feature vector to obtain a background prior key vector and a background prior value vector; The background learnable query tag, the background prior key vector and the background prior value vector are input into the background prototype resampling model to obtain a background visual perception tag feature vector.
4. The method according to claim 1, wherein: The context topology relationship coordination model includes: a relationship coordination model and a graph convolutional neural network model; and The step of inputting the context decoupling feature vector set into a context topology relationship coordination model to obtain a context topology relationship correction feature vector set includes: Determining a context adjacency matrix of the context-decoupled feature vector set through the relationship coordination model; The graph convolutional neural network model is used to perform graph symmetric normalization processing on the context adjacency relationship matrix to obtain a context topology relationship correction feature vector set.
5. The method according to claim 1, wherein: The dynamic semantic relationship aggregation model includes: a gating network, multiple expert networks and an aggregation operation network; and The step of inputting the context topology relationship correction feature vector set into a dynamic semantic relationship aggregation model to obtain a context semantic feature vector includes: Get the normally distributed noise vector; Inputting the normally distributed noise vector and the context topology relationship correction feature vector set into the gating network to obtain expert activation probability weight sets for the multiple expert networks; Screening out at least two expert activation probability weights that meet a preset activation probability condition from the expert activation probability weight set; Determining at least two expert networks corresponding to the at least two expert activation probability weights from the multiple expert networks; Inputting the context topology relationship correction feature vector set into the at least two expert networks to obtain a context semantic aggregation feature vector set; The context semantic aggregated feature vector set and the at least two expert activation probability weights are input into the aggregation operation network to obtain a context semantic feature vector.
6. The method according to claim 1, wherein: The performing weighted fusion processing on the context semantic feature vector and the context decoupling feature vector set to obtain a degraded visual feature vector includes: Normalizing the context semantic feature vector to obtain a context semantic relationship aggregation weight; Multiplying the context semantic relationship aggregation weight by each context decoupling feature vector in the context decoupling feature vector set element by element to obtain a context weight feature vector set; The context weight feature vector set is accumulated to obtain a degraded visual feature vector.
7. A device for generating simulated images of extreme environmental degradation, comprising: an acquisition unit configured to acquire degraded images, text prompt information, and object geometric layout information under extreme environments; A first input unit is configured to input the degraded image into a diffusion variational autoencoder to obtain a latent image feature vector; A first generating unit is configured to generate a foreground priori visual feature vector set and a background priori visual feature vector according to the text prompt information and the object geometric layout information; A second input unit is configured to input the foreground prior visual feature vector set and the background prior visual feature vector into a two-terminal prototype resampling model to obtain a foreground visual perception mark feature vector set and a background visual perception mark feature vector; a second generating unit configured to generate a context-decoupled feature vector set based on the potential image feature vector, the object geometric layout information, the foreground visual perception mark feature vector set, the text prompt information, and the background visual perception mark feature vector; A third input unit is configured to input the context decoupling feature vector set into a context topology relationship coordination model to obtain a context topology relationship correction feature vector set; a fourth input unit configured to input the context topology relationship correction feature vector set into a dynamic semantic relationship aggregation model to obtain a context semantic feature vector; a fusion unit configured to perform weighted fusion processing on the context semantic feature vector and the context decoupling feature vector set to obtain a degraded visual feature vector; The decoding unit is configured to perform decoding processing on the degraded visual feature vector to obtain a target degraded image.
8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Image defogging method based on color correction and context aggregation residual network
CN112991201A
Image dehazing and restoration
US20190114747A1