Infrared ship scene generation method and system based on collaborative optimization of physical prior model and data-driven algorithm, and storage medium
By combining physical prior model and data-driven algorithm, using technologies such as DeepLabV3+, ControlCom and cross-modal perceptual style transfer networks, efficient and real infrared ship images are generated, solving the problems of complex and slow speed of traditional methods and physical distortion of data-driven methods, and providing high-quality infrared ship detection data.
Patent Information
- Application Number
- CN202510286793.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-11
AI Technical Summary
The existing infrared image generation technology has a contradiction between generation efficiency and physical consistency. Traditional physical simulation methods are complex and slow, while the data-driven method generates the result physical distortion, making it difficult to generate high-quality infrared ship images in complex environments.
Using a method based on the physical prior model and data-driven algorithm to optimize the method, efficient and real infrared ship images are generated through technologies such as DeepLabV3+ semantic segmentation, ControlCom image synthesis, cross-modal perceptual style transfer network and non-uniform weighted guide filtering.
The generated infrared ship images are rich in details and the grayscale values are real, which can improve the ship recognition rate under complex sea conditions, solve the problems of high infrared image generation costs and data shortage, and provide rich maritime detection data.
Smart Images

Figure CN120298879A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of the integration of infrared image generation and artificial intelligence, and in particular, to an infrared ship scene generation method, system and storage medium based on the collaborative optimization of a physical prior model and a data-driven algorithm. Background Art
[0002] The current ship image generation technology based on infrared imaging faces significant challenges in practical applications: First, affected by the heat radiation characteristics and environmental noise interference, the generated images often have problems such as blurred target edges and lost texture details, resulting in a decrease in the recognition rate of small and weak ship targets; Second, the interference of multi-scale wave heat sources under complex sea conditions is likely to induce artifacts, causing a false alarm risk; More critically, the existing methods have insufficient adaptability to dynamic environments such as sudden changes in temperature and humidity and rain and fog occlusion, and it is difficult to ensure the imaging stability under all-weather scenarios. However, the above superficial problems essentially reflect the defects of the infrared image generation technology.
[0003] The current image translation technology from visible light to infrared images includes:
[0004] Traditional physical methods. It mainly divides into two approaches: The first is to directly use a computer to perform three-dimensional simulation modeling on the real space, solve the temperature of each part in the scene by establishing heat balance and heat conduction equations, and finally calculate the infrared intrinsic radiation of the scene and the infrared radiation atmospheric transmission through infrared physics methods to obtain the final infrared image. The second is to generate derivatives using real-shot two-dimensional visible light images. By presetting the gray value of the visible light image and performing material segmentation, temperature and radiation calculations, etc., an infrared image is obtained. This method relies on accurate environmental parameter inputs (such as temperature distribution, material emissivity), has a high modeling complexity, poor generalization, and a long debugging period;
[0005] Data-driven methods: mainly adopt generative adversarial networks (GANs) or diffusion models to learn the mapping relationship from visible light-infrared paired data. Such methods have a strong dependence on data quality and scale, and the generated results are prone to deviate from physical laws, and many distorted and blurred phenomena, as well as some unrealistic artifacts without physical model basis, are likely to occur.
[0006] Therefore, there is an urgent need to provide a new infrared image generation technology to solve the contradiction between physical consistency and generation efficiency (traditional physical simulation has high accuracy but low speed, and data-driven methods have high speed but physical distortion). Summary of the Invention
[0007] The present invention discloses an infrared ship scene generation method, system and storage medium based on the collaborative optimization of a physical prior model and a data-driven algorithm. By means of a physical regression model and an artificial intelligence generation method, a visible light background image is used to generate an infrared ship image. This method solves the contradiction between physical consistency and generation efficiency, and at the same time solves the problems of high cost of infrared ship image acquisition and data shortage. The enhancement of infrared ship data can be used in scenarios such as maritime ship detection and recognition.
[0008] To solve the above technical problems, the technical solution of the present invention is:
[0009] According to the first aspect of the technical solution of the present invention, there is provided an infrared ship scene generation method based on the collaborative optimization of a physical prior model and a data-driven algorithm, and the method includes the following steps:
[0010] S1: Preprocess the collected real infrared ship images and original visible light ship images;
[0011] S2: Use the DeepLabV3+ standard architecture to perform semantic segmentation on the preprocessed original visible light ship images, establish a ship database of different types, and obtain visible light ship foreground images;
[0012] S3: Use the ControlCom controllable image synthesis model to synthesize the background image of the preprocessed original visible light ship image and the visible light ship foreground image of a specified type into a new visible light ship image;
[0013] S4: Use the cross-modal perception style transfer network CPSTN to convert the new visible light ship image into a pseudo-infrared image;
[0014] S5: Segment the new visible light ship image and classify it according to materials, cover the segmented image mask onto the corresponding preprocessed real infrared ship image, extract the real infrared gray values of the ship part, and establish a linear regression model to output the gray values of ships of different materials;
[0015] S6: Use a multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering to optimize the pseudo-infrared image, and use the gray values output by the linear regression model to correct the gray level of the ship area, and finally generate an infrared ship scene.
[0016] Further, in the S1, the preprocessing operation includes: image correction, filtering and denoising, registration and cropping, and storing the collected real infrared ship images and original visible light ship images in pairs with the same name.
[0017] Further, S1 also includes forming a data set and dividing the data set into a training set and a validation set; wherein, the data set is divided into a training set and a validation set according to 8:2.
[0018] Further, in S2, the input is the preprocessed visible light ship image in RGB format. The encoder uses a pre-trained Xception-65 backbone network to extract multi-scale features: the shallow feature map retains high-frequency details; the deep feature map analyzes the global structure; the ASPP module fuses multi-scale context information to enhance the recognition robustness of ship edges under complex sea conditions; the decoder restores spatial details step by step through upsampling and skip connection with shallow features, and outputs a four-channel segmentation probability map, thereby establishing a ship database of different types.
[0019] Further, the high-frequency details include that the shallow feature map retains the texture of the sailing ship mast, the wake waves of the speedboat, etc.; the global structure includes that the deep feature map analyzes the arrangement of the cargo holds of the steamship, the deck layout of the fishing boat, etc.
[0020] Further, in S2, a ship database of four types, namely sailing ships, fishing boats, speedboats, and steamships, is established.
[0021] Further, in S3, the ControlCom controllable image synthesis model includes a foreground encoder and a controllable generator.
[0022] Further, the foreground encoder includes an image encoder and a global embedding module E g and a local embedding module E l .
[0023] Further, the input of the foreground encoder is the visible light ship foreground image I f of a specified type and the ship category semantics encoded by the Class Token;
[0024] The image encoder uses Patch Tokens to encode the visible light ship foreground image I f of a specified type into multi-scale features;
[0025] The global embedding module E g aggregates ship category information through the Class Token, and the MLP further maps it into high-dimensional global features to generate the global semantic features of the ship;
[0026] The local embedding module E l retains fine-grained information from the underlying features of the Patch Tokens and extracts the local detail features of the ship.
[0027] Further, the global semantic features include ship type and overall shape.
[0028] Furthermore, the local detail features include hull texture, masts, windows, etc.
[0029] Furthermore, the controllable generator is a U-Net model based on Stable Diffusion, and the inputs include:
[0030] The latent encoding of the background image of the preprocessed original visible light ship image;
[0031] Indicator Map S, which is used to control whether foreground lighting and pose need to be adjusted;
[0032] A mask, which is used to determine the bounding box information and identify the area where the foreground will be placed in the U-Net model based on Stable Diffusion.
[0033] Furthermore, each Transformer Block of the U-Net modified based on Stable Diffusion includes the following key steps:
[0034] Residual Block+Self Attention, which is used to perform convolution / residual on the features and then perform self-attention to capture the overall background context;
[0035] Global Fusion, which is used to integrate the features of the global embedding module E into the features of the U-Net modified based on Stable Diffusion in the Cross Attention of each Transformer Block; g into the features of the U-Net modified based on Stable Diffusion;
[0036] Local Enhancement: In the local enhancement module, the features of the U-Net modified based on Stable Diffusion obtain the local features corresponding to the bounding box, and then perform Cross Attention with the local embedding of the foreground, the local embedding module E; l The features of; It also includes Feature Modulation, which is used to fuse the aligned foreground embedding map with the local features;
[0037] Indicator Map(S): A 2D vector, which is combined with the local enhancement module after being extended into a local feature map;
[0038] Among them, the first channel controls whether to adjust the lighting (0 = keep, 1 = change), and the second channel controls whether to adjust the pose (0 = keep, 1 = change):
[0039] If pose = 0, the model tends to "maintain the original ship pose";
[0040] If pose = 1, then "perform a perspective transformation in the hull features";
[0041] If illumination = 1, make appropriate changes to color, brightness, etc. to match the background;
[0042] If illumination = 0, retain the original appearance of the ship.
[0043] Furthermore, the controllable generator embeds the visible-light ship foreground image of a specified type into the background image of the original visible-light ship image through a rectangular box module based on the bounding box information, and uses a residual fusion module, combined with self-attention, global semantic fusion, and local detail enhancement to optimize the illumination consistency and edge transition. Finally, a new visible-light ship image is gradually denoised through a U-Net model modified based on StableDiffusion.
[0044] Furthermore, in S3, the background image and the foreground image can be from the same image or different images.
[0045] Furthermore, in S4, the cross-modal perception style transfer network CPSTN includes a generator G(A), a generator G(B), a discriminator Dir, and a discriminator Dvis.
[0046] Furthermore, the generator G(A) and the generator G(B) adopt a U-shaped encoder-decoder architecture and contain 9 residual blocks at the bottom.
[0047] Furthermore, the generator G(A) is used to convert the new visible-light ship image I vis into a pseudo-infrared image I' ir , which is generated by extracting multi-scale features through an encoder and decoding after fusing infrared style features through residual blocks;
[0048] The generator G(B) is used to reversely generate a visible-light image I' ir from the pseudo-infrared image I' vis , thereby constructing a cyclic consistency constraint;
[0049] The discriminator Dir distinguishes the real infrared image I ir from the pseudo-infrared image I' ir through adversarial training, enabling the generator to learn the distribution characteristics of thermal radiation features, low-texture details, etc. of the infrared modality;
[0050] The discriminator Dvis discriminates the new visible-light ship image I vis from the reversely generated visible-light image I'vis the authenticity to ensure the stability of the loop path.
[0051] Further, in the step S5, the input parameters of the linear regression model are temperature, humidity, wind speed, weather, air pressure, horizontal surface radiation, normal direct radiation, and material, and the output parameter is the gray value of different materials.
[0052] Further, in the step S5, the specific steps of segmenting and classifying the new visible light ship image according to the material include:
[0053] Label the visible light ship image according to the material;
[0054] Use the labeled visible light ship image as the input to train the segmentation model;
[0055] Use the new visible light ship image as the input of the trained segmentation model;
[0056] The output is the visible light ship image segmented by the trained segmentation model.
[0057] Further, in the step S5, the specific steps of establishing a linear regression model and outputting the gray values of ships of different materials include:
[0058] According to the registered real infrared ship image and the original visible light ship image, cover the segmented new visible light ship image (mask) onto the real infrared ship image to extract the gray value of the real infrared image;
[0059] Construct a regression model according to the input parameters of the linear regression model;
[0060] After segmenting the new visible light image according to the material, obtain the gray values of ships of different materials according to the regression model.
[0061] Further, in the step S6, the multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering includes multi-scale decomposition, directional gradient enhancement, adaptive noise suppression, and weighted fusion.
[0062] Further, the steps of optimizing the pseudo-infrared image by using the multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering include:
[0063] Decompose the pseudo-infrared image into a single-layer base layer and multi-scale detail layers, where the edge structure is retained by non-local mean weights and local guidance kernels;
[0064] Design a multi-directional gradient operator for the detail layer, dynamically adjust the template parameters by combining local entropy and gray values, extract detail features of different scales and directions, and adaptively enhance the effective details through a differential gain function. Among them, dynamically suppress noise based on a noise masking model, and at the same time enhance the texture contrast according to the directional gradient response;
[0065] Combine the enhanced multi-scale detail layer with the base layer after brightness correction through entropy value weighted fusion to highlight the details in the information-rich areas, and suppress noise while enhancing the sharpness of edges such as ship contours and sea wave textures.
[0066] Further, the multi-directional gradient operator includes horizontal, vertical, and diagonal.
[0067] Further, use local variance and gradient information to distinguish noise from real details.
[0068] According to the second aspect of the technical solution of the present invention, there is provided an infrared ship scene generation system based on the collaborative optimization of a physical prior model and a data-driven algorithm. The system includes: a processor and a memory for storing executable instructions; wherein, the processor is configured to execute the executable instructions to perform the infrared ship scene generation method based on the collaborative optimization of the physical prior model and the data-driven algorithm described in any of the above aspects.
[0069] According to the third aspect of the technical solution of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the infrared ship scene generation method based on the collaborative optimization of the physical prior model and the data-driven algorithm described in any of the above aspects.
[0070] The beneficial technical effects of the present invention:
[0071] 1. For the generation of infrared images, the present invention adopts a method that combines data-driven and physical regression models to generate infrared ship images with high efficiency, authenticity, and rich infrared detail information;
[0072] 2. The research method for generating infrared ship images of the present invention can solve problems such as inconvenient collection and high cost of maritime infrared images. By using a maritime visible light background picture as the input, an infrared ship image with relatively rich detail information and relatively real gray values can be obtained, providing a more abundant data set for tasks such as maritime ship detection and recognition. Description of the Drawings
[0073] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0074] Figure 1 is a flowchart of infrared ship image generation based on physical prior and data-driven provided by an embodiment of the present invention;
[0075] Figure 2 is the interface made by this system.
[0076] The realization of the purpose of the present invention, functional features and advantages will be further described with reference to the embodiments and the drawings. Detailed implementation manners
[0077] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0078] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein.
[0079] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0080] Multiple, including two or more than two.
[0081] And / or, it should be understood that for the term "and / or" used in the present disclosure, it is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Reducing human input and assisting business automation, with the characteristics of generality, high efficiency and high precision.
[0082] The technical solution of the present invention provides an infrared ship scene generation method, system and storage medium based on the collaborative optimization of a physical prior model and a data-driven algorithm, and a solution for generating infrared ship images through an infrared ship scene generation model. First, the self-collected visible-light ship images are segmented by DeepLabV3+ to construct a ship database; secondly, the diffusion model ControlCom is used to synthesize the visible-light background and the ship foreground to generate visible-light images containing four types of ships; then, the visible-light images are converted into pseudo-infrared images through a perceptual style transfer network; subsequently, the visible-light ship images are segmented according to the material, the details are extracted, and the segmentation mask is covered on the corresponding infrared images to extract the real infrared gray values; finally, a linear regression model is established, and the gray values output by the model and the multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering are used to optimize the details and correct the gray levels of the pseudo-infrared images, so as to generate infrared ship images with rich details and real gray levels.
[0083] Specifically, the technical solution of the present invention first provides an infrared ship scene generation method based on the collaborative optimization of a physical prior model and a data-driven algorithm, and the method includes the following steps:
[0084] S1: Preprocess the collected real infrared ship images and the original visible-light ship images.
[0085] In a preferred embodiment, in the S1, the preprocessing operations include: image correction, filtering and denoising, registration and cropping, and the collected real infrared ship images and the original visible-light ship images are stored in pairs with the same name.
[0086] In a preferred embodiment, the S1 further includes forming a data set, and dividing the data set into a training set and a validation set; wherein, the data set is divided into a training set and a validation set according to 8:2.
[0087] S2: Use the DeepLabV3+ standard architecture to perform semantic segmentation on the preprocessed original visible-light ship images, establish a ship database of different types, and obtain visible-light ship foreground images.
[0088] In a preferred embodiment, in S2, the input is the preprocessed visible-light ship image in RGB format. The encoder uses a pre-trained Xception-65 backbone network to extract multi-scale features: the shallow feature map retains high-frequency details; the deep feature map analyzes the global structure; the ASPP module fuses multi-scale context information to enhance the recognition robustness of ship edges under complex sea conditions; the decoder restores spatial details step by step through upsampling and skip connections with shallow features, outputs a four-channel segmentation probability map, and then establishes a ship database of different types.
[0089] In a preferred embodiment, the high-frequency details include that the shallow feature map retains the texture of a sailboat mast, the wake of a speedboat, etc.; the global structure includes that the deep feature map analyzes the arrangement of a steamship's cargo holds, the layout of a fishing boat's deck, etc.
[0090] In a preferred embodiment, in S2, a ship database of four types, namely sailboats, fishing boats, speedboats, and steamships, is established.
[0091] S3: Use the ControlCom controllable image synthesis model to synthesize a new visible-light ship image from the background image of the preprocessed original visible-light ship image and the visible-light ship foreground image of a specified type.
[0092] In a preferred embodiment, in S3, the ControlCom controllable image synthesis model includes a foreground encoder and a controllable generator.
[0093] In a preferred embodiment, the foreground encoder includes an image encoder, a global embedding module E g and a local embedding module El.
[0094] In a preferred embodiment, the input of the foreground encoder is the visible-light ship foreground image I of a specified type f and the ship category semantics encoded by the Class Token;
[0095] The image encoder uses Patch Tokens to encode the visible-light ship foreground image I of a specified type f into multi-scale features;
[0096] The global embedding module E g aggregates ship category information through the Class Token, and the MLP further maps it into high-dimensional global features to generate the global semantic features of the ship;
[0097] The local embedding module El retains fine-grained information from the underlying features of the Patch Tokens and extracts the local detail features of the ship.
[0098] In a preferred embodiment, the global semantic features include ship type and overall shape.
[0099] In a preferred embodiment, the local detail features include hull texture, masts, windows, etc.
[0100] Here, the foreground encoder typically uses an image encoder to obtain the class token (as global information) and intermediate layer patch tokens (as local information), and then through transformations such as MLP / convolution to obtain the global embedding E g and the local embedding E l .
[0101] The global embedding (E g ) is used to extract the semantic information of the whole ship (ship type, overall shape, main structure, etc.). E g will subsequently be injected into the extended Stable Diffusion U-Net to replace the original text prompt, telling the network "what ship to generate" and the general outline. The local embedding (E l ) is used to extract finer-grained textures and details (such as the color of the hull surface, text markings, windows, deck and other fine features). The local embedding E l provides details in the local enhancement module to keep the hull in the synthesized image as consistent as possible with the appearance of the original foreground.
[0102] The core of the controllable generator is a U-Net modified based on Stable Diffusion, labeled "Diffusion U-Net" in the figure. Its inputs include:
[0103] · The latent encoding of the background image (the latent obtained through the VAE Encoder, or directly represented by small squares / feature maps in this schematic diagram);
[0104] · The Indicator Map S, used to control whether to adjust the foreground lighting and pose;
[0105] · The mask (bounding box information), used to identify the area in the U-Net where the foreground will be placed.
[0106] In each Transformer Block of this U-Net, there are mainly the following key steps:
[0107] 1. Residual Block + Self Attention
[0108] Perform convolution / residual on the features and then self-attention to capture the overall background context.
[0109] 2. Global Fusion
[0110] In the Cross Attention of each Transformer Block, the Eg (global ship feature) is incorporated into the features of the U-Net. Ensure that the network "knows" what type and general shape of ship to synthesize and make compatibility adjustments in the background (such as adapting to background lighting, sea surface style, etc.).
[0111] 3. Local Enhancement
[0112] In the local enhancement module, the features of the U-Net first obtain the local features corresponding to the bounding box through RoIAlign (or other means), and then perform Cross Attention with the local embedding El of the foreground. This can inject the detailed texture of the foreground object into this area, and further control whether the lighting and pose need to change with the Indicator Map S. Finally, there is a Feature Modulation step to fuse the aligned foreground embedding map with the local features to retain the ship details and make a natural transition on the background.
[0113] 4. Indicator Map(S)
[0114] This is a 2D vector, which is combined with the local enhancement module after being extended into a local feature map.
[0115] The first channel controls whether to adjust the lighting (0 = keep, 1 = change), and the second channel controls whether to adjust the pose (0 = keep, 1 = change): if the pose = 0, the model tends to "keep the original ship pose"; if the pose = 1, it will "perform a perspective transformation in the hull features"; if the lighting = 1, make appropriate changes to the color, brightness, etc. to match the background; if the lighting = 0, retain the original appearance of the ship.
[0116] In a preferred embodiment, in S3, the background image and the foreground image can be from the same image or different images.
[0117] In fact, the input of the model is a combination including a background image, a foreground image, a corresponding mask, a foreground position bounding box, and an indicator vector that controls the lighting and pose changes of the foreground. In the training stage, the foreground and background are often extracted from the same image, and foreground and background synthesis from different images is supported during image generation, so as to achieve flexible and controllable image synthesis.
[0118] S4: Use the cross-modal perception style transfer network CPSTN to convert the new visible-light ship image into a pseudo-infrared image.
[0119] In a preferred embodiment, in S4, the cross-modal perception style transfer network CPSTN includes a generator G(A), a generator G(B), a discriminator Dir, and a discriminator Dvis.
[0120] In a preferred embodiment, the generator G(A) and the generator G(B) adopt a U-shaped encoder-decoder architecture and contain 9 residual blocks at the bottom.
[0121] In a preferred embodiment, the generator G(A) is used to convert the new visible-light ship image I vis into a pseudo-infrared image I′ ir , which is generated by extracting multi-scale features through an encoder and decoding after fusing infrared style features through residual blocks;
[0122] The generator G(B) is used to reversely generate a visible-light image I′ ir from the pseudo-infrared image I′ vis , thereby constructing a cyclic consistency constraint;
[0123] The discriminator Dir distinguishes the real infrared image I ir from the pseudo-infrared image I′ ir through adversarial training, enabling the generator to learn the distribution characteristics such as the thermal radiation characteristics and low-texture details of the infrared modality;
[0124] The discriminator Dvis discriminates the authenticity of the new visible-light ship image I vis and the reversely generated visible-light image I′ vis to ensure the stability of the cyclic path.
[0125] S5: Segment the new visible-light ship image and classify it according to the material, cover the segmented image mask mask onto the corresponding preprocessed real infrared ship image, extract the real infrared gray values of the ship part, and establish a linear regression model to output the gray values of ships of different materials.
[0126] In a preferred embodiment, in S5, the input parameters of the linear regression model are air temperature, humidity, wind speed, weather, air pressure, surface horizontal radiation, normal direct radiation, and material, and the output parameter is the gray value of different materials.
[0127] In a preferred embodiment, in S5, segmenting the new visible-light ship image and classifying it according to the material specifically includes:
[0128] Label the visible-light ship image according to the material;
[0129] Use the marked visible-light ship image as the input to the trained segmentation model;
[0130] Use the new visible-light ship image as the input to the trained segmentation model;
[0131] The output is the visible-light ship image segmented by the trained segmentation model.
[0132] In a preferred embodiment, in S5, establishing a linear regression model and outputting the gray values of ships of different materials specifically includes:
[0133] According to the registered real infrared ship image and the original visible-light ship image, cover the segmented new visible-light ship image (mask) onto the real infrared ship image to extract the gray values of the real infrared image;
[0134] According to some physical parameters of the visible light shooting, such as temperature, humidity, wind speed, precipitation, etc., construct a regression model of these physical parameters and the real gray values;
[0135] After segmenting the new visible-light image according to the material, obtain the gray values of ships of different materials according to the physical parameters and the model already input into the model.
[0136] S6: Optimize the pseudo-infrared image by using a multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering, correct the gray level of the ship area by using the gray values output by the linear regression model, and finally generate an infrared ship scene.
[0137] In a preferred embodiment, in S6, the multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering includes multi-scale decomposition, directional gradient enhancement, adaptive noise suppression, and weighted fusion.
[0138] In a preferred embodiment, optimizing the pseudo-infrared image by using a multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering includes the following steps:
[0139] Decompose the pseudo-infrared image into a single-layer base layer and multi-scale detail layers, where the edge structure is retained through non-local mean weights and local guidance kernels;
[0140] Design a multi-directional gradient operator for the detail layers, dynamically adjust the template parameters by combining local entropy and gray values, extract detail features of different scales and directions, and adaptively enhance the effective details through a differential gain function, where the noise is dynamically suppressed based on a noise masking model, and at the same time, the texture contrast is enhanced according to the directional gradient response;
[0141] The enhanced multi-scale detail layer is combined with the brightness-corrected base layer through entropy-weighted fusion to highlight the details in information-rich regions, and suppress noise while enhancing the sharpness of edges such as ship contours and sea wave textures.
[0142] In a preferred embodiment, the multi-directional gradient operator includes horizontal, vertical, and diagonal directions.
[0143] In a preferred embodiment, local variance and gradient information are used to distinguish noise from real details.
[0144] The technical solution of the present invention further provides an infrared ship scene generation system based on the collaborative optimization of a physical prior model and a data-driven algorithm. The system includes: a processor and a memory for storing executable instructions; wherein, the processor is configured to execute the executable instructions to perform the infrared ship scene generation method based on the collaborative optimization of the physical prior model and the data-driven algorithm as described above.
[0145] The technical solution of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the infrared ship scene generation method based on the collaborative optimization of the physical prior model and the data-driven algorithm as described above.
[0146] Embodiment
[0147] This embodiment provides a method for generating an infrared ship image using a model for generating an infrared ship scene. The method includes the following steps:
[0148] Production of the data set: Preprocess the collected ship infrared images and visible light images, including image correction, filtering and denoising, registration, and cropping. Name the infrared images and visible light images in pairs and store them. Use the preprocessed ship infrared images and visible light images to produce a data set, and divide the data set into a training set and a validation set according to 8:2.
[0149] Building a ship database: The dataset used in this invention is the dataset taken by itself, and semantic segmentation of ship types is performed on visible light images. In this embodiment, the DeepLabV3+ standard architecture is used to achieve four-class semantic segmentation of visible light ship images. The network input is an RGB ship image, and the encoder uses a pre-trained Xception-65 backbone network to extract multi-scale features: the shallow feature map retains high-frequency details such as the texture of the sailboat mast and the wake of the speedboat; the deep feature map analyzes the global structure such as the arrangement of the cargo holds of the steamship and the deck layout of the fishing boat; the ASPP module fuses multi-scale context information to enhance the recognition robustness of ship edges under complex sea conditions; the decoder performs a skip connection with the shallow features through upsampling, and uses depthwise separable convolution to gradually restore spatial details, and outputs a four-channel segmentation probability map (corresponding to the categories of sailboats, fishing boats, speedboats, and steamships) to obtain the visible light foreground image of the ship and establish a ship database;
[0150] The ControlCom controllable image synthesis model is used to synthesize a new visible light image from the visible light background image and the ship foreground of the specified type. The ControlCom network includes a foreground encoder and a controllable generator. Among them, the foreground encoder consists of an image encoder, a global embedding E g and a local embedding El module. The input of the foreground encoder is the foreground image I f and the ship category semantics encoded by the Class Token; the image encoder module uses Patch Tokens to divide the image into local blocks and extract features, retaining spatial details, and encoding the foreground ship image I f into multi-scale features; the global embedding module E g aggregates the ship category information through the Class Token, and the MLP further maps it into high-dimensional global features to generate the global semantic features of the ship (such as ship type, overall shape); the local embedding El module retains fine-grained information from the underlying features of the Patch Tokens and extracts the local detail features of the ship (such as hull texture, mast, windows, etc.). The input of the controllable generator is the background image and the bounding box position. The foreground is embedded into the background through the rectangle box module, and the residual fusion module (combining self-attention, global semantic fusion, and local detail enhancement) is used to optimize the lighting consistency and edge transition. Finally, the high-fidelity image is gradually denoised through the diffusion U-Net model to ensure the natural fusion of the ship and the ocean background (such as waves, lighting), and at the same time support flexible control of the ship type and position;
[0151] Using the cross-modal perception style transfer network CPSTN, the RGB image in the visible light domain is cross-domain converted into a pseudo-infrared image. The network consists of two generators G(A), G(B) and two discriminators Dir and Dvis. The generator adopts a U-shaped encoder-decoder architecture, and its bottom contains 9 residual blocks. The role of the residual block is to enhance the network's learning ability for complex textures and geometric structures, avoid gradient disappearance, and decouple the content (the object contour of the visible light image) from the style (the thermal radiation distribution of the infrared image). The role of G(A) is to convert the visible light image I vis into a pseudo-infrared image I′ ir , which is generated by extracting multi-scale features through the encoder and decoding after fusing the infrared style features through the residual block; the role of G(B) is to reversely restore the pseudo-infrared image I′ ir into a visible light image for constructing a cyclic consistency constraint (I vis →I′ ir →I′ vis ); Dir distinguishes the real infrared image I ir from the generated image I′ ir through adversarial training, forcing the generator to learn the distribution characteristics such as the thermal radiation characteristics and low texture details of the infrared modality; Dvis discriminates the authenticity of the original visible light image I vis and the reversely generated visible light image I′ vis to ensure the stability of the cyclic path;
[0152] Segment the visible light ship image to enhance the detail information of the ship. The visible light ship image is divided according to the material, mainly including steel, fiberglass, aluminum alloy, glass, and wood;
[0153] Cover the mask of the segmented visible light ship image onto the corresponding infrared image, extract the real infrared gray value of the ship part, and establish a linear regression model. The input parameters of the regression model are air temperature, humidity, wind speed, weather, air pressure, surface horizontal radiation, normal direct radiation, and material, and the output parameter is the gray value of this material;
[0154] The details of the pseudo-infrared image are optimized by using a multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering. The multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering (NWGIF) optimizes the details of the pseudo-infrared image through the process of multi-scale decomposition → directional gradient enhancement → adaptive noise suppression → weighted fusion: First, the improved NWGIF is used to decompose the image into a single-layer base layer (low-frequency luminance information) and multi-scale detail layers (high-frequency textures). Among them, NWGIF retains the edge structure through non-local mean weights and local guidance kernels, avoiding gradient inversion of traditional filtering. Subsequently, a multi-directional gradient operator (horizontal, vertical, diagonal) is designed for the detail layers, and the template parameters are dynamically adjusted by combining local entropy and gray values to extract detail features at different scales and directions. The effective details are adaptively enhanced through a differential gain function, which dynamically suppresses noise based on a noise masking model (using local variance and gradient information to distinguish noise from real details), and at the same time enhances the texture contrast according to the directional gradient response. Finally, the enhanced multi-scale detail layers are combined with the base layer after brightness correction through entropy-weighted fusion to highlight the details in the information-rich areas. Ultimately, while enhancing the sharpness of edges such as ship contours and wave textures, noise is suppressed, and robustness is maintained under Gaussian noise interference. For the ship area, the gray values output by the regression model are used to correct the gray level of the ship area, making the detail information of the ship image richer and more in line with the gray values of real infrared images.
[0155] In summary, aiming at the problems of low modeling efficiency of physical simulation methods and physical distortion of the results generated by data-driven methods in the prior art, the present invention provides an intelligent generation method for infrared ship images based on the combination of physical prior constraints and data-driven. Another object of the present invention is to provide a large amount of infrared data for virtual perception test evaluation. Aiming at the problems of difficult acquisition of infrared ship data and poor image quality, infrared ship pictures with rich detail texture information, high image resolution, and prominent targets are provided.
[0156] The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
[0157] As certain terms are used in the specification and claims to refer to specific components. Those skilled in the art should understand that hardware manufacturers may use different terms to refer to the same component. The specification and claims do not use the difference in names as a way to distinguish components, but rather use the difference in the functions of components as the criterion for distinction. As used throughout the specification and claims, the terms "comprising" and "including" are open-ended terms and should therefore be interpreted as "comprising / including but not limited to". "Substantially" means within an acceptable error range. Those skilled in the art can solve the technical problem within a certain error range and basically achieve the technical effect. The subsequent description in the specification is of a preferred embodiment for implementing the present application, but the description is for the purpose of explaining the general principles of the present application and not for limiting the scope of the present application. The scope of protection of the present application shall be subject to what is defined by the appended claims.
[0158] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such commodity or system. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of another identical element in the commodity or system including the said element.
[0159] It should be understood that the term "and / or" used herein is merely a relationship describing the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0160] The above description shows and describes several preferred embodiments of the present application. However, as mentioned above, it should be understood that the present application is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the application concept described herein through the above teachings or the technology or knowledge in the relevant field. And the changes and variations made by those skilled in the art without departing from the spirit and scope of the present application should all be within the scope of the rights of the present application attached.
Claims
1. An infrared ship scene generation method based on the collaborative optimization of a physical prior model and a data-driven algorithm, the method comprising the following steps: S1: Preprocess the collected real infrared ship images and original visible light ship images; S2: Use the DeepLabV3+ standard architecture to perform semantic segmentation on the preprocessed original visible light ship images, establish a database of different types of ships, and obtain visible light ship foreground images; S3: Use the ControlCom controllable image synthesis model to synthesize a new visible light ship image from the background image of the preprocessed original visible light ship image and the visible light ship foreground image of a specified type; S4: Use the cross-modal perception style transfer network CPSTN to convert the new visible light ship image into a pseudo-infrared image; S5: Segment the new visible light ship image and classify it according to materials, cover the segmented image mask onto the corresponding preprocessed real infrared ship image, extract the real infrared gray values of the ship part, and establish a linear regression model to output the gray values of ships of different materials; S6: Use a multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering to optimize the pseudo-infrared image, and correct the gray values of the ship area using the gray values output by the linear regression model, and finally generate an infrared ship scene.
2. The infrared ship scene generation method according to claim 1, wherein In S2, the input is the preprocessed visible light ship image in RGB format, and the encoder uses a pre-trained Xception-65 backbone network to extract multi-scale features: The shallow feature map retains high-frequency details; the deep feature map analyzes the global structure; the ASPP module fuses multi-scale context information to enhance the recognition robustness of ship edges under complex sea conditions; the decoder restores spatial details step by step through upsampling and skip connections with shallow features, and outputs a four-channel segmentation probability map, and then establishes a database of different types of ships.
3. The infrared ship scene generation method according to claim 1, wherein In S3, the ControlCom controllable image synthesis model includes a foreground encoder and a controllable generator.
4. The infrared ship scene generation method according to claim 3, wherein, The foreground encoder includes an image encoder, a global embedding module E g and a local embedding module E l ; The input of the foreground encoder is a visible-light ship foreground image I of a specified type f and the ship class semantics encoded by the Class Token; The image encoder uses Patch Tokens to encode the visible light ship foreground image I of a specified type f into multi-scale features; The global embedding module E g Aggregates ship category information through the Class Token, and the MLP further maps it into high-dimensional global features to generate the global semantic features of the ship; The local embedding module E l Retains fine-grained information from the underlying features of Patch Tokens and extracts local detail features of the ship.
5. The infrared ship scene generation method according to claim 3, wherein The controllable generator is a U-Net model based on Stable Diffusion, and the inputs include: The latent encoding of the background image of the preprocessed original visible light ship image; An indicator vector mapping for controlling whether it is necessary to adjust the foreground illumination and pose; A mask for determining the bounding box information and identifying the area where the foreground will be placed in the U-Net model based on Stable Diffusion; Among them, the controllable generator embeds the visible light ship foreground image of a specified type into the background image of the original visible light ship image through a rectangular box module based on the bounding box information, and uses a residual fusion module, combined with self-attention, global semantic fusion and local detail enhancement to optimize the illumination consistency and edge transition, and finally gradually denoises through the U-Net model modified based on Stable Diffusion to generate a new visible light ship image.
6. The infrared ship scene generation method according to claim 1, wherein, In S4, the cross-modal perception style transfer network CPSTN includes a generator G(A), a generator G(B), a discriminator Dir, and a discriminator Dvis; among them, the generator G(A) and the generator G(B) adopt a U-shaped encoder-decoder architecture, and the bottom contains 9 residual blocks. Among them, the generator G(A) is used to convert the new visible light ship image I vis into a pseudo-infrared image I' ir , which is generated by extracting multi-scale features through an encoder, fusing infrared style features through residual blocks, and then decoding. The generator G(B) is used to inversely generate the visible light image I' from the pseudo-infrared image I'. ir vis , thereby constructing a cyclic consistency constraint. The discriminator Dir distinguishes the real infrared image I ir from the fake infrared image I' ir , enabling the generator to learn the distribution characteristics such as the thermal radiation characteristics and low-texture details of the infrared modality; The discriminator Dvis discriminates the new visible-light ship image I vis and the reversely generated visible-light image I′ vis for authenticity to ensure the stability of the loop path.
7. The infrared ship scene generation method according to claim 1, characterized in that In S6, the multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering includes multi-scale decomposition, directional gradient enhancement, adaptive noise suppression, and weighted fusion.
8. The infrared ship scene generation method according to claim 1, characterized in that Optimizing the pseudo-infrared image by using the multi-scale infrared image enhancement algorithm based on non-uniform weighted guided filtering includes the following steps: Decompose the pseudo-infrared image into a single-layer base layer and multi-scale detail layers, where the edge structure is retained by non-local mean weights and local guidance kernels. Design a multi-directional gradient operator for the detail layers, dynamically adjust the template parameters by combining local entropy and gray values, extract detail features of different scales and directions, and adaptively enhance the effective details through a differential gain function, where the noise is dynamically suppressed based on a noise masking model, and at the same time, the texture contrast is enhanced according to the directional gradient response. Combine the enhanced multi-scale detail layers with the brightness-corrected base layer through entropy value weighted fusion, highlight the details in the information-rich areas, and suppress the noise while enhancing the sharpness of edges such as ship contours and wave textures.
9. An infrared ship scene generation system based on the collaborative optimization of a physical prior model and a data-driven algorithm, the system comprising: A processor and a memory for storing executable instructions; characterized in that the processor is configured to execute the executable instructions to perform the infrared ship scene generation method based on the collaborative optimization of the physical prior model and the data-driven algorithm according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, it implements the infrared ship scene generation method based on the collaborative optimization of the physical prior model and the data-driven algorithm according to any one of claims 1 to 8.
Citation Information
Cited By
Visible light and infrared fusion-based river no-fishing ship monitoring method and system
CN120894696A
River fishing-prohibited-ship monitoring method and system based on visible light and infrared fusion
CN120894696B
Multi-modal data driven three-dimensional simulation scene construction method and device
CN121527325A