Water surface segmentation optimization method and system of neural image adaptive Transform based on stylized data driving, and computer readable storage medium

By cascading the NIAT module at the front end of the water surface segmentation model, the problem of degradation in water surface segmentation model under extreme weather conditions is solved, and efficient and accurate water surface segmentation in complex environments is achieved.

CN120047453APending Publication Date: 2025-05-27SHANDONG ZHIYANG SHANGSHUI INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510057426.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The performance of the water surface segmentation model in the prior art under extreme weather conditions has decreased, and the traditional methods are insufficient in adaptability to different weather conditions, lack flexibility and universality.

Method used

A neural image adaptive Transformer module (NIAT) based on stylized data-driven is designed. By combining data-driven methods and lightweight Transformer architecture, images in extreme weather are adaptively restored and optimized, and cascading at the front end of the water surface segmentation model.

Benefits of technology

The robustness and segmentation accuracy of the water surface segmentation model in complex environments such as foggy days, low light, overexposure and blur are significantly improved, and effective recovery from extreme conditional images to near normal visual states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047453A_ABST
    Figure CN120047453A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image data processing, and more specifically relates to a stylized data driven neural image adaptive Transform-based water surface segmentation optimization method and system, and a computer readable storage medium. The method comprises the steps of obtaining training set and test set image data, performing stylization processing on the training set image data by using a neural style migration technology, and keeping the original appearance of a test set image; the method comprises the following steps: designing an NIAT module, wherein the NIAT module comprises a front Rearrange layer, a front linear layer, z layers of serially connected Transform Blocks, a rear linear layer and a rear Rearrange layer; and cascading the NIAT module at the front end of the water surface segmentation model, training by using a training set image, and then testing by using a test set. According to the method, the problems that the performance of a segmentation technology is often greatly reduced in extreme weather and the segmentation technology is lack of flexibility and universality are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image data processing, and more specifically, relates to an optimization method, system, and computer-readable storage medium for water surface segmentation based on a stylized data-driven neural image adaptive Transformer. Background Art

[0002] In water environment monitoring and management, water surface segmentation is regarded as a key computer vision task, which involves accurately identifying and segmenting water bodies from complex backgrounds. Accurate water surface segmentation is of great significance for flood monitoring, water quality analysis, and ecological protection. However, blurring caused by extreme weather conditions such as fog, overexposure, low light, or signal transmission often seriously affects the quality of images, thereby affecting the performance of segmentation models.

[0003] Chinese invention patent CN117315240A discloses a cross-spectral image semantic segmentation method based on texture-independent features, including extracting texture-independent features, performing semantic segmentation based on texture-independent features, generating a large number of training data with the same structure but different textures through stylized images, explicitly punishing the network for learning texture-related information through a fully symmetric cumulative texture-independent loss, requiring the features extracted by the network to contain non-texture information sufficient to recover semantics through a reconstruction loss based on structural information, and biasing the model parameters better towards downstream tasks through a task-oriented fine-tuning method.

[0004] Although traditional segmentation techniques based on convolutional neural networks (CNNs) are effective in ordinary environments, their performance often drops significantly when facing challenges brought by extreme weather. This is mainly because these models usually only consider the image characteristics under normal weather conditions during design and are insufficiently adaptable to image changes under complex weather. In addition, existing technologies attempt to preprocess images through image enhancement techniques such as dehazing and brightness adjustment to improve segmentation accuracy. However, these methods not only increase the computational burden but also usually need to be adjusted for specific weather conditions, lacking flexibility and universality. Summary of the Invention

[0005] The present invention aims to overcome at least one defect of the above-mentioned prior art and provides an optimization method for water surface segmentation based on a stylized data-driven neural image adaptive Transformer, which is specifically used to optimize the water surface segmentation task.

[0006] The NIAT module designed in the present invention can effectively cascade with any segmentation model by combining data-driven methods and lightweight Transformer architectures, and adaptively restore and optimize images under extreme weather conditions. This module can not only learn the conversion from bad weather images to normal images, but also optimize the overall segmentation effect, providing a new and efficient solution for the water surface segmentation task under complex weather conditions.

[0007] The detailed technical solution of the present invention is as follows:

[0008] A method for optimizing water surface segmentation based on stylized data-driven neural image adaptive Transformer, the method comprising:

[0009] S1. First, obtain the training set and test set image data. The training set includes content images and style images, and the test set is test images. The training set is used to optimize the model parameters, and the test set is used to evaluate the performance of the model;

[0010] Then, perform stylization processing on the training set images through neural style transfer technology to simulate different extreme weather conditions and generate stylized training data;

[0011] The test set images remain unchanged and are used to evaluate the segmentation performance of the model under normal conditions and extreme weather conditions.

[0012] The neural style transfer technology uses a pre-trained model and combines two loss functions, content loss and style loss, for optimization. Among them, the content loss ensures that the generated image, that is, the synthesized image, is similar to the input content image in content, while the style loss makes the style of the generated image consistent with the input style image.

[0013] S2. To address the challenges of image processing under extreme weather conditions, a neural image adaptive Transformer module, namely the NIAT module (Neural Image Adaptive Transformer Module, NIAT), is designed;

[0014] The NIAT module adopts a lightweight Transformer architecture and is specifically designed for image restoration tasks under extreme weather conditions. The NIAT module includes a front Rearrange layer, a front linear layer, a Transformer Block connected in series with a z layer, a rear linear layer, and a rear Rearrange layer;

[0015] The front Rearrange layer: converts the input stylized training data or test set images into one-dimensional sequence data;

[0016] Front linear layer: Map the feature vector of each pixel point from 3D to 128D, that is, lift the low-dimensional features of the image to a high-dimensional feature space;

[0017] Transformer Block in series with the z layer: Each Transformer Block models the input features through the self-attention mechanism, captures the long-range dependencies between image features, and optimizes the image restoration effect layer by layer. The input and output dimensions of each Block remain the same;

[0018] Back linear layer: Remap the high-dimensional feature vector from 128D back to 3D to restore the original number of channels of the image;

[0019] Back Rearrange layer: Rearrange the processed sequence data into a two-dimensional image format and output the restored RGB image;

[0020] The advantage of the NIAT module lies in its adaptability and processing efficiency. Since the adopted network structure is relatively shallow and the traditional encoding-decoding strategy is not used, each Transformer Block is configured to output the same dimension as the input, ensuring the continuity of the information flow and the efficiency of processing. As a data-driven system, the performance and effect of the NIAT module depend on the quality and diversity of the provided training data. Through training, the module can learn and adapt to the image features under various extreme weather conditions, and achieve an effective restoration from extreme-condition images to a state close to normal vision.

[0021] S3. Cascade the NIAT module at the front end of the water surface segmentation model, train it using the training set images, and then test it through the test set.

[0022] Furthermore, the style processing of the training set images through the neural style transfer technology specifically includes:

[0023] S11. Input the training set image data, and respectively extract the high-level features of the content image and the low-level features of the style image through the pre-trained model; the high-level features are used to describe the structural information of the image, and the low-level features are used to describe the texture and style information;

[0024] S12. Initialize a copy of the content image as the initial image, and optimize the initial image through backpropagation and gradient descent to gradually minimize the content loss and the style loss, and obtain the synthesized image:

[0025] First, use the mean square error to measure the content loss L content :

[0026]

[0027] F i,jis the pixel value of the output feature map of the convolutional neural network, P i,j is the pixel value of the corresponding target content image, F i,j represents the feature value at the i-th row and j-th column of the feature map output by the generated image on this convolutional layer, P i,j represents the corresponding feature value output by the input target content image on the same convolutional layer, ensuring that the generated image is consistent with the original image in content.

[0028] Secondly, the style similarity between the generated image and the style image in the pre-trained model is measured by the mean square error of the Gram matrix:

[0029]

[0030] In formula (2), is the Gram matrix of the generated image, is the Gram matrix of the style image, N l is the number of channels of the l-th convolutional layer, and M l is the total number of elements in the feature map of this layer.

[0031] Furthermore, select style images of 4 styles: foggy day, low light, overexposure, and blur, with n images for each style, and adjust m different stylization intensities;

[0032] As the stylization intensity increases, the details of the content image are lost and the embedded style becomes extreme, generating n×m×4 stylized training data.

[0033] Furthermore, after cascading the NIAT module at the front end of the water surface segmentation model, it also includes optimization using a loss function:

[0034] S31. Configure a joint training framework, cascade the NIAT module at the front end of the water surface segmentation model, and co-train the NIAT module and the water surface segmentation model;

[0035] S32. Configure a loss function, and the total loss function L total includes an image restoration loss L restore and a segmentation loss L segment :

[0036] L total =α·L segment +β·L res tore (3);

[0037] In formula (3), α and β are parameters for adjusting the weights of the restoration loss and the segmentation loss. The segmentation loss is the loss of the water surface segmentation model, and the restoration loss is represented by the structural similarity loss. The structural similarity measures the similarity between the generated image and the content image in terms of the visual perception characteristics of the image, comprehensively considering three aspects: brightness, contrast, and structure, and is defined as follows:

[0038]

[0039] In formula (4), x and y are the generated image and the content image, and μ x and μ y are the average values of x and y respectively, and are the variances of x and y respectively, σ xy is the covariance of x and y, and c 1 and c 2 are constants added to avoid the denominator being zero;

[0040] S33. According to the overall loss function, a multi-task learning strategy is adopted to optimize the NIAT and the segmentation model during the training process.

[0041] Through tests under various extreme weather conditions, the robustness and accuracy of the NIAT module cascaded water surface segmentation model are verified. The test results show that the model can maintain a high level of segmentation performance in complex environments such as foggy days, low light, overexposure, and blurring. In practical applications, the model demonstrates good real-time processing ability and stability, and is suitable for water environment monitoring scenarios such as drone monitoring and fixed camera systems, ensuring efficient and reliable water surface segmentation results.

[0042] In another aspect of the present invention, a system for optimizing a water surface segmentation method based on a stylized data-driven neural image adaptive Transformer is provided. The system includes:

[0043] At least one processor; and

[0044] A memory that stores instructions. When the instructions are executed by the at least one processor, the at least one processor executes a method for optimizing a water surface segmentation based on a stylized data-driven neural image adaptive Transformer as described above.

[0045] In another aspect of the present invention, a computer-readable storage medium is also provided, which stores executable instructions. When the instructions are executed, the machine executes a method for optimizing a water surface segmentation based on a stylized data-driven neural image adaptive Transformer as described above.

[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0047] (1) The present invention provides an optimized method, system and computer-readable storage medium for water surface segmentation based on a stylized data-driven neural image adaptive Transformer, which enhances the model's image processing ability in extreme weather conditions through stylized training data. The NIAT module significantly improves the robustness of the water surface segmentation model and the accuracy of the segmentation model in complex environments such as foggy days, low light, overexposure and blurring.

[0048] (2) The present invention provides an optimized method, system and computer-readable storage medium for water surface segmentation based on a stylized data-driven neural image adaptive Transformer, which seamlessly integrates the image restoration and segmentation processes in the same model framework, realizing a highly integrated end-to-end solution without the need for complex image preprocessing for different extreme weather conditions, such as defogging, brightness adjustment, etc., to improve the segmentation accuracy; this integrated design reduces the dependence on external preprocessing steps of the model, improves the robustness and flexibility of the system, especially when facing multiple extreme weather conditions, it can still maintain relatively stable segmentation performance; it also simplifies the design and maintenance of the system, making it easier to promote and expand in different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a schematic flowchart of the method of the present invention.

[0050] Figure 2 is a schematic diagram of the model architecture of the present invention.

[0051] Figure 3 is a schematic diagram of the original image in the embodiment.

[0052] Figure 4 is a stylized image of the original image in Embodiment 1 under foggy style and with a stylized intensity of 0.8.

[0053] Figure 5 is a segmentation effect diagram of the foggy stylized image in Embodiment 1 passing through the NIAT module cascaded with the water surface segmentation model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The present invention will be further described below in conjunction with the drawings and embodiments.

[0055] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0056] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0057] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0058] Embodiment 1

[0059] Refer Figure 1 , this embodiment provides an optimization method for water surface segmentation based on a stylized data-driven neural image adaptive Transformer. The method includes:

[0060] S1. Obtain the training set and test set image data, and perform stylization processing on the training set image data;

[0061] First, select the standard training set and test set image data. The training set includes content images and style images, and the test set is test images. Among them, the training set is used for optimizing the model parameters, and the test set is used for evaluating the performance of the model.

[0062] Then, perform stylization processing on the training set images through neural style transfer technology to simulate different extreme weather conditions (such as foggy days, low light, overexposure, and blurring) to generate stylized training data. The test set images remain unchanged and are used for evaluating the segmentation performance of the model under normal conditions and extreme weather conditions.

[0063] Neural Style Transfer (NST) is a deep learning technology that combines the structural information of a content image and the texture and color styles of a style image to generate a synthetic image that retains the main structure of the content image and presents the artistic style of the style image.

[0064] Specifically, this embodiment selects the ImageNet dataset, and the pre-trained model is the VGG-19 model.

[0065] Preferably, the stylization processing of the training set images through neural style transfer technology specifically includes:

[0066] S11. Input the training set image data, and respectively extract the high-level features of the content image and the low-level features of the style image through a pre-trained model (such as VGG-19); the high-level features are used to describe the structural information of the image, and the low-level features are used to describe the texture and style information;

[0067] S12. Initialize a copy of the content image as the initial image, as Figure 3 shown; optimize the initial image through backpropagation and gradient descent to gradually minimize the content loss and style loss, and obtain the synthesized image, that is, the stylized image. As Figure 4 shown;

[0068] Preferably, in this embodiment, the mean square error is used to measure the content loss, which is expressed by the following formula:

[0069]

[0070] F i,j is the pixel value of the output feature map of the convolutional neural network, and P i,j is the pixel value of the corresponding target content image. F i,j represents the feature value at the i-th row and j-th column of the feature map output by the generated image on this convolutional layer, and P i,j represents the corresponding feature value output by the input target content image on the same convolutional layer. This method ensures that the generated image is consistent with the original image in content.

[0071] The style loss uses the mean square error of the Gram matrix, that is, calculates the mean square error between the Gram matrix of the generated image and the Gram matrix of the style image at the l-th layer, so as to measure the difference in style. The formula is as follows:

[0072]

[0073] Among them, is the Gram matrix of the generated image, is the Gram matrix of the style image, N l is the number of channels of the l-th convolutional layer, and M l is the total number of elements of the feature map of this layer.

[0074] Preferably, in this embodiment, 4 types of style images including foggy days, low light, overexposure, and blur are selected. 5 images are selected for each type, and 3 different stylization intensities are adjusted respectively to generate 60 stylized training data. As Figure 4 shown, it is the foggy day stylized image when the stylization intensity is 0.8.

[0075] In order to comprehensively capture and transfer the artistic style, features are extracted from the following multiple convolutional layers to calculate the style loss:

[0076] conv1_1: Capture the most basic texture and color information;

[0077] conv2_1: Provide slightly more complex texture details;

[0078] conv3_1: Capture medium-complexity texture patterns;

[0079] conv4_1: Suitable for capturing more complex style features and texture hierarchies;

[0080] conv5_1: Reflect high-level style features, including complex textures and shapes.

[0081] S2. Design the NIAT module to achieve the restoration of images under extreme conditions;

[0082] To address the challenges of image processing under extreme weather conditions, this embodiment designs a neural image adaptive Transformer module, namely the NIAT module (Neural Image Adaptive Transformer Module, NIAT).

[0083] The NIAT module adopts a lightweight Transformer architecture, specifically designed for the image restoration task under extreme weather conditions, including:

[0084] The front Rearrange layer (image-to-sequence conversion layer): Convert the input stylized training data or test set images ([3, 640, 640], i.e., 3-channel RGB images) into one-dimensional sequence data ([409,600, 3], i.e., three-dimensional feature vectors of 640×640 pixel points), providing the input format for subsequent Transformer structure processing.

[0085] The front linear layer (QLinear Layer): Map the feature vectors of each pixel point from 3 dimensions to 128 dimensions, aiming to lift the low-dimensional features of the image to a high-dimensional feature space, facilitating the Transformer to capture more image details.

[0086] The Transformer Block in series with z layers: Each Transformer Block models the input features through the self-attention mechanism, capturing the long-range dependencies between image features. Through the structure in series with z layers, the model optimizes the image restoration effect layer by layer. The input and output dimensions of each Block remain consistent ([409,600, 128]) to ensure the continuity of the information flow. Preferably, z = 5.

[0087] Post Linear Layer (HLinear Layer): Remaps the high-dimensional feature vector from 128 dimensions back to 3 dimensions to restore the original number of channels of the image.

[0088] Post Rearrange Layer (sequence-to-image conversion layer): Rearranges the processed sequence data ([409600, 3]) into a two-dimensional image format ([3, 640, 640]) and outputs the restored RGB image.

[0089] The design of the NIAT module emphasizes its adaptive ability and processing efficiency. Since the adopted network structure is relatively shallow and the traditional encoding-decoding strategy is not used, each Transformer Block is configured to output the same dimension as the input, ensuring the continuity of the information flow and the efficiency of processing. As a data-driven system, the performance and effect of the NIAT module depend on the quality and diversity of the provided training data. Through training, the NIAT module can learn and adapt to the image features under various extreme weather conditions, achieving an effective restoration from extreme-condition images to a state close to normal vision.

[0090] S3. Cascade the NIAT module at the front end of the water surface segmentation model for optimized training.

[0091] Specifically, the style processing of the training set images through neural style transfer technology specifically includes:

[0092] S31. Configure a joint training framework, cascade the NIAT module at the front end of the water surface segmentation model, as shown in the model architecture schematic diagram, so that the NIAT module and the water surface segmentation model are jointly trained; preferably, the water surface segmentation model is the deeplabv3+ model. Figure 2

[0093] S32. Configure the loss function. The total loss function L total includes the image restoration loss L restore and the segmentation loss L segment :

[0094] L total = α·L segment + β·L res tore (3);

[0095] In formula (3), α and β are parameters for adjusting the weights of the restoration loss and the segmentation loss;

[0096] The segmentation loss is the loss of the water surface segmentation model, representing the loss between the label image (the image marking the water surface range) and the image segmented by the water surface segmentation model;

[0097] ​The restoration of losses is represented using the Structural Similarity Index Measure (SSIM). The structural similarity measures the similarity between the generated image and the content image, or the similarity of the same-position regions of the generated image and the content image, by comprehensively considering three aspects: luminance, contrast, and structure, according to the visual perception characteristics of images. Its definition is as follows:

[0098]

[0099] In Equation (4), x and y are the generated image and the content image, and μ x and μ y are the average values of x and y respectively. and are the variances of x and y respectively, σ xy is the covariance of x and y, and c 1 and c 2 are constants added to avoid a zero denominator.

[0100] According to the total loss function, a multi-task learning strategy is adopted to optimize the NIAT module and the water surface segmentation model during the training process.

[0101] In this embodiment, the foggy images simulated by neural style transfer are used for testing. The result is the image after water surface segmentation by the NIAT module cascaded with the water surface segmentation model, as Figure 5 shown, and the model shows good water surface segmentation results.

[0102] Furthermore, through testing under various extreme weather conditions, the robustness and accuracy of the NIAT module cascaded with the water surface segmentation model are verified; the test results show that this model can maintain a high level of segmentation performance in complex environments such as foggy days, low light, overexposure, and blurring. In practical applications, it shows good real-time processing ability and stability, and is suitable for water environment monitoring scenarios such as unmanned aerial vehicle monitoring and fixed camera systems, ensuring efficient and reliable water surface segmentation results. The effects are presented as shown in Table 1:

[0103] Table 1: Comparison of water surface segmentation accuracy between the present invention and the prior art under different weather conditions

[0104]

[0105] Through testing under various extreme weather conditions such as foggy days, low light, overexposure, and blurring, it can be seen that the water surface segmentation accuracy of the present invention under various extreme weather conditions is 8.9% - 20.3% higher than that of the prior art.

[0106] Example 2

[0107] This embodiment provides a system for implementing an optimization method for water surface segmentation based on a stylized data-driven neural image adaptive Transformer:

[0108] at least one processor; and

[0109] a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform an optimization method for water surface segmentation of a stylized data-driven neural image adaptive Transformer as described above.

[0110] In this embodiment, the electronic device includes but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smartphones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, etc.

[0111] Embodiment 3

[0112] This embodiment also provides a computer-readable storage medium storing executable instructions that, when executed, cause the machine to perform an optimization method for water surface segmentation of a stylized data-driven neural image adaptive Transformer as described above.

[0113] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system or device is caused to read and execute the instructions stored in the readable storage medium.

[0114] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.

[0115] Examples of the readable storage medium include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or a cloud via a communication network.

[0116] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0117] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0118] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0120] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A water surface segmentation optimization method based on stylized data-driven neural image adaptive Transformer, characterized in that: The method comprises: S1. Obtain training set and test set image data, wherein the training set includes content images and style images, and the test set is test images, and use neural style transfer technology to stylize the training set image data, while the test set images remain original; S2, designing a neural image adaptive Transformer module, namely, a NIAT module, wherein the NIAT module includes a Transformer Block of a front Rearrange layer, a front linear layer, a z layer connected in series, a rear linear layer, and a rear Rearrange layer; Front Rearrange layer: converts the input stylized training data or test set images into one-dimensional sequence data; Front linear layer: maps the feature vector of each pixel from 3 dimensions to 128 dimensions, that is, elevates the low-dimensional features of the image to a high-dimensional feature space; z-layer series of Transformer Blocks: Each Transformer Block models the input features through the self-attention mechanism, captures the long-range dependencies between image features, optimizes the image restoration effect layer by layer, and keeps the input and output dimensions of each Block consistent; Post-linear layer: remaps the high-dimensional feature vector from 128 dimensions back to 3 dimensions and restores the original number of channels of the image; Post-Rearrange layer: rearranges the processed sequence data into a two-dimensional image format and outputs the restored RGB image; S3. Cascade the NIAT module to the front end of the water surface segmentation model, train it with the training set images, and then test it with the test set.

2. According to claim 1, a water surface segmentation optimization method based on stylized data driven neural image adaptive Transformer is characterized in that: The stylized processing of the training set image data using the neural style transfer technology includes: S11, input training set image data, and extract high-level features of the content image and low-level features of the style image through the pre-training model; the high-level features are used to describe the structural information of the image, and the low-level features are used to describe the texture and style information; S12. Initialize a copy of the content image as the initial image, optimize the initial image through back propagation and gradient descent, so that it gradually minimizes the content loss and style loss, and obtains the synthetic image: First, the content loss L is measured using the mean square error content : In formula (1), F i,j is the pixel value of the output feature map of the convolutional neural network, P i,j is the pixel value of the corresponding target content image, F i,j represents the feature value of the i-th row and j-th column of the feature map output by the generated image on the convolutional layer, P i,j Represents the corresponding feature value output by the input target content image on the same convolutional layer; Secondly, the style similarity between the generated image and the style image in the pre-trained model is measured by the mean square error of the Gram matrix: In formula (2), is the Gram matrix of the generated image, is the Gram matrix of the style image, N l is the number of channels of the lth convolutional layer, M l is the total number of feature map elements in layer l.

3. The water surface segmentation optimization method based on stylized data driven neural image adaptive Transformer according to claim 1 or 2, characterized in that: Select four styles of images: foggy, weak light, overexposed, and blurred, n images of each style, and adjust m different stylization intensities; As the stylization intensity increases, the details of the content image are lost and the embedded style becomes extreme, generating n×m×4 stylized training data.

4. The water surface segmentation optimization method based on stylized data driven neural image adaptive Transformer according to claim 1 or 2, characterized in that: The cascading of the NIAT module to the front end of the water surface segmentation model also includes optimizing using a loss function: S31, configure a joint training framework, cascade the NIAT module to the front end of the water surface segmentation model, so that the NIAT module and the water surface segmentation model are trained together; S32, configure the loss function, the total loss function L total Including image restoration loss L restore and segmentation loss L segment : THE total =α·L segment +β·L res bull (3); In formula (3), α and β are parameters for adjusting the weights of restoration loss and segmentation loss; Segmentation loss is the loss of the water surface segmentation model; The restoration loss is expressed as a structural similarity loss. Structural similarity measures the similarity between the generated image and the content image based on the visual perception characteristics of the image, combining brightness, contrast, and structure: In formula (4), x and y are the generated image and content image, μ x and μ y are the means of x and y respectively, and are the variances of x and y, σ xy is the covariance of x and y, c1 and c2 are constants added to avoid the denominator being zero; S33. Based on the overall loss function, a multi-task learning strategy is used to optimize the NIAT and segmentation models during the training process.

5. A water surface segmentation optimization system based on a stylized data-driven neural image adaptive Transformer, characterized in that: The system comprises: processor; a memory having stored thereon a computer program executable on the processor; Wherein, when the computer program is executed by the processor, the steps of a water surface segmentation optimization method based on a stylized data-driven neural image adaptive Transformer are implemented as described in any one of claims 1 to 4.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Cross-spectrum image semantic segmentation method based on texture irrelevant features

    CN117315240A