Sea-land target multi-modal image segmentation method, system, device, medium and product

By combining Swin Transformer and deep-separable convolutional network, a regularized Hu moment loss function is introduced, and the multimodal target segmentation model is optimized, which solves the problems of boundary blur and detail loss, and improves the fineness and robustness of sea and land target segmentation.

CN120495315APending Publication Date: 2025-08-15NAVAL AVIATION UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510629548.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the existing deep learning model, in dealing with the segmentation task of ship and aircraft targets, there are problems of blurred boundaries and loss of details, which affect the segmentation accuracy and robustness.

Method used

Swin Transformer's feature extraction network and deep separable convolutional network are combined to build a multimodal target segmentation model, and a regularized Hu moment loss function is introduced. Through the combination of mean square variance loss and regularized Hu moment loss function, the model training process is optimized, and boundary capture ability and shape feature learning are enhanced.

Benefits of technology

Effectively reduce boundary blurring, retain the small structure of the target, improve the fineness and robustness of the segmented images, especially in small-objective segmentation tasks in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495315A_ABST
    Figure CN120495315A_ABST
Patent Text Reader

Abstract

The invention discloses a sea-land target multi-modal image segmentation method, system and device, a medium and a product, and relates to the field of remote sensing image analysis, and the method comprises the steps: carrying out the preprocessing of an obtained target image data set, and then according to the preprocessed data set, a Swin Transform-based feature extraction network and a depth separable convolutional network, carrying out the segmentation of a multi-modal image of a sea-land target, and carrying out the segmentation of the multi-modal image of the sea-land target. Determining a trained multi-modal target segmentation model, and segmenting the to-be-segmented multi-modal image pair by using the trained multi-modal target segmentation model; wherein the training process of the multi-modal target segmentation model comprises the steps of determining an overall loss function by utilizing a mean square error loss function and a regularization Hu moment loss function according to a real mask and a prediction mask, updating the multi-modal target segmentation model, and dynamically adjusting adaptive weight parameters in the overall loss function; according to the invention, the boundary blur phenomenon can be reduced, the fine structure of the target is reserved, and the fineness of the segmented image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of remote sensing image analysis, and in particular to a method, system, equipment, medium and product for multimodal image segmentation of land and sea targets. Background Art

[0002] In the field of computer vision and remote sensing image analysis, the segmentation of ships and aircraft targets is a critical task, which is of great significance for applications such as traffic management and safety. Traditional edge detection and threshold segmentation methods often lack generalization and adaptability when faced with complex and changing environments. In recent years, deep learning methods, especially convolutional neural networks (CNNs), have achieved remarkable success in image segmentation tasks due to their powerful feature extraction capabilities and end-to-end learning framework. However, existing deep learning models mainly rely on pixel-level loss functions, such as cross-entropy loss function or Dice loss function. These loss functions may lead to blurred boundaries or loss of details when dealing with objects with complex shapes, which restricts further improvement of algorithm performance.

[0003] To address this issue, several algorithms have been studied. Some have introduced edge-aware loss functions to enhance the model's sensitivity to target edges. This approach increases the weight of edge pixels when calculating the loss, or designs a specialized edge detection module that integrates with the segmentation network to more accurately capture target boundary information. Other studies have employed multi-scale feature fusion strategies, integrating feature maps at different scales to enhance the model's understanding of object details and overall structure, thereby alleviating the problem of detail loss to a certain extent. Segmentation methods based on attention mechanisms have also been proposed. By dynamically adjusting the level of attention paid to different image regions, the model can focus more closely on the target area, improving segmentation accuracy and boundary clarity. These improved methods have significantly improved the segmentation of complex targets such as ships and aircraft, providing new ideas and directions for future research. However, these methods do not explicitly learn refined edge features, which hinders further improvement in the performance of these algorithms.

[0004] Therefore, there is an urgent need to provide a new multimodal image segmentation method or system for sea and land targets, which can effectively reduce boundary blur, retain the fine structure of the target, and improve the fineness of the segmented image. Summary of the Invention

[0005] The purpose of this application is to provide a multimodal image segmentation method, system, equipment, medium and product for sea and land targets, which can reduce boundary blur, retain the fine structure of the target, and improve the fineness of the segmented image.

[0006] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides a method for multimodal image segmentation of land and sea targets, the method comprising: Obtain a target image dataset; the target image dataset includes: a training set and a test set; Preprocessing the target image dataset to obtain a preprocessed dataset; the preprocessed dataset includes: a multimodal image pair and a corresponding true mask; the multimodal image pair includes: an infrared image and a corresponding visible light image; the preprocessing includes: image cropping and mask annotation; According to the preprocessed data set, a trained multimodal target segmentation model is determined based on a Swin Transformer feature extraction network and a depthwise separable convolutional network; the multimodal target segmentation model takes a multimodal image pair as input and a predicted mask as output; the training process of the multimodal target segmentation model specifically includes: determining an overall loss function based on the true mask and the predicted mask using a mean square error loss function and a regularized Hu moment loss function; and training the multimodal target segmentation model using the overall loss function to obtain a trained multimodal target segmentation model; The trained multimodal object segmentation model is used to segment the multimodal image pairs to be segmented.

[0007] Optionally, obtaining a target image dataset specifically includes: The target image dataset is divided into a ratio of 7:3 to determine the training set and test set.

[0008] Optionally, determining an overall loss function based on the true mask and the predicted mask using a mean square error loss function and a regularized Hu moment loss function specifically includes: Determining a pixel-level loss between the true mask and the predicted mask using a mean square error loss function; Extracting contour point sequences of the true mask and the predicted mask respectively, and determining corresponding true contours and predicted contours; Determining a corresponding Hu moment according to the true contour and the predicted contour; According to the Hu moment corresponding to the true contour and the Hu moment corresponding to the predicted contour, the corresponding regularized Hu moment is determined respectively; Determining a regularized Hu moment loss function according to the regularized Hu moment corresponding to the true contour and the regularized Hu moment corresponding to the predicted contour; and determining a regularized Hu moment loss using the regularized Hu moment loss function; An overall loss function is determined based on the pixel-level loss and the regularized Hu moment loss.

[0009] Optionally, determining the regularized Hu moment loss by using the regularized Hu moment loss function specifically includes: Using the formula Determine the regularized Hu moment loss ;in, K is the number of moments used, To seek peace, k is the number of the regularized Hu moment used, The first k A regularized Hu moment, The true contour k A regularized Hu moment.

[0010] Optionally, determining an overall loss function according to the pixel-level loss and the regularized Hu moment loss specifically includes: Using the formula Determine the overall loss function ;in, is the pixel-level loss, is the regularized Hu moment loss, is an adaptive weight parameter used to balance the importance of the pixel-level loss and the regularized Hu moment loss.

[0011] Optionally, the training of the multimodal object segmentation model using the overall loss function to obtain a trained multimodal object segmentation model specifically includes: According to the overall loss function, the multimodal target segmentation model is updated through the back propagation algorithm and the Adam optimizer; and the adaptive weight parameters in the overall loss function are dynamically adjusted; until the number of updates is reached, the trained multimodal target segmentation model is obtained; the dynamic adjustment is to use the formula =0.2+e / 500*0.6 Determine adaptive weight parameters , increase the weight of the regularized Hu moment loss; where, e is the number of updates.

[0012] In a second aspect, the present application provides a multimodal image segmentation system for land and sea targets, the multimodal image segmentation system for land and sea targets comprising: A target image data set acquisition module is used to acquire a target image data set; the target image data set includes: a training set and a test set; A preprocessing module, configured to preprocess the target image dataset to obtain a preprocessed dataset; the preprocessed dataset includes: a multimodal image pair and a corresponding true mask; the multimodal image pair includes: an infrared image and a corresponding visible light image; the preprocessing includes: image cropping and mask annotation; A multimodal target segmentation model establishment module is used to determine a trained multimodal target segmentation model based on the preprocessed data set, a feature extraction network based on SwinTransformer, and a depth-wise separable convolutional network; the multimodal target segmentation model takes a multimodal image pair as input and a predicted mask as output; the training process of the multimodal target segmentation model specifically includes: determining an overall loss function based on the true mask and the predicted mask using a mean square error loss function and a regularized Hu moment loss function; and training the multimodal target segmentation model using the overall loss function to obtain a trained multimodal target segmentation model; The image segmentation module is used to segment the multimodal image pairs to be segmented using the trained multimodal target segmentation model.

[0013] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above-mentioned methods for multimodal image segmentation of land and sea targets.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-mentioned methods for multimodal image segmentation of land and sea targets.

[0015] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned methods for multimodal image segmentation of land and sea targets.

[0016] According to the specific embodiments provided in this application, this application has the following technical effects: The present application provides a method, system, device, medium and product for multimodal image segmentation of land and sea targets. By introducing regularized Hu moments into the loss function, a new loss function based on regularized Hu moments (Regularized HuMoments-based Loss, RHML) is determined, and a multimodal target segmentation model is trained. The low-order moments in the Hu moments can align the segmentation mask and the prediction mask from an overall macro perspective, and the high-order moments can enhance the ability to capture the details of the target boundary. The regularized Hu moments can not only effectively reduce the boundary blur phenomenon, but also better preserve the fine structure of the target, thereby further improving the segmentation fineness and improving the image segmentation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 Schematic diagram of a flow chart of a multimodal image segmentation method for sea and land targets in one embodiment of the present application; Figure 2 This is a schematic diagram of the segmentation process of the multimodal target segmentation model in one embodiment of the present application; Figure 3 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0021] In an exemplary embodiment, Figure 1 As shown, a multimodal image segmentation method for land and sea targets is provided, and the multimodal image segmentation method for land and sea targets includes the following S1-S4. S1: Obtain the target image dataset.

[0022] Satellites, drones, or ground-based cameras are used to collect infrared and visible light images of ships or aircraft. This dataset is then divided into a training set and a test set in a 7:3 ratio. The training set is used to train and fine-tune the multimodal object segmentation model, while the test set is used to evaluate its generalization and segmentation accuracy.

[0023] S2: Preprocessing the target image dataset to obtain a preprocessed dataset.

[0024] Preprocessing the multimodal image pairs in the target image dataset; the preprocessing includes: image cropping and mask annotation; Specifically, all multimodal image pairs were cropped to a standard size of 512×512, ensuring that the pixel values of each channel were between 0 and 1. This yielded a multimodal image pair consisting of an infrared image and a corresponding visible light image. The multimodal image pairs were manually annotated using an image annotation tool. The specific locations of the aircraft or ship targets in the infrared and visible light images of the multimodal image pair were annotated, and a corresponding ground-truth mask was generated. Pixels in the image that belonged to the target were labeled as 1 in the mask, while pixels that did not belong to the target were labeled as 0 in the mask.

[0025] S3: According to the preprocessed data set, a trained multimodal object segmentation model is determined based on the feature extraction network of Swin Transformer and the depthwise separable convolutional network.

[0026] S3 specifically includes: S31: Establish a multimodal object segmentation model.

[0027] The input of the multimodal object segmentation model is a multimodal image pair, and the output is a predicted mask. In order to enable the multimodal object segmentation model to learn features more fully, a feature extraction network based on Swin Transformer (Shifted Windows Transformer) is designed, and the multimodal object segmentation model is lightweighted using a deep separable convolutional network. The flow chart of the multimodal object segmentation model is shown below. Figure 2 As shown in the figure, after the visible light image and the infrared image pass through a depthwise separable convolution DwSConv (Depthwise Separable Convolution), a 256*256*64 feature map is obtained, which is then fused and spliced. After passing through two Swin Transformer layers and a depthwise separable convolution layer, a 32*32*128 feature map is obtained. After four deconvolution operations, each deconvolution operation includes a deconvolution layer Deconv, a deformable convolution layer DfConv (Deformable Convolution) and a convolution layer Conv, and finally a 512*512*1 output is obtained. After passing through an activation function Sigmoid, the predicted mask is finally obtained.

[0028] Visible light images provide rich texture and color information, while infrared images are insensitive to lighting changes and can provide effective information in low light or complex backgrounds. By fusing visible light and infrared images, this application can simultaneously utilize the complementary information of the two modalities, enhancing the robustness of the multimodal object segmentation model in different environments. In addition, by using depthwise separable convolution to perform preliminary feature extraction on the two modal images before fusion, the computational effort is reduced while retaining important feature information. Swin Transformer can capture global and local features at different scales through the local window self-attention mechanism, which is suitable for image segmentation tasks. Compared with traditional convolutional neural networks, Swin Transformer performs better in processing long-distance dependencies and can more effectively capture the contextual information of the target. After two Swin Transformer layers and one depthwise separable convolution layer, it can gradually extract more abstract features from shallow to deep layers, and finally generate a 32*32*128 feature map. The above feature map contains rich semantic information, which is helpful for subsequent segmentation tasks. Depthwise separable convolution greatly reduces the number of parameters and computation by decomposing the standard convolution into depthwise convolution and pointwise convolution, making the network more lightweight and suitable for deployment on resource-constrained devices. Deformable convolution can adaptively adjust the shape of the convolution kernel to better adapt to the geometric transformation of the target and enhance the robustness of the model to target deformation. Through depthwise separable convolution and deformable convolution, the model's ability to learn features is enhanced while reducing the number of model parameters. Finally, through multimodal feature fusion, Swin Transformer's global feature extraction and deconvolution detail recovery enable the network to achieve high-precision target segmentation in complex scenarios.

[0029] S32: Use the mean square error loss function and the regularized Hu moment loss function to train the multimodal target segmentation model to obtain a trained multimodal target segmentation model.

[0030] S32 specifically includes: S321: Construct the mean squared error (MSE) loss function to calculate the pixel-level loss.

[0031] The input of the mean square error loss function is the true mask and the predicted mask, and the output is the MSE loss value, that is, the pixel-level loss. The mean square error loss function is a commonly used pixel-level loss function that measures the difference between the predicted value and the true value. Its calculation formula is as follows: .

[0032] in, is the pixel-level loss, is the true mask label, To predict the mask label, N is the total number of pixels of the mask image, j is the pixel number of the mask image.

[0033] S322: Construct a loss function of the regularized Hu moment to calculate the regularized Hu moment loss.

[0034] The mean square error loss function effectively measures the pixel-level difference between the predicted mask and the true mask, ensuring that the predictions generated by the multimodal object segmentation model are as close to the true value as possible. However, the mean square error loss function only focuses on pixel-level differences and cannot capture the shape characteristics of the object, which can easily lead to blurred boundaries or loss of details.

[0035] To this end, a loss function based on regularized Hu moments is constructed. This loss function combines shape features and pixel-level similarity to improve segmentation accuracy. Hu moments are a set of descriptors based on geometric moments. By integrating the grayscale values of an image, a set of coefficients reflecting the object shape is obtained. Hu moments are invariant to translation, rotation, and scaling, and can effectively describe the shape characteristics of complex objects.

[0036] The input of the regularized Hu moment loss function is the real mask and the predicted mask, and the output is the regularized Hu moment loss. The calculation process is as follows: The contour point sequences of the real mask and the predicted mask are extracted respectively to obtain the corresponding contours.

[0037] Contour points are pixels in the real mask or predicted mask that are 1 and none of the 8 adjacent pixels are all 1. Starting from the top leftmost point, extract contour points in clockwise order to form a closed polygon; record these points as a sequence , ,… , where each is a two-dimensional coordinate, is the number of points in the contour point sequence, is the sequence number. For specific operations, you can use the findContours function of OpenCV.

[0038] For the extracted true contour and predicted contour, their regularized Hu moments are calculated respectively; for each contour, the regularized Hu moments of order 0 to 7 are calculated according to the following steps: 1. Calculate the original moment , the calculation formula is as follows: .

[0039] in, is the two-dimensional coordinates of a contour point sequence, is the number of points in the contour point sequence, For beg Power, For beg power, where and The values are all the order of the moment, ranging from 0, 1, ..., 7. The values are derived from the order of the moment, and the values are different when calculating Hu moments of different orders.

[0040] 2. Calculate the central moment , the calculation formula is as follows: .

[0041] in, , , and Represents the value in the brackets and Power.

[0042] 3. Normalize the central moment to eliminate the influence of scale and normalize the central moment The calculation formula is: .

[0043] in, for the reason and The regulatory factor that determines for and The central moment value when both are 0 can be calculated by step 2, representing of Power, , for and Determine the central moment value.

[0044] 4. Calculate the Hu moment. The Hu moment is based on the 7 invariant moments of the normalized central moment. The calculation formulas are as follows: .

[0045] .

[0046] .

[0047] .

[0048] .

[0049] .

[0050] .

[0051] Among them, according to The corner mark can be obtained and The value of, such as the subscript 20 represents , =0 ,Right now The square of 0th power.

[0052] 5. Calculate the regularized Hu moment; in order to enhance the stability of the Hu moment, take the logarithm of the Hu moment and take the absolute value. The calculation formula is as follows: .

[0053] in, is the order number, =1, 2, ..., 7.

[0054] The above operations are performed on the contour line sequence of the real mask and the corresponding predicted mask of the input image, that is, the contour point sequences of the real mask and the predicted mask are extracted respectively, and the corresponding real contour and predicted contour are determined; the corresponding Hu moment is determined according to the real contour and the predicted contour; the corresponding regularized Hu moment is determined according to the Hu moment corresponding to the real contour and the Hu moment corresponding to the predicted contour; the first k The regularized Hu moment is denoted as , the predicted contour k The regularized Hu moment is denoted as , k is the number of the regularized Hu moment.

[0055] 6. Calculate the regularized Hu moment loss In order to measure the shape difference between two contours, the Euclidean distance between the regularized Hu moments is defined as the loss function. The specific calculation formula is as follows: .

[0056] in, K Indicates the number of moments used (7 is used here, and the specific value can be determined according to the experimental results of different data sets). k is the regularized Hu moment number.

[0057] S323: Construct a combined loss function. By weighted summation, the regularized Hu moment loss and pixel-level loss are combined to obtain the overall loss function. , the calculation formula is as follows: .

[0058] in, is an adaptive weight parameter used to balance the importance of the two losses. The model can flexibly weigh the importance of the two losses during training, thereby improving training efficiency and the quality of segmentation results.

[0059] S324: Using the overall loss function to train the multimodal target segmentation model to obtain a trained multimodal target segmentation model.

[0060] The initial learning rate was set to 0.0002, and cosine annealing was used to gradually reduce the learning rate to accelerate convergence and prevent overfitting. The batch size was set to 64, meaning 64 images were input for each training run. The number of training runs was set to 500, with model parameters updated once per iteration. The Adam optimizer was used.

[0061] The multimodal object segmentation model is trained according to the following process: 1. Input image pairs: Input 64 aircraft target image pairs into the multimodal target segmentation model to obtain the segmentation prediction results of each image .

[0062] 2. Calculating Loss: Using Segmentation Prediction Results And the true mask, respectively calculate the pixel-level loss and regularized Hu moment loss, and then determine the overall loss function, and then back-transfer to optimize the multimodal target segmentation model.

[0063] 3. Back propagation: The back propagation algorithm is used to calculate the gradient of the overall loss function with respect to the feature extraction network parameters constructed by S31, and the Adam optimizer is used to update the multimodal object segmentation model.

[0064] 4. Dynamically adjust weights: During training, dynamically adjust the weights of regularized Hu moment loss and pixel-level loss For the e Training, setting =0.2+e / 500*0.6 , to gradually increase the weight of the regularized Hu moment loss, so that the model pays more attention to the alignment of shape features.

[0065] Repeat the above process 500 times to finally obtain the trained model.

[0066] This application designs a loss function based on the regularized Hu moment, which can obtain a set of coefficients by performing a moment transformation on the complex representation of the contour. These coefficients can describe the main features of the shape and are invariant to translation, rotation and scaling. This application also further improves the adaptability to scale changes through normalization operations. Therefore, even if the target moves in position, changes angle or scales in size in the image, the regularized Hu moment can still remain consistent, making the model more robust when dealing with targets of different postures. At the same time, the regularized Hu moment can effectively capture the topological structure of the shape, not just the difference at the pixel level. By incorporating the regularized Hu moment loss into the loss function, the model is prompted to output a prediction mask that is not only close to the true value at the pixel level, but also more similar in shape features, thereby improving the prediction effect on small targets and complex shape targets. In addition, the low-order part of the regularized Hu moment mainly describes the overall structure of the shape, while the high-order part describes the details. By retaining the previous k A regularized Hu moment is used to force the model to focus on the main features of the shape and ignore some unnecessary detail noise. On the other hand, in order to better adjust the training process, the adaptive weight parameters are set. , which is used to balance the regularized Hu moment loss and pixel-level loss. By dynamically adjusting the value of the weight, the multimodal image segmentation model can flexibly weigh the importance of the two losses during the training process. Therefore, the loss function provided in this application can effectively improve the performance of the model in shape features, especially when dealing with small target segmentation tasks, and can better capture the shape and topological structure of the target, which not only enhances the robustness and generalization ability of the model, but also improves the training efficiency and the quality of the segmentation results.

[0067] S4: Segment the multimodal image pairs to be segmented using the trained multimodal target segmentation model.

[0068] The multimodal image to be segmented is input into the trained multimodal target segmentation model to obtain the segmentation result.

[0069] This application designs a method for aircraft target segmentation based on regularized Hu moments. By constructing a loss function based on regularized Hu moments, the edge shapes predicted by the multimodal image segmentation model are aligned with the true labels. At the same time, normalization is used to effectively capture the topological structure and overall shape of the target while enhancing the multimodal image segmentation model's adaptability to translation, rotation, and scaling. By designing a feature extraction network based on Swin Transformer and a depthwise separable convolutional network, the multimodal image segmentation model is enabled to effectively extract feature information at different scales, ultimately improving its performance in small target segmentation tasks.

[0070] This application addresses the shortcomings of existing segmentation methods in dealing with complex shapes and boundary details, and improves the robustness and segmentation quality of multimodal image segmentation models. Through detailed step descriptions and formula derivations, it provides a valuable reference for subsequent research. This application can significantly improve segmentation accuracy and robustness when dealing with aircraft target segmentation tasks, especially when dealing with small targets in complex backgrounds. Future work can further explore how to combine other shape descriptors (such as Zernike moments, Fourier descriptors, etc.) to enhance the feature representation capabilities of the model, or introduce more prior knowledge (such as physical constraints, geometric constraints, etc.) to improve segmentation results. In addition, it is also possible to study how to apply this method to other types of remote sensing image segmentation tasks to expand its scope of application.

[0071] In an exemplary embodiment, a multimodal image segmentation system for land and sea targets is provided, comprising: A target image data set acquisition module is used to acquire a target image data set; the target image data set includes: a training set and a test set; A preprocessing module, configured to preprocess the target image dataset to obtain a preprocessed dataset; the preprocessed dataset includes: a multimodal image pair and a corresponding true mask; the multimodal image pair includes: an infrared image and a corresponding visible light image; the preprocessing includes: image cropping and mask annotation; A multimodal target segmentation model establishment module is used to determine a trained multimodal target segmentation model based on the preprocessed data set, a feature extraction network based on SwinTransformer, and a depth-wise separable convolutional network; the multimodal target segmentation model takes a multimodal image pair as input and a predicted mask as output; the training process of the multimodal target segmentation model specifically includes: determining an overall loss function based on the true mask and the predicted mask using a mean square error loss function and a regularized Hu moment loss function; and training the multimodal target segmentation model using the overall loss function to obtain a trained multimodal target segmentation model; The image segmentation module is used to segment the multimodal image pairs to be segmented using the trained multimodal target segmentation model.

[0072] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 3As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store a multimodal target segmentation model. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multimodal image segmentation method for land and sea targets is implemented.

[0073] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above-mentioned method embodiments when executing the computer program.

[0074] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0075] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0076] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0077] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0078] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0079] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A multimodal image segmentation method for sea and land targets, characterized in that: The multimodal image segmentation method for land and sea targets includes: Obtain a target image dataset; the target image dataset includes: a training set and a test set; Preprocessing the target image dataset to obtain a preprocessed dataset; the preprocessed dataset includes: a multimodal image pair and a corresponding true mask; the multimodal image pair includes: an infrared image and a corresponding visible light image; the preprocessing includes: image cropping and mask annotation; According to the preprocessed data set, a trained multimodal target segmentation model is determined based on a Swin Transformer feature extraction network and a depthwise separable convolutional network; the multimodal target segmentation model takes a multimodal image pair as input and a predicted mask as output; the training process of the multimodal target segmentation model specifically includes: determining an overall loss function based on the true mask and the predicted mask using a mean square error loss function and a regularized Hu moment loss function; and training the multimodal target segmentation model using the overall loss function to obtain a trained multimodal target segmentation model; The trained multimodal object segmentation model is used to segment the multimodal image pairs to be segmented.

2. The multimodal image segmentation method for land and sea targets according to claim 1, characterized in that: The acquiring of the target image dataset specifically includes: The target image dataset is divided into a ratio of 7:3 to determine the training set and test set.

3. The multimodal image segmentation method for land and sea targets according to claim 1, characterized in that: The method of determining an overall loss function based on the true mask and the predicted mask using a mean square error loss function and a regularized Hu moment loss function specifically includes: Determining a pixel-level loss between the true mask and the predicted mask using a mean square error loss function; Extracting contour point sequences of the true mask and the predicted mask respectively, and determining corresponding true contours and predicted contours; Determining a corresponding Hu moment according to the true contour and the predicted contour; According to the Hu moment corresponding to the true contour and the Hu moment corresponding to the predicted contour, the corresponding regularized Hu moment is determined respectively; Determining a regularized Hu moment loss function according to the regularized Hu moment corresponding to the true contour and the regularized Hu moment corresponding to the predicted contour; and determining a regularized Hu moment loss using the regularized Hu moment loss function; An overall loss function is determined based on the pixel-level loss and the regularized Hu moment loss.

4. The multimodal image segmentation method for land and sea targets according to claim 3, characterized in that: The method of determining the regularized Hu moment loss using the regularized Hu moment loss function specifically includes: Using the formula Determine the regularized Hu moment loss ;in, K is the number of moments used, To seek peace, k is the number of the regularized Hu moment used, The first k A regularized Hu moment, The true contour k A regularized Hu moment.

5. The multimodal image segmentation method for land and sea targets according to claim 3, characterized in that: Determining the overall loss function based on the pixel-level loss and the regularized Hu moment loss specifically includes: Using the formula Determine the overall loss function ;in, is the pixel-level loss, is the regularized Hu moment loss, is an adaptive weight parameter used to balance the importance of the pixel-level loss and the regularized Hu moment loss.

6. The multimodal image segmentation method for land and sea targets according to claim 5, characterized in that: The multimodal object segmentation model is trained using the overall loss function to obtain a trained multimodal object segmentation model, specifically comprising: According to the overall loss function, the multimodal target segmentation model is updated through the back propagation algorithm and the Adam optimizer; and the adaptive weight parameters in the overall loss function are dynamically adjusted; until the number of updates is reached, the trained multimodal target segmentation model is obtained; the dynamic adjustment is to use the formula =0.2+e / 500*0.6 Determine adaptive weight parameters , increase the weight of the regularized Hu moment loss; where, e is the number of updates.

7. A multimodal image segmentation system for land and sea targets, characterized by: The multimodal image segmentation system for land and sea targets includes: A target image data set acquisition module is used to acquire a target image data set; the target image data set includes: a training set and a test set; A preprocessing module, configured to preprocess the target image dataset to obtain a preprocessed dataset; the preprocessed dataset includes: a multimodal image pair and a corresponding true mask; the multimodal image pair includes: an infrared image and a corresponding visible light image; the preprocessing includes: image cropping and mask annotation; A multimodal target segmentation model establishment module is used to determine a trained multimodal target segmentation model based on the preprocessed data set, a feature extraction network based on SwinTransformer, and a depth-wise separable convolutional network; the multimodal target segmentation model takes a multimodal image pair as input and a predicted mask as output; the training process of the multimodal target segmentation model specifically includes: determining an overall loss function based on the true mask and the predicted mask using a mean square error loss function and a regularized Hu moment loss function; and training the multimodal target segmentation model using the overall loss function to obtain a trained multimodal target segmentation model; The image segmentation module is used to segment the multimodal image pairs to be segmented using the trained multimodal target segmentation model.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal image segmentation method for land and sea targets according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multimodal image segmentation method for land and sea targets according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the multimodal image segmentation method for land and sea targets according to any one of claims 1 to 6 is implemented.