An Interactive Image Segmentation Method and System Based on Internal and External Guidance

Through an interactive image segmentation method based on internal and external guidance, the network module and pyramid feature processor with a coarse-to-fine structure are used to solve the problems of low segmentation accuracy and large annotation workload in complex scenarios, and efficient and accurate image segmentation is achieved.

CN114693927BActive Publication Date: 2025-05-27XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210302843.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2025-05-27
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

The existing interactive image segmentation method is difficult to achieve high-precision segmentation in complex and changing scenarios, and the annotation workload is large and the efficiency is low.

Method used

An interactive image segmentation method based on internal and external guidance is adopted to form the 5-channel input data of the network through preprocessing of the training set data, and the network module with the coarse-to-fine structure is used for convolution processing and decoding, combining the pyramid feature processor and cross-layer connection to perform global semantic acquisition and feature fusion.

Benefits of technology

It improves the accuracy and efficiency of image segmentation, is suitable for target segmentation in complex and changing scenarios, and reduces the annotation workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693927B_ABST
    Figure CN114693927B_ABST
Patent Text Reader

Abstract

The present invention discloses an interactive image segmentation method and system based on internal and external guidance. The method forms 5-channel input data of the network as training set data by preprocessing the image to be trained; performs convolution processing on the preprocessed training set data, decodes the convolved data, and then fuses the features of each layer after decoding through convolution operations to obtain a segmentation result; corrects the obtained segmentation result and inputs it into the bottom-layer pyramid feature module of the network segmentation model, and then uses the obtained segmentation result and the corresponding training set to train the corrected network segmentation model. The trained network segmentation model is used to implement image segmentation. The present invention uses the decoded and fused segmentation result and the corresponding training set to train the network segmentation model, and uses the trained network segmentation model to perform interactive image segmentation. Cross-stage feature aggregation is adopted in the process of network encoding and decoding. The present invention is applicable to target segmentation in complex and variable scenarios and can improve the segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an interactive image segmentation method and system based on internal and external guidance. Background Art

[0002] In the past few years, semantic and instance segmentation have made revolutionary progress in different fields, such as general scenes, autonomous driving, aerial images, and medical diagnosis. Successful segmentation models are usually built on a large amount of high-quality training data. However, the process of creating the pixel-level training data required to build these models is usually expensive, laborious, and time-consuming. Therefore, how to quickly extract the object of interest and reduce the annotation workload is one of the main research topics in such tasks.

[0003] The purpose of interactive image segmentation is to segment the target object of interest with the least amount of user interaction input. Interactive segmentation allows annotators to quickly extract the object of interest by providing some user inputs, such as bounding boxes or clicks, which is an effective method to reduce the annotation workload. It has practical applications in many fields, such as image editing and medical image analysis. In recent years, with the popularization of data-driven deep learning technologies, in some fields, the demand for pixel-level annotations has increased sharply, such as salient object detection, semantic segmentation, instance segmentation, camouflaged object detection, and image and video processing. There is a great need for efficient interactive segmentation technologies to reduce the annotation cost. Therefore, more and more researchers are conducting extensive explorations in this field.

[0004] In the early stage, most traditional interactive segmentation methods mainly utilized manually extracted features, and some research methods focused very much on boundary properties. Subsequently, graph model-based methods became more popular, that is, modeling the interactive segmentation task as a graph segmentation optimization problem and solving it with the well-known minimum cut / maximum flow algorithm. The classic method based on graph segmentation is GrabCut, which uses a Gaussian mixture model as the color model and a bounding box as the input, simplifying the segmentation process. Later, the random walk algorithm was improved, and a new higher-order formula was introduced, with an additional soft label consistency constraint. Then a fault-tolerant method was provided, allowing users to make some incorrect interactions. These methods based on low-level features cannot adapt to object segmentation in complex and variable scenarios. Summary of the Invention

[0005] The purpose of the present invention is to provide an interactive image segmentation method and system based on internal and external guidance to overcome the deficiencies of the prior art.

[0006] An interactive image segmentation method based on internal and external guidance includes the following steps:

[0007] S1. Preprocessing of the training set data: Use the ground truth to generate internal and external guiding points for the images to be trained as imitated artificial interaction points. Crop the images to be trained to the size of the circumscribed rectangle formed by the external guiding points. Then, generate 2D Gaussian centers for the internal and external guiding points, create two independent heatmaps for the internal and external guiding points, and connect the obtained heatmaps with the RGB input images to form the 5-channel input data of the network.

[0008] S2. Input the processed training set data into the network module with a coarse-to-fine structure for convolution processing. Decode the data after convolution processing, and add a pyramid feature processor at the deepest layer to obtain global semantics. At the same time, connect the low-level semantics of the convolution processing with the high-level semantics in the decoding stage at the same downsampling ratio through cross-layer connections. Then, fuse the features of each layer after decoding through a certain number of convolution operations to obtain the segmentation result.

[0009] S3. After obtaining the segmentation result, the image correction operation will be input into the model through a separate correction module. The correction is divided into external guiding point correction and internal guiding point correction. Then, generate 2D Gaussian centers for the internal and external guiding points, create two independent heatmaps for the internal and external guiding points, input the two correction heatmaps into the correction module for convolution operations, and then input them into the pyramid feature module at the bottom layer of the network.

[0010] S4. Use the decoded and fused segmentation result and the corresponding training set to train the network segmentation model, and use the trained network segmentation model to segment interactive images.

[0011] Furthermore, use the backpropagation strategy to optimize the parameters of the network, update the network parameters according to the value of the loss function, so that the loss function continuously decreases until it converges to the set value, and complete the training of the network segmentation model.

[0012] Furthermore, in the process of network encoding and decoding, cross-stage feature aggregation is adopted. From the downsampling and upsampling of the previous stage to the downsampling process of the current stage, two types of SEPA rate information flows are introduced, and 1×1 convolutions are added to each process, so as to better utilize the prior information and extract more discriminative features.

[0013] Furthermore, a channel attention module is used in the model to explicitly implement the dependency relationship of feature channels, automatically learn the importance of each channel feature, and then assign a weight value to each feature channel with this importance, so that the network focuses on certain feature channels, that is, enhance the feature channels useful for the current task and suppress the feature channels that are not very useful for the current task.

[0014] Furthermore, a feature pyramid module is applied to each layer in CorseNet to better obtain global semantic information.

[0015] Furthermore, the dataset is augmented by means of random cropping, Gaussian blur, contrast enhancement, or mirror flipping.

[0016] Furthermore, the user correction module is separated from the main network, and the heatmap obtained by user correction is input into the bottom layer encoded by CorseNet, which is more conducive to transmitting the user's intention into the model.

[0017] Furthermore, when the user input information is insufficient, expansion operations can be performed according to the external points already input by the user. New external guiding points are automatically generated at a distance of 1% of the diagonal length of the image extending outward from the four extreme points of the object to the outside of the circumscribed rectangle of the object. At the same time, new external guiding points are automatically generated outside the external guiding points, that is, at a distance of 1% of the diagonal length of the image. The automatically generated guiding points can provide more prior knowledge for the model to achieve better segmentation results.

[0018] Furthermore, the features are encoded and decoded through a network model with a coarse-to-fine structure to better learn the boundary information of the object to be segmented, which is conducive to obtaining better segmentation results.

[0019] An interactive image segmentation system based on internal and external guidance includes a data preprocessing module, a segmentation module, and a correction module;

[0020] The data preprocessing module uses the ground truth to generate internal and external guiding points of the training image as artificial interaction points for imitation, crops the training image to the size of the circumscribed rectangle formed by the external guiding points, then generates 2D Gaussian centers for the internal guiding points and the external guiding points, creates two independent heatmaps for the internal guiding points and the external guiding points, and connects the obtained heatmaps with the RGB input image to form the 5-channel input data of the network;

[0021] The segmentation module inputs the processed training set data into a network module with a coarse-to-fine structure for convolution processing, decodes the data after convolution processing, adds a pyramid feature processor in the deepest layer to obtain global semantics, and at the same time connects the low-level semantics of the convolution processing with the high-level semantics in the decoding stage at the same downsampling ratio through cross-layer connection. Then, the features of each layer after decoding are fused after a certain number of convolution operations to obtain the segmentation result. The network segmentation model is trained using the segmentation result after decoding and fusion and the corresponding training set, and the trained network segmentation model is used for the segmentation of interactive images;

[0022] The correction module is used for re-segmentation when the user is not satisfied with the segmentation result.

[0023] After the user obtains the segmentation result, the correction operation on the image will be input into the model through a separate correction module. The user's correction is divided into external guiding point correction and internal guiding point correction. Then, 2D Gaussian centers are generated for the internal guiding points and external guiding points, and two independent heatmaps are created for the internal guiding points and external guiding points. The two corrected heatmaps are input into the correction module for convolution operations, and then input into the bottom-layer pyramid feature module of the network. After that, the network is trained to output the final segmentation result.

[0024] Compared with the prior art, the present invention has the following beneficial technical effects:

[0025] The present invention provides an interactive image segmentation method based on internal and external guidance. The method forms 5-channel input data of the network as the training set data by preprocessing the image to be trained; performs convolution processing on the preprocessed training set data, decodes the convolved data, and then fuses the features of each layer after decoding through convolution operations to obtain the segmentation result; corrects the obtained segmentation result and inputs it into the bottom-layer pyramid feature module of the network segmentation model, and then uses the obtained segmentation result and the corresponding training set to train the corrected network segmentation model. The trained network segmentation model is used to implement image segmentation. The present invention uses the decoded and fused segmentation result and the corresponding training set to train the network segmentation model, and uses the trained network segmentation model to perform interactive image segmentation. Cross-stage feature aggregation is adopted in the process of network encoding and decoding. The present invention is applicable to target segmentation in complex and variable scenarios and can improve the segmentation accuracy.

[0026] Furthermore, a channel attention module is used in the model to explicitly implement the dependency relationship of feature channels. The importance degree of each channel feature is obtained through an automatic learning method, and then a weight value is assigned to each feature channel with this importance degree, enabling the network to focus on certain feature channels, that is, enhancing the feature channels useful for the current task and suppressing the feature channels that are not very useful for the current task. The network model with a coarse-to-fine structure is used to encode and decode the features to better learn the boundary information of the object to be segmented and obtain better segmentation accuracy.

[0027] Specifically, from the downsampling and upsampling in the previous stage to the downsampling process in the current stage, two types of SEPA rate information flows are introduced, and 1×1 convolutions are added to each process, so as to better utilize the prior information and extract more discriminative features.

[0028] Furthermore, by jointly using cross-entropy loss, IOU loss, and deep supervision loss, the backpropagation of gradients is promoted, the model convergence is strengthened, and the model training effect is further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 FIG. is a flowchart of the implementation of the internal and external guidance-based interactive image segmentation method in an embodiment of the present invention.

[0030] Figure 2 FIG. is a network structure diagram of the internal and external guidance-based interactive image segmentation model in an embodiment of the present invention.

[0031] Figure 3 FIG. is a cross-stage feature aggregation module diagram of the internal and external guidance-based interactive image segmentation model in an embodiment of the present invention.

[0032] Figure 4 FIG. is a channel attention module diagram of the internal and external guidance-based interactive image segmentation model in an embodiment of the present invention.

[0033] Figure 5 FIG. is a segmentation effect diagram of the internal and external guidance-based interactive image segmentation model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0036] As Figure 1 shown, an internal and external guidance-based interactive image segmentation method is provided to achieve better segmentation accuracy, and specifically includes the following steps;

[0037] S1, Data preprocessing: Specifically, use the ground truth to generate internal and external guiding points for the training images to imitate human interaction points. The internal guiding points are the four extreme points of the object to be segmented, namely the topmost, bottommost, leftmost, and rightmost points. The external guiding points are the four vertices of the minimum bounding rectangle of the object to be segmented. Then, crop the training images to the size of the bounding rectangle formed by the external guiding points. Next, generate 2D Gaussian centers for the internal and external guiding points, create two independent heatmaps for the foreground (internal guiding points) and background (external guiding points), and concatenate the obtained heatmaps with the RGB input images to form the 5-channel input data of the network, that is, the preprocessed training set data;

[0038] S2, Input the processed training set data into the network module with a coarse-to-fine structure for convolution processing. Decode the convolved data, and add a pyramid feature processor at the deepest layer to obtain global semantics. At the same time, connect the low-level semantics of the convolution processing with the high-level semantics in the decoding stage at the same downsampling ratio through cross-layer connections. Then, fuse the features of each layer after decoding through a certain number of convolution operations to finally obtain the segmentation result;

[0039] S3, After obtaining the segmentation result, the image correction operation will be input into the model through a separate correction module. The correction is divided into external guiding point correction and internal guiding point correction. Then, generate 2D Gaussian centers for the internal and external guiding points, create two independent correction heatmaps for the internal and external guiding points, input the two correction heatmaps into the correction module for convolution operations, and then input them into the bottom pyramid feature module of the network segmentation model;

[0040] Specifically, after obtaining the segmentation result, the user's image correction operation will be input into the model through a separate correction module. The user correction is divided into external guiding point correction and internal guiding point correction; separate the user correction module from the main network, and input the heatmap obtained by the user correction into the bottom layer after encoding by CorseNet, which is more conducive to transmitting the user's intention to the model.

[0041] S4, Use the decoded and fused segmentation result and the corresponding training set to train the network segmentation model that inputs the correction heatmap, and use the trained network segmentation model to segment interactive images.

[0042] During the training process, use the backpropagation strategy to optimize the parameters of the network, use the loss function to assist in training, and update the network parameters according to the value of the loss function, so that the loss function continuously decreases until it converges to the set value. At this time, the training ends, and the trained network segmentation model is saved.

[0043] When the user input information is insufficient, expand the operation according to the external points already input by the user, and automatically generate new external guiding points at 1% of the length of the diagonal of the picture extended outward from the four extreme points of the object to the outside of the circumscribed rectangle of the object. At the same time, new external guiding points are automatically generated outside the external guiding points, that is, at 1% of the length of the diagonal of the picture. The automatically generated guiding points can provide more prior knowledge for the model to achieve better segmentation results.

[0044] Specifically, this application uses a publicly available dataset as the dataset and divides the dataset into a training set and a test set.

[0045] In the process of network decoding, cross-stage feature aggregation is adopted. From the downsampling and upsampling of the previous stage to the downsampling process of the current stage, two SEPA rate information flows are introduced, and 1×1 convolutions are added to each process, so as to better utilize prior information and extract more discriminative features.

[0046] In the network segmentation model, a channel attention module is used to explicitly implement the dependence relationship of feature channels, obtain the importance degree of each channel feature through automatic learning, assign a weight value to each feature channel according to the importance degree of each channel feature, and let the network focus on certain feature channels, that is, enhance the feature channels useful for the current task and suppress the feature channels that are not very useful for the current task.

[0047] Segmentation module design: Adopt a design from coarse segmentation to fine segmentation to solve the problem that the segmentation boundary of the segmentation algorithm is not fine enough. Among them, a cascade structure is adopted. Specifically, the segmentation network consists of two subnets. The first subnet, the coarse segmentation network adopts the FPN design, and through horizontal connection, gradually fuses the deep semantic information with the shallow low-level details. Different from CPN, a pyramid scene parsing module is also added to the deepest layer to enrich the representation with global context information. For the second subnet, the fine segmentation network, its goal is to restore the lost boundary details. It is realized through a multi-scale fusion structure, which fuses information at different levels of the coarse segmentation network through upsampling and cascade operations; more convolutional blocks are applied to the deeper features to achieve a better balance between accuracy and efficiency.

[0048] At the same time, connect the low-level semantics in the encoding stage with the high-level semantics in the decoding stage at the same downsampling ratio through cross-layer connection, and finally decode to obtain the segmentation result.

[0049] The dataset is augmented by means of random cropping, Gaussian blur, contrast enhancement, or mirror flipping. The training is carried out using the training data. During the training process, the backpropagation strategy is used to optimize the parameters of the network, and the loss function is used to assist the training. The loss functions used include cross-entropy loss and IOU loss to assist the training, and the backpropagation is used to optimize the parameters of the network.

[0050] Cross-entropy is the most commonly used loss in image segmentation algorithms. It compares each pixel with the ground truth map one by one, and its formula is expressed as follows:

[0051]

[0052] In the formula: is the number of pixels in the entire three-dimensional image, is the true label of the th element, where 0 represents the background and 1 represents the foreground, represents the probability that the network predicts that this pixel belongs to the foreground.

[0053] The formula for the IOU loss is expressed as follows:

[0054]

[0055] In the formula: intersection is the intersection area, and union is the area of the union part.

[0056] According to the value of the loss function, the network parameters are updated to make the loss function continuously decrease until it converges to a smaller value. At this time, the training ends, and the trained network model is saved; the saved trained model is used to construct an internal and external guidance-based interactive image segmentation network model.

[0057] An internal and external guidance-based interactive image segmentation system includes a data preprocessing module, a segmentation module, and a correction module;

[0058] The data preprocessing module uses the ground truth to generate internal and external guidance points for the training images to be used as imitating manual interaction points. The internal guidance points are the four extreme points of the top, bottom, left, and right of the object to be segmented, and the external guidance points are the four vertices of the minimum circumscribed rectangle of the object to be segmented. Then, the training image to be used is cropped to the size of the circumscribed rectangle formed by the external guidance points. Then, 2D Gaussian centers are generated for the internal and external guidance points, and two independent heatmaps are created for the foreground, that is, the internal guidance points, and the background, that is, the external guidance points. The obtained heatmaps are concatenated with the RGB input image to form the 5-channel input data of the network;

[0059] The segmentation module inputs the processed training set data into the network module with a coarse-to-fine structure for convolution processing, decodes the data after convolution processing, adds a pyramid feature processor at the deepest layer to obtain global semantics, and connects the low-level semantics of the convolution processing with the high-level semantics in the decoding stage at the same downsampling ratio through cross-layer connections. Then, the features of each layer after decoding are fused after a certain number of convolution operations to obtain the segmentation result. The network segmentation model is trained using the decoded and fused segmentation result and the corresponding training set, and the trained network segmentation model is used for the segmentation of interactive images;

[0060] The correction module is used for re-segmentation when the user is not satisfied with the segmentation result. The correction operation of the user on the image after obtaining the segmentation result will be input into the model through a separate correction module. The user's correction is divided into external guidance point correction and internal guidance point correction. Then, 2D Gaussian centers are generated for the internal guidance points and external guidance points, and two independent heatmaps are created for the internal guidance points and external guidance points. The two corrected heatmaps are input into the correction module for convolution operations, and then input into the pyramid feature module at the bottom layer of the network. After that, the network is trained to output the final segmentation result.

[0061] An interactive image segmentation method based on internal and external guidance of the present invention generates internal and external guidance points of the image to be trained according to the groundtruth in the dataset, and then generates 5-channel input data for the segmentation network. The segmentation network adopts a design from coarse segmentation to fine segmentation to solve the problem that the segmentation boundary of the segmentation algorithm is not fine enough; a cascade structure is adopted, and the segmentation network consists of two subnets.

[0062] The first subnet, the coarse segmentation network adopts a design similar to FPN, and gradually fuses the deep semantic information with the shallow low-level details through horizontal connections. Different from CPN, a pyramid scene parsing module is also added at the deepest layer to enrich the representation with global context information.

[0063] The second subnet, the fine segmentation network, aims to recover the lost boundary details. This is achieved through a multi-scale fusion structure that fuses information at different levels of the coarse segmentation network through upsampling and cascading operations. Similar to CPN, more convolutional blocks are also applied to the deeper features to achieve a better balance between accuracy and efficiency.

[0064] By jointly using the cross-entropy loss and the IOU loss, the backpropagation of the gradient is promoted, the model convergence is strengthened, and the model training effect is further improved;

[0065] This application has achieved competitive IOU results on the public dataset PASCAL, and the performance is better than several popular interactive image segmentation methods.

[0066] Embodiment

[0067] An interactive image segmentation method based on internal and external guidance, comprising the following steps:

[0068] S1. Preprocess the PASCAL source data to make it suitable for model training, and divide it into a training set and a test set.

[0069] The specific work process is as follows:

[0070] (1.1) Use the publicly available dataset PASCAL as the dataset;

[0071] (1.2) Generate internal and external guidance points for the training images using the ground truth of the dataset in step (1.1) as imitating manual interaction points. The internal guidance points are the four extreme points of the top, bottom, left, and right of the object to be segmented, and the external guidance points are the four vertices of the minimum bounding rectangle of the object to be segmented;

[0072] (1.3) Crop the training images to the size of the bounding rectangle formed by the external guidance points;

[0073] (1.4) Generate 2D Gaussian centers for the internal and external guidance points for the data processed in step (1.3), create two independent heatmaps for the foreground, i.e., the internal guidance points, and the background, i.e., the external guidance points, and connect the obtained heatmaps with the RGB input image to form the 5-channel input data of the network.

[0074] S2. Then, according to the characteristics of the preprocessed dataset, in order to make the network have better segmentation accuracy, the network is processed in a network module with a coarse-to-fine structure. The specific work process is as follows:

[0075] (2.1) For the input data obtained in step (1.4), first input the input data into CorseNet for coarse segmentation to extract more general features;

[0076] (2.2) For CorseNet described in step (2.1), adopt a design similar to FPN, and gradually fuse the deep semantic information with the shallow low-level details through lateral connections. Different from CPN, a pyramid scene parsing module is also added to the deepest layer to enrich the representation with global context information;

[0077] (2.3) Then, input the decoded data of each layer of CorseNet into the second sub-network FineNet for fine segmentation, and its goal is to recover the lost boundary details;

[0078] (2.4) For the FineNet described in step (2.3), this is achieved through a multi-scale fusion structure that fuses information at different levels of the coarse segmentation network through upsampling and cascading operations. Similar to CPN, more convolutional blocks are also applied to deeper features to achieve a better trade-off between accuracy and efficiency.

[0079] S3. Design cross-stage feature aggregation during the network encoding and decoding process. From the downsampling and upsampling of the previous stage to the downsampling process of the current stage, two types of SEPA rate information flows are introduced, and 1×1 convolutions are added to each process, so as to better utilize prior information and extract more discriminative features, as Figure 3 shown. The specific workflow is as follows:

[0080] (3.1) For each scale, two independent information flows are introduced from the downsampling and upsampling units of the previous stage to the downsampling process of the current stage. It should be noted that 1×1 convolutions are added to each flow, combined with the downsampling features of the current stage, and three components are added to generate a fusion result. Through this design, more discriminative representations can be fully utilized to extract prior information at the current stage;

[0081] S4. Design a channel attention module to explicitly implement the dependence relationship of feature channels, obtain the importance of each channel feature through automatic learning, and then use this importance to assign a weight value to each feature channel, so that the network focuses on certain feature channels, that is, enhance the feature channels useful for the current task and suppress the feature channels that are not very useful for the current task, as Figure 4 shown. The specific workflow is as follows:

[0082] (4.1) Given an input x with the number of feature channels C', a feature with the number of feature channels C is obtained through a series of convolutions (Ftr) and other transformations.

[0083] (4.2) Squeeze (Fsq), through global pooling, compresses the two-dimensional feature (HxW) of each channel into a real number, which is achieved through global average pooling. This belongs to a kind of feature compression in the spatial dimension. Because this real number is calculated based on all feature values, it has an average receptive field to some extent and keeps the number of channels unchanged. Therefore, after the squeeze operation, it becomes 1x1xC;

[0084] (4.3), excitation (Fex), generates a weight value for each feature channel through parameters. How this weight value is generated is crucial. It models the correlation between channels by forming a Bottleneck structure with two fully connected layers and outputs a weight value with the same number of input and output features.

[0085] S5. Decode the output obtained by the data through the operations described in step (2.2). At the same time, connect the low-level semantics in the encoding stage with the high-level semantics in the decoding stage at the same downsampling ratio through cross-layer connections, and finally decode to obtain the segmentation result.

[0086] S6. For the network model described in step S5, use cross-entropy loss and IOU loss during the training process to promote gradient backpropagation, strengthen model convergence, and further improve the training effect.

[0087] S7. For the trained internal and external guidance interactive image segmentation model, use the test image as input to obtain the segmentation result, as Figure 2 shown. The specific workflow is as follows:

[0088] (7.1), for the internal and external guidance interactive image segmentation model described in step S5, use the test set described in step (1.4) as input to obtain the segmentation result of the model.

[0089] (7.2), compare the object contour obtained by the interactive segmentation of the internal and external guidance interactive image segmentation model described in step (7.1) with the ground truth label, and find that the internal and external guidance interactive image segmentation model described in step (7.1) has achieved excellent segmentation results, especially outstanding in terms of the IOU index, as Figure 5 shown.

Claims

1. An interactive image segmentation method based on internal and external guidance, characterized in that, it includes the following steps: S1, preprocess the image to be trained to form 5-channel input data of the network as training set data; S2, perform convolution processing on the preprocessed training set data, decode the convolved data, input the processed training set data into the network module for convolution processing, decode the convolved data, and add a pyramid feature processor at the deepest layer to obtain global semantics. At the same time, connect the low-level semantics of the convolution processing with the high-level semantics in the decoding stage at the same downsampling ratio through cross-layer connection. Then, after a certain number of convolution operations on the features of each decoded layer, fuse them to obtain the segmentation result; S3, input the corrected segmentation result into the bottom pyramid feature module of the network segmentation model, and then use the obtained segmentation result and the corresponding training set of the segmentation result to train the corrected network segmentation model, and use the trained network segmentation model to implement image segmentation.

2. An interactive image segmentation method based on internal and external guidance according to claim 1, characterized in that, using groundtruth to generate internal and external guidance points of the image to be trained as artificial interaction point imitations, cropping the image to be trained to the size of the circumscribed rectangle formed by the external guidance points, then generating 2D Gaussian centers for the internal and external guidance points, creating two independent heatmaps for the internal and external guidance points, and connecting the obtained heatmaps with the RGB input image to form 5-channel input data of the network.

3. An interactive image segmentation method based on internal and external guidance according to claim 1, characterized in that, using the backpropagation strategy to optimize the parameters of the network, updating the network parameters according to the value of the loss function, so that the loss function continuously decreases until it converges to the set value, and completing the training of the network segmentation model.

4. An interactive image segmentation method based on internal and external guidance according to claim 1, characterized in that, cross-stage feature aggregation is adopted during the network decoding process, from the downsampling and upsampling of the previous stage to the downsampling process of the current stage.

5. An interactive image segmentation method based on internal and external guidance according to claim 1, characterized in that, a channel attention module is set in the network segmentation model, and the importance degree of each channel feature is obtained through automatic learning, and a weight value is assigned to each feature channel according to the importance degree of each channel feature.

6. An interactive image segmentation method based on internal and external guidance according to claim 1, characterized in that, decode the convolved data, add a pyramid feature processor at the deepest layer to obtain global semantics. At the same time, connect the low-level semantics of the convolution processing with the high-level semantics in the decoding stage at the same downsampling ratio through cross-layer connection. After convolution operations on the features of each decoded layer, fuse them to obtain the segmentation result.

7. An interactive image segmentation method based on internal and external guidance according to claim 1, characterized in that, The dataset is augmented by means of random cropping, Gaussian blurring, contrast enhancement or mirror flipping.

8. An interactive image segmentation method based on internal and external guidance according to claim 2, characterized in that the internal guidance points are the four extreme points of the object to be segmented, namely the uppermost, lowermost, leftmost and rightmost points, and the external guidance points are the four vertices of the minimum bounding rectangle of the object to be segmented.

9. An interactive image segmentation method based on internal and external guidance according to claim 1, characterized in that the features are encoded and decoded by a network model with a coarse-to-fine structure.

10. An interactive image segmentation system based on internal and external guidance, characterized in that it includes a data preprocessing module, a segmentation module and a correction module; the data preprocessing module is used to preprocess the training images to form 5-channel input data of the network as the training set data; the segmentation module is used to perform convolution processing on the preprocessed training set data, decode the data after convolution processing, input the processed training set data into the network module for convolution processing, decode the data after convolution processing, and add a pyramid feature processor to the deepest layer to obtain global semantics. At the same time, the low-level semantics of the convolution processing are connected with the high-level semantics in the decoding stage at the same downsampling ratio through cross-layer connections, and then the features of each layer after decoding are fused after a certain number of convolution operations to obtain the segmentation result; the correction module corrects the obtained segmentation result and inputs it to the bottom-layer pyramid feature module of the network segmentation model, and then uses the obtained segmentation result and the training set corresponding to the segmentation result to train the corrected network segmentation model to realize image segmentation.

Citation Information

Patent Citations

  • Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field

    AU2020103901A4

  • Method for interactive segmenting an object on an image and electronic computing device implementing the same

    WO2021150017A1