A method for segmenting human images based on feature contours

By obtaining the feature outline of the portrait image, using Mean-Shift preprocessing and pre-training network to generate high-quality three-value maps, the problem of low portrait segmentation accuracy in the prior art is solved, and efficient and accurate portrait segmentation is achieved.

CN115272378BActive Publication Date: 2025-07-11XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210622266.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-07-11
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

The existing portrait segmentation methods have problems with low segmentation accuracy and low automation when generating three-value maps, especially the edge transparency error, which affects the segmentation effect.

Method used

By obtaining the feature outline of the portrait image, using Mean-Shift preprocessing to generate edge curves, determine the effective domain and contour points, and combine the pretrained semantic segmentation network and fine segmentation network to generate high-quality three-value maps.

Benefits of technology

The generation efficiency and fineness of the three-value map are improved, the accuracy of portrait segmentation is improved, the edge clarity is improved, error is reduced, and the degree of automation is high.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272378B_ABST
    Figure CN115272378B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for segmenting a human image based on a feature contour, comprising: obtaining a human image to be processed; performing Mean-Shift preprocessing on the human image to be processed to obtain an edge curve; determining the effective domain of each pixel point on the edge curve, and determining contour points according to the effective domain; determining a preset number of feature contours according to a plurality of contours formed by the contour points, the area and perimeter of each contour; inputting the human image to be processed and the feature contours into a semantic segmentation network to obtain a ternary image after semantic segmentation of the human image to be processed; inputting the human image to be processed and the ternary image into a fine segmentation network so that the fine segmentation network predicts the transparency of the pixel points in the unknown area of the ternary image to obtain a segmentation result of the human image to be processed. The present invention can automatically generate a high-quality ternary image, which not only improves the generation efficiency of the ternary image and improves the fineness of the ternary image, but also is beneficial to making the segmentation of the human image more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to a method for segmenting human images based on feature contours. Background Art

[0002] Image segmentation belongs to the category of computer vision. It originally only referred to the semantic segmentation of images, and now it includes image semantic segmentation and fine segmentation. Among them, the image fine segmentation technology is mainly applied to close-range segmentation tasks with human figures as the target.

[0003] Currently, most image fine segmentation methods are based on deep learning and can be divided into two types: with trimap and without trimap. Among them, the image segmentation method with trimap needs to first obtain a fine trimap and then use it together with the original Figure 1 image as the input of the network. Through the calculation and prediction of the network, a corresponding transparency mask map (alpha matting) is generated. The image segmentation method without trimap directly predicts the foreground, background pixels and their transparencies from the original image, and it is difficult to guarantee the segmentation accuracy. Therefore, the image segmentation method with trimap has become a research hotspot in this field.

[0004] Trimap usually consists of three parts: the foreground area, the determined background area, and the unknown area. Image fine segmentation is to accurately predict the unknown area of the trimap to obtain the final alpha matting. Therefore, a high-quality trimap is crucial for the efficiency and accuracy of fine segmentation. In related technologies, the generation methods of trimap include manual marking and automatic generation algorithms. The tools for manual marking include photoshop, autocad, and visio, etc. Although the marking accuracy can be improved a lot by humans, it takes a lot of time and manpower. The automatic generation method of trimap is to continue the erosion and dilation operations on the result of rough segmentation. However, the erosion and dilation operations generally use a convolution kernel of a fixed size, and the quality of the trimap obtained with different kernel sizes is also different. In practical applications, it is not highly available and lacks universality.

[0005] After obtaining the trimap, the Deep Image Matting algorithm is used for human figure segmentation in the prior art. Specifically, first, an encoder-decoder network is used to predict the rough alpha matting of the human figure, and then a shallow network is used to describe and correct the blurred edges of the alpha matting more precisely. However, there are large errors in the transparency of the edges of the human figures segmented by the above method, which will affect the automation degree of the algorithm and require an additional loss function during correction. Therefore, the automation level and segmentation accuracy of the existing human figure segmentation methods need to be improved. Summary of the Invention

[0006] To solve the above problems existing in the prior art, the present invention provides a method for segmenting a human image based on a feature contour. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] The present invention provides a method for segmenting a human image based on a feature contour, including:

[0008] Obtaining a human image to be processed;

[0009] Performing Mean-Shift preprocessing on the human image to be processed to obtain an edge curve;

[0010] Determining the effective domain of each pixel point on the edge curve, and determining contour points from all pixel points of the edge curve according to the effective domain;

[0011] According to the multiple contours formed by the contour points, calculating the area and perimeter of each contour respectively, and determining a preset number of feature contours according to the area and the perimeter;

[0012] Inputting the human image to be processed and the feature contour into a pre-trained semantic segmentation network, so that the semantic segmentation network generates a low-level feature stream, generating a feature map according to the feature contour and the human image to be processed, further restoring a first image according to the feature map and the low-level feature stream, and classifying the pixel points in the first image into a foreground region, a background region or an unknown region to obtain a ternary map after semantic segmentation of the human image to be processed;

[0013] Inputting the human image to be processed and the ternary map into a pre-trained fine segmentation network, so that the fine segmentation network predicts the transparency of the pixel points in the unknown region of the ternary map to obtain a fine segmentation result of the human image to be processed.

[0014] In an embodiment of the present invention, the step of determining the effective domain of each pixel point on the edge curve and determining contour points from all pixel points of the edge curve according to the effective domain includes:

[0015] According to the expression y = f(x) of the edge curve and a curvature detection algorithm, determining a discrete curvature of a pixel point P i on the edge curve:

[0016] S(p i ) = f i+1 -f i

[0017] where P i represents the i-th pixel point on the edge curve, and S(P i ) represents the pixel point Pi Discrete curvature;

[0018] Take a pixel point P on the edge curve i Adjacent pixel point P i-k , P i+k , and calculate the length l when k = 1, 2, 3K respectively And the vertical distance d from P ik To i To Vertical distance d ik ; Wherein, Represents the chord between pixel points P i-k , P i+k ;

[0019] When the Length l ik And P i To Vertical distance d ik Satisfy the preset conditions, determine the effective domain radius and effective domain D(P corresponding to the pixel point P according to the maximum value k of k i i ); i )

[0020] For a pixel point P on the edge curve, when it satisfies the condition |S(P i i )| ≥ |S(P j )|, keep the pixel point; otherwise, delete the pixel point and the pixel point with discrete curvature of 0; where |i - j| ≤ k i / 2; i

[0021] For the remaining pixel points on the edge curve, when the effective domain radius k corresponding to the pixel point P i i = 1 and there is P i Or P i-1 In the effective domain D(P i+1 ), then delete the pixel point P that satisfies the condition |S(P i )| ≤ |S(P i-1 )| or |S(P i )| ≤ |S(P i+1 )| i ; i ;

[0022] For the remaining pixel points on the edge curve, if there are more than 2 pixel effective domains, delete the pixel points P outside the endpoints of the effective domain i ; If there is an effective domain containing 2 pixel points, when |S(P i )| > |S(P i+1 )|, delete the pixel point Pi+1 , and when |S(P i )| < |S(P i+1 )|, delete the pixel point P i ;

[0023] Determine the remaining pixel points as contour points.

[0024] In an embodiment of the present invention, the preset conditions include a first preset condition and a second preset condition; wherein,

[0025] The first preset condition is: l ik ≥ l i,k+1 ;

[0026] The second preset condition is:

[0027]

[0028] In an embodiment of the present invention, when the length l ik and P i to vertical distance d ik meet the preset conditions, according to the maximum value k i of k, determine the effective domain radius and effective domain D(P i ) corresponding to the pixel point P i ) steps, including:

[0029] When the length l ik meets the first preset condition and / or the length l ik and the vertical distance d i from P meet the second preset condition, determine the maximum value k ik of k as the radius of the effective domain of the pixel point P i , and determine the effective domain D(P i ) of the pixel point P i according to the radius. i )

[0030] In an embodiment of the present invention, the steps of calculating the area and perimeter of each contour respectively according to the multiple contours composed of the contour points, and determining a preset number of characteristic contours according to the area and the perimeter, include:

[0031] For the multiple contours composed of the contour points, number them respectively and establish a network structure;

[0032] Calculate the area s and perimeter l of each of the contours, and obtain two arrays: L = [l1, l2, K, lm and S = [s1, s2, ..., s m , where m represents the number of the contour;

[0033] When the corresponding elements in the two arrays satisfy the condition (l m ≥ a) ∪ (s m ≥ b), the contour numbered m is determined as the characteristic contour; where a represents a preset perimeter threshold, and b represents a preset area threshold.

[0034] In an embodiment of the present invention, the semantic segmentation network includes: a first encoder, an Atrous Spatial Pyramid Pooling (ASPP) module, a first decoder, and a Softmax layer. The first encoder includes multiple Blocks of the MobileNetV3-mid network; where

[0035] The first encoder is used to generate a low-level feature stream and generate a sub-feature map based on the input characteristic contour and the person image to be processed;

[0036] The ASPP module is used to extract different-scale features in the sub-feature map and generate a feature map based on the sub-feature map and the different-scale features;

[0037] The first decoder is used to restore a first image according to the feature map and the low-level feature stream, and the size of the first image is the same as that of the person image to be processed;

[0038] The Softmax layer is used to classify each pixel point in the first image into a foreground region, a background region, or an unknown region, and output a ternary map.

[0039] In an embodiment of the present invention, the fine segmentation network includes a U-net and a Contextual Attention Module (CAM). The U-net includes a second encoder and a second decoder; where

[0040] The second encoder is used to extract the transparency feature of the person image to be processed and generate an Alpha feature stream after obtaining the ternary map and the person image to be processed;

[0041] The CAM module is used to predict the transparency of the pixel points classified into the unknown region according to the transparency of the pixel points classified into the foreground region and the background region

[0042] The second decoder is used to restore the original resolution of the person image to be processed and predict a transparency mask map according to the Alpha feature stream, and obtain the segmentation result of the person image to be processed.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] The present invention provides a method for segmenting a human image based on a feature contour. First, a prior generation of the human feature contour is performed, and the feature contour is used to mark the human image to be processed. Then, a semantic segmentation network is used to predict a ternary map through three-class classification, and a high-quality alpha matting is further segmented by using a fine segmentation network. This method can automatically generate a high-quality ternary map, which not only improves the generation efficiency of the ternary map and improves the fineness of the ternary map, but also is conducive to making the subsequent segmentation of the human image more accurate.

[0045] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Description of the Drawings

[0046] Figure 1 is a flowchart of a method for segmenting a human image based on a feature contour provided by an embodiment of the present invention;

[0047] Figure 2 is a schematic diagram of a method for segmenting a human image based on a feature contour provided by an embodiment of the present invention;

[0048] Figure 3a is a result diagram of a contour detection method in the related art;

[0049] Figure 3b is an example diagram of a feature contour provided by an embodiment of the present invention;

[0050] Figure 4 is a schematic diagram of a semantic segmentation network provided by an embodiment of the present invention;

[0051] Figure 5 is a schematic diagram of a Block in the MobileNetV3 network provided by an embodiment of the present invention;

[0052] Figure 6 is a schematic diagram of the Last Stage structure in the MobileNetV3-mid network provided by an embodiment of the present invention;

[0053] Figure 7 is a schematic diagram of a fine segmentation network provided by an embodiment of the present invention;

[0054] Figure 8 is a schematic diagram of each Block module in the fine segmentation network provided by an embodiment of the present invention;

[0055] Figure 9 is a schematic diagram of a CAM module provided by an embodiment of the present invention;

[0056] Figure 10 is a comparison diagram of the generation results of the ternary map provided by an embodiment of the present invention;

[0057] Figure 11 It is a comparison diagram of the human image segmentation results provided by the embodiments of the present invention. Detailed implementation manners

[0058] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0059] Figure 1 It is a flowchart of a method for segmenting a human image based on a feature contour provided by an embodiment of the present invention. Figure 2 It is a schematic diagram of a method for segmenting a human image based on a feature contour provided by an embodiment of the present invention. Please refer to Figure 1-2 As shown, an embodiment of the present invention provides a method for segmenting a human image based on a feature contour, including:

[0060] S1. Obtain a human image to be processed;

[0061] S2. Perform Mean-Shift preprocessing on the human image to be processed to obtain an edge curve;

[0062] S3. Determine the effective domain of each pixel point on the edge curve, and determine the contour points from all pixel points of the edge curve according to the effective domain;

[0063] S4. Calculate the area and perimeter of each contour respectively according to the multiple contours composed of the contour points, and determine a preset number of feature contours according to the area and perimeter;

[0064] S5. Input the human image to be processed and the feature contour into a pre-trained semantic segmentation network, so that the semantic segmentation network generates a low-level feature stream, generates a feature map according to the feature contour and the human image to be processed, further restores a first image according to the feature map and the low-level feature stream, and classifies the pixel points in the first image into a foreground region, a background region or an unknown region to obtain a ternary map of the human image to be processed after semantic segmentation;

[0065] S6. Input the human image to be processed and the ternary map into a pre-trained fine segmentation network, so that the fine segmentation network predicts the transparency of the pixel points in the unknown region of the ternary map to obtain the segmentation result of the human image to be processed.

[0066] In the above method for segmenting a human image, Mean-Shift preprocessing can be performed on the human image to be processed first. Specifically, define the Mean-Shift vector: set the dimension of the space to d, there are N pixel points in the human image to be processed, denoted as pixel = 1,..., N, and the definition of the Mean-Shift vector of point x is:

[0067]

[0068] S h is a set of points, representing pixel points that satisfy the following conditions:

[0069] S h (x)≡{y:(y - x) T (y - x) ≤ h 2}

[0070] where k represents the number of pixel points that satisfy the above conditions among all pixel points, and S h can also be regarded as a high - dimensional sphere, and h is its radius. Then, create a spherical space with any point P0 as the origin, and S p and C r are the radii in the physical space and the color space respectively, where S p = 1 and C r = 70. Each point in the sphere has a color vector to the point P0. Calculate the sum of all vectors, and iteratively calculate the vector sum of each point. The termination condition of the iteration is that the end point of the vector sum is exactly the center point P N .

[0071] Optionally, in the above step S3, the step of determining the effective domain of each pixel point on the edge curve and determining the contour points from all pixel points on the edge curve according to the effective domain includes:

[0072] S301. Determine the discrete curvature of a pixel point P i on the edge curve according to the expression y = f(x) of the edge curve and a curvature detection algorithm:

[0073] S(p i ) = f i+1 - f i

[0074] where P i represents the i - th pixel point on the edge curve, and S(P i ) represents the discrete curvature of the pixel point P i .

[0075] S302. Take the pixel points P i adjacent to the pixel point P i-k and P i+k on the edge curve, and calculate the length l when k = 1, 2, 3...K respectively ik and the vertical distance d i from P ; where, ik represents the pixel point P and P i-k , P i+kThe chord between.

[0076] Optionally, in step S302,

[0077] S303. When the length l ik and P i to the vertical distance d ik satisfies the preset conditions, determine the effective domain radius and the effective domain D(P i corresponding to the pixel point P i ) according to the maximum value k i ) of k;

[0078] Specifically, the preset conditions include a first preset condition and a second preset condition; among them, the first preset condition is: l ik ≥l i,k+1 , and the second preset condition is:

[0079]

[0080] Optionally, when the length l ik satisfies the first preset condition and / or the length l ik and the vertical distance d i from P satisfies the second preset condition, determine the maximum value k ik of k as the radius of the effective domain of the pixel point P i , and determine the effective domain D(P i ) of the pixel point P according to the radius. i ) of P i ).

[0081] Among them, A and B respectively represent the first preset condition and the second preset condition.

[0082] S304. For a pixel point P i on the edge curve, when it satisfies the condition |S(P i )|≥|S(P j )|, retain the pixel point; otherwise, delete the pixel point and the pixel points with a discrete curvature of 0; where, |i - j|≤k i / 2;

[0083] S305. For the remaining pixel points on the edge curve, when the effective domain radius k i corresponding to the pixel point P i = 1 and there is P i or P i-1 in the effective domain D(Pi+1 , then delete the pixel points P that satisfy the condition |S(P i )| ≤ |S(P i-1 )| or |S(P i )| ≤ |S(P i+1 )|; i ;

[0084] S306. For the remaining pixel points on the edge curve, if there is a valid domain with more than 2 pixel points, then delete the pixel points P outside the endpoints of the valid domain i ; if there is a valid domain containing 2 pixel points, then when |S(P i )| > |S(P i+1 )|, delete the pixel point P i+1 , and when |S(P i )| < |S(P i+1 )|, delete the pixel point P i ;

[0085] S307. Determine the remaining pixel points as contour points.

[0086] In the above step S4, the step of calculating the area and perimeter of each contour respectively according to the multiple contours composed of contour points and determining a preset number of characteristic contours according to the area and perimeter includes:

[0087] S401. For the multiple contours composed of contour points, number them respectively and establish a network structure;

[0088] S402. Calculate the area s and perimeter l of each contour, and obtain two arrays: L = [l1, l2, K, l m and S = [s1, s2, K, s m , where m represents the number of the contour;

[0089] S403. When the corresponding elements in the two arrays satisfy the condition (l m ≥ a) ∪ (s m ≥ b), determine the contour numbered m as a characteristic contour; where a represents a preset perimeter threshold and b represents a preset area threshold.

[0090] Figure 3a is a result diagram of a contour detection method in the related art, Figure 3bIt is an example diagram of the feature contour provided by the embodiment of the present invention. It should be noted that in this embodiment, ultimately 2-4 feature contours need to be determined. The perimeter threshold a and the area threshold b can be flexibly set according to actual needs. If the number of determined contours is greater than 4, the perimeter threshold a and the area threshold b are increased, and the adjustment step size is 1; conversely, if the number of determined contours is less than 2, the perimeter threshold a and the area threshold b are correspondingly decreased. Through such iterative calculation, the feature contours of the portrait can be obtained, and the interference contours inside and outside the portrait can be excluded to a certain extent.

[0091] Figure 4 It is a schematic diagram of the semantic segmentation network provided by the embodiment of the present invention. Figure 5 It is a schematic diagram of a Block in the MobileNetV3 network provided by the embodiment of the present invention. Figure 6 It is a schematic diagram of the Last Stage structure in the MobileNetV3-mid network provided by the embodiment of the present invention. Please refer to Figure 4-6 , in this embodiment, the semantic segmentation network includes: a first encoder, an Atrous Spatial Pyramid Pooling (ASPP) module, a first decoder, and a Softmax layer. The first encoder includes multiple Blocks of the MobileNetV3-mid network; among them,

[0092] The first encoder is used to generate a low-level feature stream and generate a sub-feature map according to the input feature contour and the person image to be processed;

[0093] The ASPP module is used to extract features of different scales in the sub-feature map and generate a feature map according to the sub-feature map and the features of different scales;

[0094] The first decoder is used to restore the first image according to the feature map and the low-level feature stream, and the size of the first image is the same as that of the person image to be processed;

[0095] The Softmax layer is used to classify each pixel point in the first image into the foreground area, the background area, or the unknown area and output a ternary map.

[0096] In this embodiment, for the semantic segmentation network, a MobileNet Block with an SE module is added, and its network structure is as Figure 4 shown. Among them, the input image feature block first performs a 1*1 pointwise convolution, and then performs a 3*3 depthwise convolution. After each convolution, BN and the activation function ReLU6 operations are used for normalization and non-linear processing. Subsequently, an attention weight is added through an SE Block. The SE Block mainly performs two operations: Squeeze (compression F ex ) and Excitation (excitation F sq ).

[0097] In MobileNetV3, the activation function of the block adopts the non-linear activation function NL, and the hard-swish activation function is used in the second fully-connected layer. Its calculation formula is:

[0098]

[0099] It should be noted that in this embodiment, the ReLU6 activation function with higher precision is still used in the Blocks in the first half of the MobileNetV3 network, and the h-swish activation function is used in the second half.

[0100] Furthermore, the overall structure of the MobileNetV3 network is shown in Table 1. Taking the resolution of the to-be-processed person image as 320*320 as an example, when it is input into the network, since the number of convolution kernels is reduced from 32 to 16 in the first convolution, the change in the number of channels at the head of the network structure can be avoided, the corresponding number of parameters can be reduced, about 2 ms of time can be saved, and there is almost no impact on the accuracy. In Table 1, "input" represents the shape of the feature map input to the current layer, "out" represents the output channel size, "Bottleneck" represents a block of the MobileNetV2 network, "SE" means whether there is an SE Block, "NL" is the type of activation function, and "s" is the stride. The size of the convolution kernel of the DW convolution is connected behind the Bottleneck, and "NBN" means not using the normalization operation.

[0101] Table 1

[0102]

[0103]

[0104] Compared with the existing MobileNetV3-large network, the MobileNetV3-mid network provided by the present invention has 2 fewer Bottlenecks, but the number of times of feature extraction has not decreased, and it is still 5 times. It just uses one less Bottleneck during the two feature extractions respectively. In the image semantic segmentation task, the speed of MobileNetV3-mid has been effectively improved compared with MobileNetV3-large, and the loss of accuracy can be ignored.

[0105] Furthermore, as shown in Table 2, when using the MobileNetV3-mid network structure as the backbone network in this embodiment, the ASPP module is added to capture the multi-scale features of the image, so the Last Stage can be streamlined. The streamlined LastStage structure is as Figure 6As shown, after the first Conv convolution, the pooling layer and subsequent convolution operations are discarded, and the above operations are directly performed in the subsequent ASPP.

[0106] As shown in Table 2, the feature map output by the first 1×1 convolution of the structure is the deep feature output of 10×10×480. It is input into the ASPP module. The ASPP module can extract features of different scales. The dilated convolution used therein expands the receptive field of the network, avoiding the problem of information loss during the pooling process. The dilated convolution enables the output of the convolution to contain more extensive and larger information.

[0107] Specifically, the ASPP module used in the present invention combines a 1×1 convolution, three dilated convolutions with dilation rates of 6, 12, and 18, and a global information layer in parallel. The output channels of each branch are all 256. Finally, the outputs of these parallel layers are concatenated together, and then a 1×1 convolution is connected for output. The ASPP module uses depthwise separable dilated convolutions with different expansion rates for feature extraction. The design of the 5 parallel modules is shown in Table 2. To accelerate the network convergence speed, in the improved ASPP module of this embodiment, batch normalization BN and ReLU activation functions are used after each convolution layer. The outputs of the five building blocks are respectively denoted as a0, a1, a2, a3, and a4. The above five building blocks are superimposed, that is, a0 + a1 + a2 + a3 + a4. After superimposition, the size is 10×10×1280. Then, the number of channels is changed to 256 through a 1×1 convolution. Finally, the output is 10×10×256, which is y. By improving the ASPP module, the present invention can capture more global information using a larger sampling rate, which helps to improve the segmentation effect.

[0108] Table 2

[0109]

[0110] Table 3

[0111]

[0112]

[0113] As shown in Figure 3, the semantic segmentation network further includes a first decoder. For the first decoder, as shown in Table 3, first, the low-level feature stream of 80*80*24 from the MobileNetV3-mid module is processed. The number of channels is adjusted to 64 through a 1*1 convolution, followed by BN and ReLU operations. Then, the output y of the ASPP module is upsampled to enlarge the feature map, and the size is changed to 80*80*256 using resize_images. Next, the two are superimposed, and the superimposed feature map is subjected to two 3*3 convolutional kernels and two 1*1 convolutions. Finally, the size is changed to 320*320*3 using resize_images.

[0114] For the obtained feature map of 320*320*3, the softmax layer reshapes it to a size of 102400*3, and then performs a three-classification operation on 102400 pixel points to determine whether each pixel point belongs to the background area, the foreground area, or the unknown area, completing semantic segmentation.

[0115] Exemplarily, the formula processed by the softmax function is:

[0116]

[0117] In the formula, y i represents the output image block of size 102400*3. After softmax calculation, its probability distribution is obtained. n represents the number of classes, and here n is 3.

[0118] In this embodiment, three colors, white, black, and gray, are used to represent the three categories in the ternary graph. By traversing each pixel point of the person image to be processed, the probability of which category the pixel point belongs to is calculated to be the highest, and then the pixel point is classified into the corresponding area.

[0119] Figure 7 is a schematic diagram of the fine segmentation network provided by an embodiment of the present invention. Figure 8 is a schematic diagram of each Block module in the fine segmentation network provided by an embodiment of the present invention. Figure 9 is a schematic diagram of the CAM module provided by an embodiment of the present invention. Please refer to Figure 7-9 , the fine segmentation network includes a U-net and a context attention CAM module. The U-net includes a second encoder and a second decoder; among them,

[0120] The second encoder is used to extract the transparency feature of the person image to be processed and generate an Alpha feature stream after obtaining the ternary graph and the person image to be processed.

[0121] The CAM module is used to predict the transparency of the pixels classified into the unknown region according to the transparency of the pixels classified into the foreground region and the background region.

[0122] The second decoder is used to restore the original resolution of the person image to be processed and predict the transparency mask map according to the Alpha feature stream, so as to obtain the segmentation result of the person image to be processed.

[0123] Specifically, the structure of the fine segmentation network is as Figure 7 shown, and the processing process of the image therein is as follows:

[0124] Use the second encoder of the fine segmentation network to extract the features of the person image to be processed. The overall structure of the fine segmentation network adopts U-net, which is shaped like the letter U, and the downward part is the second encoder. Optionally, the second encoder has five layers, and each layer is short-circuited through a short cut block. In this way, the second decoder will combine the features of the second encoder before the upsampling block, which can avoid more convolutions on the feature map after the second encoder. This design method can preserve more low-level features of the image and is helpful for retaining more detailed texture information when predicting and generating alpha matting. In addition, in this embodiment, two layer short cut blocks are used to align the channels of the second encoder feature map to achieve feature fusion. In addition, the original input of the image is directly connected to the last convolutional layer of the network after passing through the short cut, so that the features will not generate any computational operations in the backbone network, thus well preserving the detail features and gradient information.

[0125] The structure of each block is as Figure 8 shown. The short cut block module consists of two 3*3 Stride Conv; the Res Block consists of two 3*3 ordinary convolutions and one 1*1 convolution; the four downsampling convolution blocks Res Down Block consist of two 3*3 convolutions, one average pooling operation and one 1*1 convolution to extract the depth features of the image.

[0126] The structure of the CAM module is as Figure 9 shown. The CAM module has two input feature streams, namely the alpha prediction feature stream and the low-level feature stream. The CAM module is used to predict the opacity of the unknown region according to the information around the foreground region and the background region. The CAM module performs affinity on the learned low-level features and transfers the high-level opacity information based on this appearance information, and can effectively use the rich features generated by the network.

[0127] It should be noted that the low-level feature stream is the feature information of the intermediate features of the semantic segmentation network after passing through the Feature change module, while the Alpha prediction feature stream is the information flow transmitted by the second encoder in the fine segmentation network, including unknown regions and known regions. Since the low-level feature stream is the same size as the Alpha prediction feature stream, in the low-level feature stream, the information of the Alpha prediction feature stream is used to divide the detailed features into two parts, namely known regions (foreground regions, background regions) and unknown regions. A 3*3 window is used to cut the entire image feature stream, and then the cut feature map blocks are reshaped into one-dimensional features, which are regarded as convolution kernels to calculate the similarity with the feature map of the unknown region:

[0128]

[0129] In the formula, U x,y represents the feature map of the unknown region centered at (x,y), I x′,y′ represents the image feature block centered at (x′,y′), S (x,y),(x′,y′) represents the similarity between the two regions, U x,y ∈μ is also an element in the image feature block set τ (μ∈τ), and the constant λ is the penalty hyperparameter, set to -10 4 , which can avoid too large similarity between the unknown region and its location.

[0130] Starting from (x′,y′), perform scaled softmax according to the following formula to obtain the attention weight attention score of each feature block:

[0131] a (x,y),(x′,y′) = soft max(ω(μ,κ,x′,y′)s (x,y),(x′,y′) )

[0132]

[0133] clamp(φ) = min(max(φ,0.1),10)

[0134] Among them, ω() represents the weight function, and κ = I - μ is the known region block in the image feature block set. However, the size of the unknown region in the ternary map is uncertain, and there may be a large number of unknown regions. Therefore, in this embodiment, the size of each feature block weight is set according to the size of the known region, and the designed function is

[0135] When the known region is relatively large, the feature blocks in the known region can carry more and more accurate detailed information to distinguish the foreground and background. At this time, the feature blocks in the known region can be weighted with large weight values; conversely, if the known region is relatively small, the feature blocks in the known region will only provide a small amount of appearance information, which is not enough to predict the opacity. Therefore, a smaller weight value can be given to the feature blocks in the known region. At the same time, the Alpha prediction feature stream is also divided in the same way and reshaped into a convolution kernel, and then a deconvolution operation is performed. After that, the unknown region will be reproduced on the alpha feature map, and the final result is fused with the original alpha feature and passed down for training.

[0136] Furthermore, the second decoder of the fine segmentation network is used to restore the image size. The second decoder is the upward part of the U-shaped network and also has 5 layers. Please continue to refer to Figure 10 , where the upsampling convolution block Res Block up of the first four layers of the decoder consists of two 3*3 convolutions, a Nearest Upsample operation, and a 1*1 convolution. The features output by the upsampling module are processed by a Res module, and then an element-wise feature addition operation is performed with the underlying information from the short cut block. The role of the Feature change module is to adjust the size and channels of the intermediate features, with the aim of obtaining the same size and number of channels as the other input feature map of the CAM module. The intermediate feature of the semantic segmentation network is 80*80*24, and only one 3*3 convolution operation with a Stride of 2 is needed to output the size as 40*40, and then a 1*1 convolution operation is performed to adjust the number of channels to 64. If the intermediate feature of 80*80*24 is not used, three 3*3 convolutions with a Stride of 2 need to be performed on the original image and then the channels are adjusted, which increases the computational burden of the network. Such a design makes full use of the semantic segmentation network, strengthens the connection between the two sub-networks, and reduces unnecessary computational redundancy. The last layer of the decoder consists of a Deconv and a Conv operation. All convolution operations in the figure are added with spectral normalization SN and batch normalization operation BN to accelerate convergence.

[0137] It should be noted that the fine segmentation network predicts the depth alpha feature around the human portrait contour. Therefore, the supervision signal used in the training process is the alpha mask map of the unknown region in the ternary map, which is defined as the absolute difference between the alphamatting of the unknown region and the Ground Truth. The loss function of the fine segmentation network is:

[0138] L d =m μ ||α p -αg ||1

[0139] Among them, m μ is a binary mask indicating whether it is an unknown area of a three - valued graph. m μ = 1 indicates that the pixel point is in the unknown area, and m μ = 0 indicates that the pixel point is in the known area, and α p represents the predicted alpha matting, and α g represents the alpha matting of GroundTruth, and m μ blocks the determined area, playing an attention role, which can effectively simplify the process of loss calculation and feedback, and m μ the prediction of the area where m = 0 may not be accurate enough. Especially for the internal details of the portrait foreground, the prediction is too difficult and will consume a large amount of computing resources, affecting the segmentation efficiency of the entire network.

[0140] The portrait fine - segmentation algorithm based on feature contours of the present invention consists of a semantic segmentation network and a fine - segmentation sub - network. Each network has its own responsibilities, with independent prediction tasks and loss functions, and is closely connected at the same time, which is an end - to - end calculation and training process. Therefore, a joint loss function is required to calculate the loss of the entire network. This loss Loss is the weighted sum of the loss function L t of the semantic segmentation network and the loss function L d of the fine network:

[0141] Loss = εL t +(1 - ε)L d

[0142] Among them, ε represents the weight of the loss function, which takes the value of 0.1 during training, emphasizing the fineness of the alpha matting map while using a high - quality three - valued graph as a constraint.

[0143] Next, the present invention further illustrates the above - mentioned method for segmenting human - body images based on feature contours through experiments.

[0144] Figure 10 is a comparison graph of the generation results of the three - valued graph provided by the embodiment of the present invention. Refer to Figure 10, the experimental results of generating a ternary map by the semantic segmentation network based on feature contours of the present invention are analyzed as follows: The first column is the human image to be processed, the second column is the Ground Truth of the images in the first column, and the third column is the ternary map obtained by performing erosion and dilation operations on the second column. It can be seen that the area distribution of the unknown regions in this ternary map is relatively balanced and only related to the size of the convolution kernel. For some relatively complex contour regions (such as hair, clothing edges, etc.), more regions need to be marked as unknown. At this time, if the convolution kernel of the morphological operation is increased, it will cause the global unknown regions of the ternary map to expand, affecting the efficiency and accuracy of subsequent fine segmentation. The fourth column is the ternary map after semantic segmentation and reclassification using DeeplabV3+. It can be seen that the edges of the human figure at this time are not complete and clear enough, and there are problems with missing detail features. Compared with the fifth column, more positions are predicted as unknown regions. The fifth column is the ternary map generated by the semantic segmentation network proposed by the present invention. It can be seen that after generating the preprocessing using feature contours, the area of the unknown regions in the ternary map is smaller, and the edges of the ternary map are clearer, with fewer rough and misjudgment situations. Coupled with the attention mechanism and the ASPP module, the accuracy and segmentation efficiency of high-frequency features are improved, resulting in the quality of the ternary map in the fifth column being significantly higher than that in the fourth column. Manually annotate the ternary maps in the dataset using photoshop, as shown in Figure 10 shown. When annotating, most of the unknown regions use 10 pixel units, and both the annotated area and accuracy can be artificially controlled, but the overall operation efficiency is very low. The average time for operating an image using photoshop is 3 minutes, while the average speed of generating a ternary map using the automated semantic segmentation network of the invention is 22.3 fps. The generation speed at the millisecond level is incomparable to manual operation.

[0145] As shown in Table 4, the present invention uses the mean intersection over union (MIoU) and the percentage of unknown region pixels in the image (UP) to analyze and verify the data:

[0146] Table 4

[0147]

[0148] The ternary map generated by the Feature Contour-based Semantic Segmentation Network (FCSS) used in the present invention is the closest to the manually annotated one. The average proportion of the area of the unknown region in the ternary map is 3.88%, which is nearly 30% less than that of the ternary map generated by the DeeplabV3+ algorithm. Taking the manually annotated ternary map as the Ground Truth, the intersection over union (IoU) of the effects of each algorithm and the Ground Truth is calculated. The foreground intersection over union (fgIou), background intersection over union (bgIou), unknown region intersection over union (unIou), and mean intersection over union (MIoU) of the semantic segmentation algorithm based on feature contours in this paper are all the highest, and the classification performance of the algorithm is significantly improved. The number of parameters of the DeeplabV3+ algorithm using MobileNetV3-large as the backbone network is 5.6M. The present invention proposes MobileNetV3-mid and designs the semantic segmentation network FCSS with MobileNetV3-mid as the backbone network. The number of parameters of the entire network model is only 4.8M, and the processing efficiency is also improved. The feature contour generation preprocessing algorithm proposed by the present invention can well predict most of the contour edges. Coupled with the attention module introduced in the semantic segmentation network design, the edge information can be accurately located and strengthened. By fusing information of different sizes through ASPP, the classification performance is also greatly improved. From both subjective and objective evaluation indicators, the semantic segmentation network based on feature contours proposed by the present invention has a better segmentation effect, and the generated high-quality ternary map provides a good foundation for subsequent fine segmentation.

[0149] Figure 11 It is a comparison diagram of the human image segmentation results provided by the embodiments of the present invention. Refer to Figure 11 , the overall effect analysis of the feature contour-based human portrait segmentation is as follows: Two datasets, Deep Image Matting and PM-100, are used for model training and verification. Figure 11 In, the first column is the human portrait image to be processed; the seventh column is the Ground Truth of the original image segmentation map; the second column is the segmentation effect of DCNN; the third column is the segmentation effect of MODNet; the fourth column is the segmentation effect of the Deep ImageMatting algorithm, which uses the ternary map generated by erosion and dilation as the input. Using Deeplabv3+ as the semantic segmentation network to generate the ternary map, and then combining with the CAM-based fine segmentation network of the present invention, the segmentation effect is shown in the fifth column. The sixth column is the segmentation effect diagram of the algorithm of the present invention. The small red-framed box in the segmentation effect diagram is the magnified detail diagram. The present invention selects the regions with obvious distinguishability in each group of images for display. From Figure 11As can be seen from the figure, the accuracy of segmentation increases from left to right, and the visual effect is gradually improved. For close-up high-definition portraits, there are many fine features such as hair, and the segmentation effect of the above algorithms can reach the hair level. However, there are still big differences in the amount of details that the algorithms can retain. The portrait edge details obtained by the DCNN algorithm are the least, and the portrait features are easily lost as a whole block, and the stability is insufficient. For the segmentation map using the MODNet algorithm, since the ternary map is not used as input, the segmentation effect is slightly worse, the fine edge features are relatively fuzzy, multiple hairs are often predicted as a fuzzy hair, there are many fuzzy shadows, and the areas with more dark parts such as shadows will be judged as portraits, and the segmentation effect will also be unstable. The segmentation effect of the Deep Image Matting algorithm is significantly better than the previous algorithms, and the detailed features are also the most retained, and the edges of the portraits are very sharp and clear. However, through careful observation, it is found that its alpha matting map has the phenomenon of over-segmentation. Some of the segmented hairs are thicker than those in the Ground Truth, the details are too sharp and not soft enough, and some features inside the portrait are often judged as backgrounds, which has certain misjudgment problems. Especially for semi-transparent wedding images, the transparency judgment is not as good as the algorithm of the present invention, and the overly complex network structure uses a lot of parameters, which will cause some images with a resolution of more than 2500*4000 in the verification set to burst the GPU memory when processing. The segmentation effect of the fifth column is close to that of the sixth column, and the details it retains are slightly less than those in the sixth column. Under a more complex background, the edge of the portrait outline is misjudged as the background, resulting in insufficient edge continuity. Overall, the segmentation effect of the algorithm of the present invention is the closest to the Ground Truth, with clear edges, more details retained, and the portrait is neither blurred nor too stiff, and there is rarely any over-segmentation. In addition, the portrait segmentation effect is relatively stable for different object distances, and it has strong adaptability to different environments such as too dark or too bright.

[0150] As shown in Table 5, the present invention calculates the MSE, SAD, Gradient, and Connectivity values ​​in the validation set, and the objective analysis of the segmentation effect is as follows:

[0151] Table 5

[0152] Algorithm Trimap <![CDATA[MSE(×10 -1 )]]> SAD Grad Conn DIM-Trimapless - 0.110 70.31 70.06 70.05 DCNN - 0.079 122.40 129.57 121.80 MODNet - 0.041 50.05 42.31 - DIM+DE √ 0.014 50.04 31.00 50.08 DIM+FCSS √ 0.013 46.03 36.32 49.22 Alpha GAN √ 0.030 52.40 38.00 - Ours+Deeplabv3+ √ 0.0108 41.12 22.46 36.74 Ours+FCSS √ 0.0091 35.22 18.81 44.23

[0153] Ours represents the CAM-based fine segmentation network proposed in this paper. Figure 11It can be seen that the segmentation algorithm proposed by the present invention is superior to the algorithm of the control group in terms of evaluation indicators. Whether the algorithms in the figure use ternary maps also has a great influence on the segmentation accuracy. Among them, DIM-Trimapless means that after the network structure of the Deep Image Matting algorithm deletes the trimap channel, the network is allowed to directly predict the alpha value, and the accuracy of the algorithm decreases significantly. DIM+DE means that the DIM algorithm uses the ternary map generated by corrosion and expansion as input; while DIM+FCSS uses the semantic segmentation algorithm proposed by the present invention to generate a ternary map and input it into the DIM network; Alpha GAN is also a segmentation algorithm that requires a ternary map. The MSE and SAD of the MODNet algorithm, which does not require a ternary map by design, perform better than other algorithms without ternary maps. However, if the semantic segmentation network FCSS proposed in this article is added to MODNet, that is, the ternary map channel is added during network training, the effect will also be significantly improved. In the PPM-100 data set, MSE is reduced by about 50%. At the same time, the quality of the ternary graph input during training will also affect the accuracy of segmentation. For comparison, the ternary graph generated by the Deeplabv3+ algorithm is input into the DIM network and the CAM-based fine segmentation network, and the MSEs are 0.013 and 0.0112, respectively. However, after using the ternary graph generated by the FCSS algorithm of the present invention as input, the MSE on the PPM-100 data set is reduced by 18.8% and 12.5%, respectively. The fine segmentation algorithm (Ours) of the present invention reduces the MSE by an average of 27% relative to the DIM algorithm, and Ours+FCSS has the best segmentation effect in these groups of experiments. The experimental results under the two data sets are relatively similar, and both can objectively prove the advantages and disadvantages of the algorithms.

[0154] It can be seen from the above embodiments that the beneficial effects of the present invention are:

[0155] The present invention provides a method for segmenting a human image based on a feature contour. Firstly, a priori generation of a human feature contour is performed, and the human image to be processed is marked with the feature contour. Then, a semantic segmentation network is used to classify and predict a ternary image, and a fine segmentation network is used to segment high-quality alpha matting. The method can automatically generate a high-quality ternary image, which not only improves the generation efficiency of the ternary image and improves the fineness of the ternary image, but also helps to make the subsequent human image segmentation more accurate.

[0156] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0157] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0158] Although the present application has been described in connection with various embodiments, however, in the process of implementing the claimed present application, those skilled in the art can understand and realize other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality of cases. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0159] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0161] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for segmenting a human image based on a feature contour, characterized in that Including: Obtain a person image to be processed; Perform Mean-Shift preprocessing on the person image to be processed to obtain an edge curve; Determine the effective domain of each pixel point on the edge curve, and determine contour points from all pixel points of the edge curve according to the effective domain; Calculate the area and perimeter of each contour respectively according to the multiple contours composed of the contour points, and determine a preset number of characteristic contours according to the area and the perimeter; Input the person image to be processed and the characteristic contours into a pre-trained semantic segmentation network, so that the semantic segmentation network generates a low-level feature stream, generates a feature map according to the characteristic contours and the person image to be processed, further restores a first image according to the feature map and the low-level feature stream, and classify the pixel points in the first image into a foreground area, a background area or an unknown area to obtain a ternary map of the person image to be processed after semantic segmentation; Input the person image to be processed and the ternary map into a pre-trained fine segmentation network, so that the fine segmentation network predicts the transparency of the pixel points in the unknown area of the ternary map to obtain the segmentation result of the person image to be processed; The semantic segmentation network includes: a first encoder, an Atrous Spatial Pyramid Pooling (ASPP) module, a first decoder and a Softmax layer, and the first encoder includes multiple Blocks of MobileNetV3-mid network; wherein, The first encoder is used to generate a low-level feature stream and generate a sub-feature map according to the input characteristic contours and the person image to be processed; The ASPP module is used to extract different-scale features in the sub-feature map and generate a feature map according to the sub-feature map and the different-scale features; The first decoder is used to restore a first image according to the feature map and the low-level feature stream, and the size of the first image is the same as that of the person image to be processed; The Softmax layer is used to classify each pixel point in the first image into a foreground area, a background area or an unknown area and output a ternary map; The fine segmentation network includes a U-net and a Contextual Attention Module (CAM), and the U-net includes a second encoder and a second decoder; wherein, The second encoder is used to extract the transparency feature of the person image to be processed and generate an Alpha feature stream after obtaining the ternary map and the person image to be processed; The CAM module is used to predict the transparency of the pixel points classified into the unknown area according to the transparency of the pixel points classified into the foreground area and the background area; The second decoder is used to restore the original resolution of the person image to be processed and predict a transparency mask map according to the Alpha feature stream to obtain the segmentation result of the person image to be processed.

2. The method for segmenting a human image based on a feature profile according to claim 1, wherein The step of determining the effective domain of each pixel point on the edge curve and determining contour points from all pixel points of the edge curve according to the effective domain includes: Determine the discrete curvature of a pixel point P on the edge curve according to the expression y = f(x) of the edge curve and the curvature detection algorithm i : S(p i ) = f i+1 -f i where P i represents the i-th pixel point on the edge curve, and S(P i ) represents the discrete curvature of the pixel point P i . Take a pixel point P on the edge curve i adjacent pixel point P i-k 、P i+k , and calculate the length l of when k = 1, 2, 3... ik and the vertical distance d from P i to ; where, ik denotes the chord between pixel points P 、P i-k 、P i+k ; When the length l ik and the perpendicular distance d from P i to satisfy the preset conditions, the effective domain radius and the effective domain D(P ik corresponding to the pixel point P are determined according to the maximum value k of k i ; i ) i ; For a pixel point P on the edge curve i , when it satisfies the condition |S(P i )| ≥ |S(P j )|, retain this pixel point; otherwise, delete this pixel point and pixel points with a discrete curvature of 0; where |i - j| ≤ k i / 2; For the remaining pixel points on the edge curve, when pixel point P i corresponding effective domain radius k i = 1 and there exists P i in the effective domain D(P i-1 ), or P i+1 , then delete the pixel point P i that satisfies the condition |S(P i-1 )| ≤ |S(P i )| or |S(P i+1 )| ≤ |S(P i ); For the remaining pixel points on the edge curve, if there is a valid domain with more than two pixel points, delete the pixel points P outside the endpoints of the valid domain i ; if there is a valid domain containing two pixel points, when |S(P i )| > |S(P i+1 )|, delete the pixel point P i+1 , and when |S(P i )| < |S(P i+1 )|, delete the pixel point P i ; Determine the remaining pixel points as contour points.

3. The method for segmenting a human image based on a feature contour according to claim 2, wherein The preset conditions include a first preset condition and a second preset condition; among them, The first preset condition is: l ik ≥ l i,k+1 ; the second preset condition is:

4. The method for segmenting a human image based on a feature contour according to claim 3, wherein When the length l ik and P i to the vertical distance d ik meet the preset conditions, according to the maximum value k of k i determine the pixel point P i corresponding effective domain radius and effective domain D(P i ) steps, including: When the length l ik satisfies the first preset condition and / or the length l ik and the vertical distance d from P i to satisfies the second preset condition, the maximum value k ik of k is determined as the radius of the effective domain of the pixel point P i , and the effective domain D(P i ) of the pixel point P is determined according to the radius i . i ) 5. The method for segmenting a human image based on a feature profile according to claim 2, wherein The step of respectively calculating the area and perimeter of each contour according to the multiple contours formed by the contour points, and determining a preset number of characteristic contours according to the area and the perimeter includes: For the multiple contours formed by the contour points, number them respectively and establish a network structure; Calculate the area s and perimeter l of each of the said contours to obtain two arrays: L = [l1, l2,..., l m and S = [s1, s2,..., s m , where m represents the contour number; When the corresponding elements in two arrays satisfy the condition (l m ≥ a) ∪ (s m ≥ b), the contour numbered m is determined as the feature contour; where a represents a preset perimeter threshold and b represents a preset area threshold.

Citation Information

Patent Citations

  • Semantic segmentation method based on improved PSPNet

    CN112365514A

  • Complex background image matting method based on deep learning semantic segmentation

    CN112749624A