An object segmentation method based on a non-local feature aggregation neural network

Through the target segmentation method based on non-local feature aggregation neural network, the problem of long time of manual calibration data set and the limitations of RoI feature extraction of traditional instance segmentation algorithms is solved, and the target segmentation effect with high precision, low data volume and strong environmental adaptability is achieved.

CN115223080BActive Publication Date: 2025-06-10NANJING NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210831695.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-15
Publication Date
2025-06-10
Estimated Expiration
2042-07-15

AI Technical Summary

Technical Problem

In the prior art, manual calibration data sets are time-consuming, training an instance segmentation model requires a large amount of data and a long training time, and the RoI feature extraction part in the traditional instance segmentation algorithm has limitations.

Method used

The target segmentation method based on non-local feature aggregation neural network is adopted, and the target segmentation network is optimized by building a non-local feature aggregation neural network model, combining the 80-classified pre-trained model of the coco dataset, training the target segmentation network, and optimizing the key points of the target outline using the BAS-DP lightweight algorithm.

Benefits of technology

The target segmentation method is realized with high accuracy, small amount of training data, strong environmental adaptability and good segmentation effect, reducing manual labeling costs and improving the performance and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223080B_ABST
    Figure CN115223080B_ABST
Patent Text Reader

Abstract

The present invention discloses an object segmentation method based on a non-local feature aggregation neural network, including: collecting a target video, extracting the original frame images of the video to obtain a small sample data set; building a non-local feature aggregation neural network model, training the non-local feature aggregation neural network model to obtain an object segmentation network; collecting the target video again, calculating the segmentation quality evaluation score of each object in the image; judging the image quality according to the segmentation quality evaluation score, continuing to train the low-quality images, and retaining the target contour key points in the high-quality images; optimizing the target contour key points through the BAS-DP lightweight algorithm to obtain the object segmentation result. The present invention has the advantages of high accuracy, small amount of training data, strong environmental adaptability and good segmentation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning and computer vision, and particularly relates to an object segmentation method based on a non-local feature aggregation neural network. Background Art

[0002] With the development of the modernization process, big data, as a new and potentially valuable asset, is greatly influencing and changing fields such as agriculture, commerce, finance, and healthcare. With the development of deep learning, the number of layers of neural networks is getting deeper and deeper. The generalization of models depends on the amount of data they are trained on, and more training data is needed to support the learning of the models. Therefore, obtaining enough labeled data is a major challenge today. Currently, large image datasets include datasets such as ImageNet and COCO, and these image datasets are all labeled for the objects in the images manually. This method is time-consuming and laborious, and the quality of the data marked under a fatigued state is relatively low. It is difficult to quickly and conveniently obtain a dataset with high quality, large quantity, and meeting the requirements of actual scenarios only through manual annotation. Semi-supervised learning uses a large amount of unlabeled data and simultaneously uses a small amount of labeled data to perform pattern recognition work. Adopting semi-supervised learning requires minimizing the cost of manual annotation while bringing relatively high accuracy. Therefore, semi-supervised learning has been receiving increasing attention.

[0003] Instance segmentation is a classic task in the field of computer vision. It divides the entire image into pixel groups and then labels and classifies them.

[0004] He Kaiming et al. proposed the Mask R-CNN algorithm based on Faster R-CNN. This algorithm adds a mask branch to Faster R-CNN to predict the mask, and uses RoIAlign instead of RoI Pooling to extract the features of the region of interest (RoI), thereby improving the mask segmentation quality. However, the segmentation quality score and the classification score are shared. In fact, there will be a deviation when using the classification score to evaluate the segmentation quality. Then Huang et al. proposed a segmentation quality scoring mechanism, using the intersection over union (MaskIoU) of the predicted mask and the ground truth mask to measure the segmentation quality. A MaskIoU Head branch is added based on Mask R-CNN, and the MaskIoU is multiplied by the classification score to obtain the mask score, so that the mask output by the network is more complete. However, there is a lack of global information of the features before the features are input into the MaskIoU Head branch, and there is a large deviation between the predicted mask and the ground truth mask when the detected object is small. Summary of the Invention

[0005] Objective of the Invention: In order to overcome the deficiencies in the prior art, such as the long time for manually calibrating the data set, the large amount of data and long training time required for training the instance segmentation model, and the limitations in the RoI feature extraction part of the traditional instance segmentation algorithm, a target segmentation method based on a non-local feature aggregation neural network is provided, which has the advantages of high accuracy, small amount of training data, strong environmental adaptability, and good segmentation effect.

[0006] Technical Solution: To achieve the above objective, the present invention provides a target segmentation method based on a non-local feature aggregation neural network, including the following steps:

[0007] S1: Collect the target video, extract the original frame images of the video to obtain a small sample data set, manually segment all the targets, label the targets of different categories with masks of different colors, and divide the training set and the validation set;

[0008] S2: Build a non-local feature aggregation neural network model, and perform transfer learning in combination with the 80-class pre-trained model of the coco data set to train the non-local feature aggregation neural network model to obtain a target segmentation network;

[0009] S3: Collect the target video again, extract the frame images, input them into the target segmentation network, obtain the target categories, classification scores, mask contour key points, and mask scores in each image, and calculate the segmentation quality evaluation scores of each target in the image;

[0010] S4: Judge the image quality according to the segmentation quality evaluation scores. If the segmentation quality evaluation scores of all the targets in an image are greater than the set value, then this image is a high-quality image; otherwise, it is a low-quality image. The low-quality images are continuously trained, and the target contour key points in the high-quality images are retained;

[0011] S5: Optimize the target contour key points through the BAS-DP lightweight algorithm to obtain the target segmentation result.

[0012] Further, the method for building the non-local feature aggregation neural network model in step S2 is as follows:

[0013] A1: Input the obtained small sample data set into the backbone neural network of the non-local feature aggregation neural network model for processing to obtain the feature map of the input image;

[0014] A2: Input the obtained feature map into the Region Proposal Network (RPN) of the non-local feature aggregation neural network model to obtain the Region of Interest (RoI) candidates, and use RoIAlign to extract the RoI features and align the RoI features;

[0015] A3: Further process the RoI using the RoI feature aggregation network to extract the RoI global features on the feature map;

[0016] A4: Input the candidate region into the R-CNN Head network to obtain the target region class scores and the bounding box regression positions, and at the same time input it into the Mask Head network to obtain the predicted mask of the target, separate the target from the complex environment and predict its contour key points;

[0017] A5: Input the RoI features obtained in step A3 and the predicted mask obtained in step A4 into the MaskIoU Head network, regress the true mask and the predicted mask, calculate the intersection over union ratio of the predicted mask and the true mask, that is, the MaskIoU value, to obtain the target mask score, and complete the construction of the non-local feature aggregation neural network model.

[0018] Furthermore, the backbone neural network in step A1 includes a Residual Network ResNet50 and a Feature Pyramid Network FPN; the Residual Network ResNet50 consists of a 7×7 input convolution, followed by a 3×3 max pooling and then 16 residual blocks, each residual block is composed of three convolutional layers of 1×1, 3×3, and 1×1, so there are a total of 50 layers of network; the Residual Network ResNet50 is divided into 5 stages, and the outputs of stage1 to stage4 are four different-scale feature maps [C2, C3, C4, C5]. The feature pyramid has five layers. After extracting features from the first layer, they are passed layer by layer to the fifth layer, but the scale is halved layer by layer to generate different-scale feature maps, and then the adjacent feature maps are subtracted to obtain new feature maps, corresponding to the output of 5 different-scale feature maps in the Feature Pyramid Network FPN network.

[0019] Furthermore, the specific process of step A2 is as follows:

[0020] B1: Input the obtained feature map into the RPN to generate 9 target boxes with preset aspect ratios and areas for each position through a sliding window; after a convolution with a kernel size of 3×3, through a convolution with a kernel size of 1×1 and an output channel number of 36 and a convolution with a kernel size of 1×1 and an output channel number of 18 respectively. The result obtained by the former includes four values, namely the horizontal and vertical coordinates of the center point of the target box and the length and width of the target box. The result obtained by the latter is cropped and filtered and then passed through the softmax activation function to determine whether the target box belongs to the foreground or the background, and the coordinates of the target box belonging to the foreground are corrected to generate RoI;

[0021] B2: Align the RoI using RoIAlign; find 400 RoIs on the original image, map these RoIs back to the feature map, and use the following formula:

[0022]

[0023] Among them, w and h represent the width and height of the RoI respectively; k a is the scale of the feature layer to which this RoI should belong; k 0 is the scale of the feature layer mapped when w = 224 and h = 224, and it is generally taken as 4, that is, corresponding to P 4 ;

[0024] Divide the length and width of each RoI by the stride to obtain the image size of the RoI mapped to the feature map. If the size mapped to the feature map is a floating point number at this time, no rounding operation is performed and the floating point number is retained; if the region of interest on the feature map is to be aligned to 7×7, assuming the size mapped to the feature map is n×n, then the n×n region is divided into 49 parts, each part with a size of (n / 7)×(n / 7). Let the number of sampling points be 4, and then divide the (n / 7)×(n / 7) small region into four equal parts. Take the pixel at the center point position of each part, and use the bilinear interpolation method to calculate the pixel values of four points. Take the maximum value of the four pixel values as the pixel value of this small region, and so on. The 49 pixel values obtained from the 49 small regions form a feature map with a size of 7×7.

[0025] Furthermore, the specific process of step A3 is as follows:

[0026] The RoI feature aggregation network includes three sub-modules: a preprocessing model, an aggregation module, and a post-processing module; the preprocessing module is a convolutional layer with a convolutional kernel size of 5×5, which can further expand the receptive field of the feature map. Assume that the output after preprocessing is U = [u 1 , u 2 , u 3 , u 4 ∈ R H×W , where H and W are the height and width of the feature map respectively, and u 1 , u 2 , u 3 , u 4 are the features of the RoI mapped to four different scales of the original feature map [C2, C3, C4, C5] respectively. The result after aggregation operation is represented by the following formula:

[0027] X = sum(U)·ε(F s (U))

[0028]

[0029]

[0030] Among them, in the formula, m and n are the row and column respectively, k is the number of different scales, and the function F s(·) is an aggregation function, ε(·) is a ReLU activation function, sum(U) is the sum of RoI features at four scales. The final aggregated feature X is obtained by multiplying the sum of the results after global average pooling of the RoI features at four scales respectively by the result after summing the RoI features at four scales;

[0031] Due to the non-local property of the graph convolutional neural network (GCN), the present invention uses GCN as the non-local network of the post-processing module, where each graph node represents a single pixel on the feature map, and the output Z of the non-local module g can be expressed as:

[0032] Z g = σ(AXW g ) + X

[0033] where A represents the adjacency matrix of the adjacent relationship of graph nodes, W g is a weight-learnable output transformation matrix implemented by 1×1 convolution, σ(·) is a non-linear function composed of BatchNorm normalization and ReLU activation function, and a residual connection of the aggregated feature X output by an aggregation module is added;

[0034] The constructed adjacency matrix A is expressed as:

[0035] A = softmax(θ(x i )) T φ(x j ))

[0036] where x i and x j are two paired graph nodes, and each pair of graph nodes has pairwise similarity. θ(·) and are two trainable transformation functions implemented by 1×1 convolution. θ(x i ) and are the outputs after 1×1 convolution, and softmax is the activation function.

[0037] Furthermore, the specific process of step A4 is as follows:

[0038] C1: Input the obtained feature map into the R-CNN Head for object classification and location regression; the input shape of the R-CNN Head is the RoI feature of 7×7×256, followed by two fully connected layers of 1024 dimensions, and then use a fully connected layer of C dimensions to classify the object into C categories, and use the softmax activation function to obtain the scores of each category, and use a fully connected layer of 4 dimensions to obtain the regression of the object bounding box;

[0039] C2: Input the obtained feature map into the Mask Head to obtain the predicted mask of the target; the Mask R-CNN Head inputs RoI features with a shape of 14×14×256, followed by four convolutional layers, and then followed by a transposed convolutional layer to double the size of the feature map, i.e., 28×28×256. Then, use one convolutional layer to obtain a mask with features of 28×28×C, and use the sigmoid activation function to obtain the predicted mask.

[0040] Further, in step A5, the MaskIoU Head consists of three convolutional layers, one max pooling layer, and two fully connected layers. Concatenate the predicted mask with the RoI feature map and input it into the MaskIoU Head to obtain the target mask score, that is, the MaskIoU value of the target mask.

[0041] Further, step S3 specifically includes the following steps:

[0042] D1: Collect the target video again, extract one frame of image per second without manual annotation to obtain a large number of unlabeled images;

[0043] D2: Input all the unlabeled images into the target segmentation network to obtain the target category, classification score, mask contour key points, and mask score in each image. Multiply the classification score by the mask score to obtain the segmentation quality evaluation score for each target in the image. The segmentation quality evaluation score is represented by the following formula:

[0044] S mask =S cls ×S IoU

[0045] where S mask is the segmentation quality evaluation score, S cls is the classification score, and S IoU is the mask score.

[0046] Further, step S5 specifically includes the following steps:

[0047] E1: For the edge contour closed curve of the output image of the target segmentation network, first initialize the parameters of the Douglas–Peucker algorithm (DP) and set the distance threshold D threshold ;

[0048] E2: Calculate the two points M and N that are farthest apart on the closed curve, connect the two points M and N to form a line segment D MN , set the optimization function f(·) of the distance D q ;

[0049] E3: Initialization of parameters of the Beetle Antennae Search algorithm (BAS), including the initialization of the step size attenuation factor E ta , step size S tep , ratio v of the antennae, number of iterations n t and the number m of parameters to be optimized t ;

[0050] E4: Calculate the position x of the left antenna of the beetle l and its optimization function f(x l ), and the position x of the right antenna r and its optimization function f(x r ). The calculation formulas are as follows:

[0051]

[0052] In the formula, d ir is the direction of the random beetle, d 0 is the distance between the left and right antennae, and sign(·) is the sign function

[0053]

[0054] If f(x l ) > f(x r ), the beetle iterates to the next position x = x l , d = f(x l ), otherwise x = x r , d = f(x r );

[0055] E5: Loop and execute step E4 for n t times, and the optimized function value corresponding to the final position x of the beetle is the optimal solution;

[0056] E6: Find the point q on the remaining points of the closed curve that has the maximum distance from the line segment D MN , and calculate the distance D MN between the point q and the line segment D q ; Compare the size of this distance D q with the pre - given threshold D threshold . If D q is less than D threshold , then take the line segment D MN as the approximation of the curve, and this section of the curve is processed. If the distance D q is greater than D threshold , then use the point q to divide the curve into two segments M q and N q(Points M, N, and q are called key points), and the processing from step E3 to step E4 is performed on the two curves respectively; after all the curves are processed, the contour key points are connected in sequence to form a polygon as an approximation of the original closed curve, thereby improving the segmentation effect.

[0057] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0058] 1. The present invention proposes a non-local feature aggregation neural network, adding an aggregation module after RoIAlign to obtain a global feature with effective information of different scales of FPN, and using a graph convolutional neural network to learn this global feature, making full use of the effective information of RoI, thereby improving the performance and generalization ability of the model.

[0059] 2. Based on the longhorn beetle antennae algorithm and the Graham-Puk algorithm, the present invention proposes a BAS-DP lightweight algorithm to optimize the key points of the target contour, converting the contour information surrounding the target into the best polygon surrounding the target, and improving the segmentation quality.

[0060] 3. The present invention provides a target segmentation method, and still ensures the segmentation quality in complex environments such as occlusions, small targets, and multiple types.

[0061] 4. The present invention trains the model in a semi-supervised learning manner, with low implementation cost, small training data volume, strong environmental adaptability, and good segmentation effect. Description of the Drawings

[0062] Figure 1 is the overall flowchart of the present invention;

[0063] Figure 2 is the overall network structure diagram of the model of the present invention;

[0064] Figure 3 is the structure diagram of the backbone network and the feature pyramid network of the present invention;

[0065] Figure 4 is the residual module diagram of the backbone network of the present invention;

[0066] Figure 5 is the structure diagram of the region generation network of the present invention;

[0067] Figure 6 is the structure diagram of the RoI feature aggregation network of the present invention;

[0068] Figure 7 is the segmentation result diagram of the existing method Mask R-CNN;

[0069] Figure 8 is the segmentation result diagram of the existing method Mask Scoring R-CNN;

[0070] Figure 9 This is the segmentation result diagram of the present invention. Specific implementation manners

[0071] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.

[0072] The present invention provides a target segmentation method based on a non-local feature aggregation neural network, as Figure 1 shown, which includes the following steps:

[0073] Step 1: Collect the target video of the required scene, perform frame splitting on the video, extract one frame of image per second, manually segment the selected target, obtain a small sample dataset, and divide it into a training set and a validation set at a ratio of 8:2;

[0074] Step 2: Build a non-local feature aggregation neural network model, and perform transfer learning in combination with the 80-class pre-trained model of the coco dataset, and train the non-local feature aggregation neural network model to obtain a target segmentation network;

[0075] Step 3: Collect a video for a longer time again, extract one frame of image per second, input it into the target segmentation network, obtain the target category, classification score, mask contour key points and mask score in each image, and calculate the segmentation quality evaluation score of each target in the image;

[0076] Step 4: Judge the image quality according to the segmentation quality evaluation score. If the segmentation quality evaluation scores of all targets in an image are greater than 0.9, then this image is a high-quality image, otherwise it is a low-quality image. The low-quality images are continued to be trained, and the target contour key points in the high-quality images are retained;

[0077] Step 5: Propose a BAS-DP lightweight algorithm based on the beetle antennae search algorithm and the Graff-Purk algorithm to optimize the target contour key points, thereby improving the segmentation effect.

[0078] The structure of the non-local feature aggregation neural network model in step 2 of the method of the present invention is as Figure 2 shown, and the specific process of step 2 is as follows:

[0079] Step 2.1: Input the small sample dataset obtained in step 1 into the backbone network of the non-local feature aggregation neural network model for processing, so as to obtain the feature map of the input image;

[0080] Step 2.2: Input the feature map obtained in Step 2.1 into the Region Proposal Network (RPN) of the non-local feature aggregation neural network model to obtain Regions of Interest (RoIs), and use RoIAlign to extract and align RoI features;

[0081] Step 2.3: Further process the RoIs using the RoI feature aggregation network to extract RoI global features on the feature map;

[0082] Step 2.4: Input the candidate regions in Step 2.3 into the R-CNN Head network to obtain the target region class scores and the bounding box regression positions, and at the same time input them into the Mask Head network to obtain the predicted mask of the target, separate the target from the complex environment and predict its contour key points;

[0083] Step 2.5: Input the RoI features obtained in Step 2.3 and the predicted mask obtained in Step 2.4 into the MaskIoUHead network to regress the true mask and the predicted mask, calculate the intersection over union of the predicted mask and the true mask, that is, the MaskIoU value, to obtain the target mask score;

[0084] Step 2.6: After the non-local feature aggregation neural network model is built, perform transfer learning by combining the 80-class pre-trained model of the coco dataset to train the non-local feature aggregation neural network model to obtain the target segmentation network.

[0085] As Figure 3 shown, the non-local feature aggregation neural network model in Step 2.1 includes a backbone network and a feature pyramid network part. Step 2.1 is specifically as follows:

[0086] The backbone network ResNet50 consists of five stages from Stage0 to Stage4. Stage0 is the preprocessing part of the input image, which is a convolutional layer with a convolutional kernel size of 7×7, a stride of 2, and an output channel of 64, followed by a max pooling layer with a size of 3×3 and a stride of 2. Stage1 to Stage4 are respectively composed of 3, 4, 6, and 3 residual blocks. The first residual block in each stage is composed of Conv Block ( Figure 4 the right half part in Figure 4It consists of the left - middle part. Each stage outputs feature maps [C2, C3, C4, C5] of different sizes, which are used as the input of the Feature Pyramid Network (FPN). The sizes of C2 to C5 are reduced by 4, 8, 16, and 32 times respectively compared to the original image. C5 passes through a convolutional layer with a convolutional kernel size of 1×1, a stride of 1, and an output channel number of 256 to obtain M5. After C4 undergoes the same convolutional operation, the result is added to the result of upsampling M5 once to obtain M4. Similarly, M3 and M2 are obtained. After M2 - M5 pass through convolutional layers with a convolutional kernel size of 3×3, a stride of 1, and an output channel number of 256 respectively, P2 - P5 are obtained. After P5 passes through a max - pooling layer with a size of 1×1 and a stride of 2, P6 is obtained. [P2, P3, P4, P5, P6] are used as the input of the subsequent Region Proposal Network (RPN).

[0087] Step 2.2 is specifically as follows:

[0088] As Figure 5 shown, it is the RPN network part of the present invention. First, it passes through a convolution with a convolutional kernel size of 3×3, and then passes through a convolution with a convolutional kernel size of 1×1 and an output channel number of 36 and a convolution with a convolutional kernel size of 1×1 and an output channel number of 18 respectively. The result of the former includes four values, namely the horizontal and vertical coordinates of the center point of the target box and the length and width of the target box. The result of the latter, after being cropped and filtered, passes through the Softmax function to determine whether the target box belongs to the foreground or the background, and corrects the coordinates for the target boxes belonging to the foreground to generate RoIs. Then, the RoIs are aligned through RoIAlign and mapped back to the original feature map, which is used as the input of the subsequent RoI feature aggregation network.

[0089] Step 2.3 is specifically as follows:

[0090] As Figure 6 shown, it is the RoI feature aggregation network of the present invention. The RoI feature aggregation network includes three sub - modules: a pre - processing model, an aggregation module, and a post - processing module. The pre - processing model is a convolutional layer with a convolutional kernel size of 5×5, which can further expand the receptive field of the feature map. Assume that the output after pre - processing is U = [u 1 ,u 2 ,u 3 ,u 4 ∈R H×W , where H and W are the height and width of the feature map respectively, and u 1 ,u 2 ,u 3 ,u 4 are the features of the RoI mapped to four different scales of the original feature maps [C2, C3, C4, C5] respectively. The result after aggregation operation can be expressed by the following formula:

[0091] X = sum(U)·ε(F s(U))

[0092]

[0093]

[0094] Wherein, in the formula, m and n are the row and column respectively, k is the number of different scales, and the function F s (·) is an aggregation function, ε(·) is a ReLU activation function, sum(U) is the sum of the RoI features of four scales, and the final aggregated feature X is obtained by multiplying the sum of the results after global average pooling of the RoI features of four scales respectively by the result after summing the RoI features of four scales;

[0095] Due to the non-local property of the graph convolutional neural network (GCN), the present invention uses GCN as the non-local network of the post-processing module, where each graph node represents a single pixel on the feature map. The output Z of the non-local module g can be expressed as:

[0096] Z g =σ(AXW g ) + X

[0097] Wherein, A represents the adjacency matrix of the adjacent relationship of graph nodes. W g is a weight-learnable output transformation matrix, which can be implemented by 1×1 convolution. σ(·) is a non-linear function composed of BatchNorm normalization and ReLU activation function. And a residual connection of the aggregated feature X output by an aggregation module is added.

[0098] The constructed adjacency matrix A can be expressed as:

[0099] A = softmax(θ(x i )) T φ(x j ))

[0100] Wherein, x i and x j are two paired graph nodes, and each two graph nodes have paired similarity. θ(·) and are two trainable transformation functions, which are implemented by 1×1 convolution. θ(x i ) and are the outputs after 1×1 convolution, and softmax is the activation function;

[0101] Step 2.4 specifically includes the following steps:

[0102] Step 2.4.1: Input the feature map obtained in Step 2.2 into Figure 2In the R-CNN Head, object classification and location regression are performed. The R-CNN Head takes RoI features with an input shape of 7×7×256, followed by two fully connected layers with 1024 dimensions. Then, a fully connected layer with C dimensions is used to classify the object into C categories, and the softmax activation function is adopted to obtain the scores for each category. A fully connected layer with 4 dimensions is used to obtain the regression of the object bounding box;

[0103] Step 2.4.2: Input the feature map obtained in Step 2.2 into Figure 2 In the Mask Head, the predicted mask of the object is obtained. The Mask R-CNN Head takes RoI features with an input shape of 14×14×256, followed by four convolutional layers, and then followed by a transposed convolutional layer to double the size of the feature map, that is, 28×28×256. Then, a convolutional layer is used to obtain a mask with a feature of 28×28×C, and the sigmoid activation function is used to obtain the predicted mask.

[0104] Step 2.5 is specifically as follows:

[0105] Figure 2 The MaskIoU Head consists of 3 convolutional layers, 1 max pooling layer, and two fully connected layers. The predicted mask is input into a convolutional layer with a kernel size of 1×1, a stride of 1, and an output channel number of 1, and then followed by a max pooling layer with a size of 2×2 and a stride of 2 to change the feature map of the predicted mask to 14×14×1. This feature map is concatenated with the RoI features of 14×14×256 to obtain a feature map of 14×14×257 as the input of the MaskIoU Head, followed by three convolutional layers (the part marked as ×3 in the figure is omitted), where each convolutional layer has a kernel size of 3×3, a stride of 1, and an output channel number of 256, and then followed by a max pooling layer and two fully connected layers to output the MaskIoU of the C-class mask.

[0106] To verify the actual effect of the method of the present invention, in this embodiment, the segmentation method of the present invention is compared with the existing segmentation method, and the segmentation result diagrams for different methods as shown in Figure 7 、 Figure 8 、 Figure 9 are obtained. Figure 7 is the segmentation result diagram of the existing method MaskR-CNN. Figure 7 Among the three farthest people (the farther the target is, the smaller it is), the middle one is not recognized. The car accuracy is 0.85. It can be seen that when the target is small, the target cannot be accurately recognized, and the segmentation accuracy is low; Figure 8 is the segmentation result diagram of the existing method Mask Scoring R-CNN. In the figure, the segmentation bounding box of the right leg of the rightmost person among the three farthest people is inaccurate. The car accuracy is 0.92. It can be seen that when the target is small, there is a certain deviation between the segmentation bounding box and the actual target;Figure 9 For the segmentation method of the present invention, the segmentation bounding box is relatively accurate, the car accuracy is 0.98, the segmentation accuracy is further improved, and the segmentation quality is good. It can be seen that the segmentation method of the present invention has good effects.

Claims

1. A target segmentation method based on a non-local feature aggregation neural network, characterized in that, it includes the following steps: S1: Collect the target video, extract the original frame images of the video to obtain a small sample data set, manually segment all targets, label the targets of different categories with masks of different colors, and divide the training set and the validation set; S2: Build a non-local feature aggregation neural network model, and perform transfer learning by combining the 80-class pre-trained model of the coco data set to train the non-local feature aggregation neural network model to obtain a target segmentation network; S3: Collect the target video again, extract the frame images, input them into the target segmentation network, obtain the target categories, classification scores, mask contour key points and mask scores in each image, and calculate the segmentation quality evaluation scores of each target in the image; S4: Judge the image quality according to the segmentation quality evaluation scores. If the segmentation quality evaluation scores of all targets in an image are greater than the set value, then this image is a high-quality image, otherwise it is a low-quality image. The low-quality images are continuously trained, and the target contour key points in the high-quality images are retained; S5: Optimize the target contour key points through the BAS-DP lightweight algorithm to obtain the target segmentation result; The method for building the non-local feature aggregation neural network model in step S2 is: A1: Input the obtained small sample data set into the backbone neural network of the non-local feature aggregation neural network model for processing to obtain the feature map of the input image; A2: Input the obtained feature map into the Region Proposal Network (RPN) of the non-local feature aggregation neural network model to obtain the Region of Interest (RoI) candidates, and use RoIAlign to extract the RoI features and align the RoI features; A3: Use the RoI feature aggregation network to further process the RoI and extract the RoI global features on the feature map; The specific process of step A3 is: The RoI feature aggregation network includes three sub-modules: a preprocessing model, an aggregation module, and a post-processing module; The preprocessing module is a convolutional layer with a convolutional kernel size of 5×5, which can further expand the receptive field of the feature map. Assume that the output after preprocessing is U = [u 1 , u 2 , u 3 , u 4 ∈ R H×W , where H and W are the height and width of the feature map respectively, and u 1 , u 2 , u 3 , u 4 are the features of the RoI mapped to four different scales of the original feature map [C2, C3, C4, C5] respectively. The result after the aggregation operation is represented by the following formula: X = sum(U)·ε(F s (U)) Among them, in the formula, m and n are the row and column respectively, k is the number of different scales, and the function F s (·) is an aggregation function, ε(·) is a ReLU activation function, sum(U) is the sum of the RoI features at four scales. The final aggregated feature X is obtained by multiplying the sum of the results after global average pooling of the RoI features at four scales respectively by the result after summing the RoI features at four scales.

2. The target segmentation method based on a non-local feature aggregation neural network according to claim 1, characterized in that, the backbone neural network in step A1 includes a Residual Network (ResNet50) and a Feature Pyramid Network (FPN); the Residual Network ResNet50 consists of a 7×7 input convolution, followed by a 3×3 max pooling and then 16 residual blocks, each residual block is composed of a three-layer convolutional layer of 1×1, 3×3, and 1×1, so there are a total of 50 layers of network; the Residual Network ResNet50 is divided into 5 stages, and the outputs of stage1 to stage4 are four different-scale feature maps of [C2, C3, C4, C5]. The feature pyramid has five layers. After extracting features from the first layer, they are passed layer by layer to the fifth layer, but the scale is halved layer by layer to generate different-scale feature maps, and then the adjacent feature maps are subtracted to obtain a new feature map, corresponding to obtaining the output of 5 different-scale feature maps in the Feature Pyramid Network FPN network.

3. The target segmentation method based on a non-local feature aggregation neural network according to claim 1, characterized in that, the specific process of step A2 is: B1: Input the obtained feature map into the RPN to generate 9 target boxes with preset aspect ratios and areas for each position through a sliding window; after a convolution with a kernel size of 3×3, respectively pass through a convolution with a kernel size of 1×1 and an output channel number of 36 and a convolution with a kernel size of 1×1 and an output channel number of 18. The result obtained by the former includes four values, namely the horizontal and vertical coordinates of the center point of the target box and the length and width of the target box. The result obtained by the latter is cropped and filtered and then passed through the softmax activation function to determine whether the target box belongs to the foreground or the background, and the coordinates of the target box belonging to the foreground are corrected to generate RoI; B2: Use RoIAlign to align the RoI; Find 400 RoIs on the original image, map these RoIs back to the feature map, and use the following formula: where w and h represent the width and height of the RoI respectively; k a is the scale of the feature layer to which this RoI should belong; k 0 is the scale of the feature layer mapped when w = 224 and h = 224; Divide the length and width of each RoI by the stride to obtain the image size of the RoI mapped to the feature map. If the size mapped to the feature map is a floating point number at this time, no rounding operation is performed and the floating point number is retained; if the region of interest on the feature map is to be aligned to 7×7, assuming the size mapped to the feature map is n×n, then the n×n region is divided into 49 parts, each part with a size of (n / 7)×(n / 7). Let the number of sampling points be 4, and then the (n / 7)×(n / 7) small region is evenly divided into four parts, and the pixel value at the center point position of each part is taken. The pixel values of four points are calculated using bilinear interpolation, and the maximum value of the four pixel values is taken as the pixel value of this small region. By analogy, the 49 pixel values obtained from the 49 small regions form a feature map with a size of 7×7.

4. A target segmentation method based on a non-local feature aggregation neural network according to claim 1, characterized in that, In step A3, a GCN is used as the non-local network of the post-processing module, where each graph node represents a single pixel on the feature map, and the output Z of the non-local module g is expressed as: Z g = σ(AXW g ) + X Among them, A represents the adjacency matrix of the graph node adjacency relationship, and W g is an output transformation matrix with learnable weights, implemented by 1×1 convolution. σ(·) is a non-linear function composed of BatchNorm normalization and ReLU activation function, and an aggregated feature X of the output of an aggregation module is added with a residual connection; The constructed adjacency matrix A is expressed as: A = softmax(θ(x i ) T φ(x j )) where x i and x j are two paired graph nodes, and every two graph nodes have paired similarity. θ(·) and are two trainable transformation functions implemented by 1×1 convolution. θ(x i ) and are the outputs after 1×1 convolution, and softmax is the activation function.

5. A target segmentation method based on a non-local feature aggregation neural network according to claim 1, characterized in that, The method for building the non-local feature aggregation neural network model in step S2 further includes: A4: Input the candidate region into the R-CNN Head network to obtain the target region class score and the bounding box regression position, and at the same time input it into the MaskHead network to obtain the predicted mask of the target, separate the target from the complex environment and predict the contour key points of the target; A5: Input the RoI feature obtained in step A3 and the predicted mask obtained in step A4 into the MaskIoU Head network to regress the real mask and the predicted mask, calculate the intersection over union ratio of the predicted mask and the real mask, that is, the MaskIoU value, to obtain the target mask score, and complete the building of the non-local feature aggregation neural network model.

6. A target segmentation method based on a non-local feature aggregation neural network according to claim 5, characterized in that, The specific process of step A4 is: C1: Input the obtained feature map into the R-CNN Head for object classification and location regression; the R-CNN Head takes RoI features with an input shape of 7×7×256, followed by two fully connected layers with 1024 dimensions each, and then uses a fully connected layer with C dimensions to classify the objects into C categories, and adopts the softmax activation function to obtain the scores of each category, and uses a fully connected layer with 4 dimensions to obtain the regression of the object bounding box; C2: Input the obtained feature map into the Mask Head to obtain the predicted mask of the object; the MaskR-CNN Head takes RoI features with an input shape of 14×14×256, followed by four convolutional layers, and then followed by a transposed convolutional layer to double the size of the feature map, that is, 28×28×256, and then uses a convolutional layer to obtain a mask with a feature of 28×28×C, and uses the sigmoid activation function to obtain the predicted mask.

7. The object segmentation method based on the non-local feature aggregation neural network according to claim 1, characterized in that, in step A5, the MaskIoU Head consists of 3 convolutional layers, 1 max pooling layer and two fully connected layers, splices the predicted mask with the RoI feature map, inputs it into the MaskIoU Head, and obtains the object mask score, that is, the MaskIoU value of the object mask.

8. The object segmentation method based on the non-local feature aggregation neural network according to claim 1, characterized in that, step S3 specifically includes the following steps: D1: Collect the target video again, extract one frame of image per second without manual annotation to obtain unlabeled images; D2: Input all the unlabeled images into the object segmentation network to obtain the object category, classification score, mask contour key points and mask score in each image, multiply the classification score by the mask score to obtain the segmentation quality evaluation score of each object in the image, and the segmentation quality evaluation score is represented by the following formula: S mask = S cls × S IoU Among them, S mask is the segmentation quality assessment score, S cls is the classification score, and S IoU is the mask score.

9. The object segmentation method based on the non-local feature aggregation neural network according to claim 1, characterized in that, step S5 specifically includes the following steps: E1: For the edge contour closed curve of the output image of the target segmentation network, first initialize the parameters of the Douglas-Peucker algorithm and set the distance threshold D threshold ; E2: Calculate the two points M and N that are farthest apart on the closed curve, and connect the two points M and N to form a line segment D MN , and set the distance D q 's optimization function f(·); E3: BAS parameter initialization, including initializing the step decay factor E ta , step size S tep , required ratio v, number of iterations n t and the number of parameters m to be optimized t ; E4: Calculate the position x of the left antenna of the longhorn beetle l and its optimization function f(x l ), and the position x of the right antenna r and its optimization function f(x r ). The calculation formula is as follows: where d ir is the direction of the random longhorn beetle, and d 0 is the distance between the two antennae on the left and right, and sign(·) is the sign function If f(x l ) > f(x r ), then the longhorn beetle iterates to the next position x = x l , d = f(x l ), otherwise x = x r , d = f(x r ); E5: Repeat step E4 for a total of n t times to obtain the optimal solution as the optimized function value corresponding to the final position x of the longhorn beetle; E6: Find the point q among the remaining points on the closed curve that has the maximum distance from the line segment D MN and calculate the distance D MN from point q to the line segment D q ; Compare this distance D q with the pre-given threshold D threshold . If D q is less than D threshold , then use the line segment D MN as the approximation of the curve, and this section of the curve is processed. If the distance D q is greater than D threshold , then divide the curve into two segments M q and N q using point q, and perform the processing from step E3 to step E4 on each of the two segments of the curve respectively; After all the curves are processed, sequentially connect each contour key point to form a polygon as an approximation to the original closed curve.

Citation Information

Patent Citations

  • Deep learning image segmentation method and device based on sparse features

    CN110544256A

  • Retinal vessel segmentation method in fundus image and computer readable storage medium

    CN112233135A