Object extraction method and apparatus
Through technologies such as semantic segmentation and point-of-interest generation network, the object's point-of-interest and target mask diagrams are generated, which solves the problem of poor accuracy in rod object extraction in the existing technology, and achieves more efficient and accurate object extraction.
Patent Information
- Application Number
- CN202011238741.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-11-09
AI Technical Summary
The prior art has poor accuracy when extracting rod-shaped objects in images, mainly because the rod-shaped objects are slender and the detection effect of the surrounding frame is poor.
By inputting the image into the semantic segmentation feature network, a semantic segmentation result graph is generated, and a point of interest generation network, point of interest control network and instance selection network are used to generate the object's interest point and target mask graph, thereby determining the location of the object.
It effectively avoids the shortcomings of enclosing frame detection and improves the accuracy and effectiveness of object extraction, especially when dealing with elongated shaped objects.
Smart Images

Figure CN114494712B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to computer technology, and in particular, to an object extraction method and device. Background Art
[0002] During the process of map building, it is usually necessary to extract rod-shaped objects in image information, such as street lamp poles, utility poles, etc.
[0003] Currently, when extracting rod-shaped objects in image information, related technologies can first detect the boundaries of the bounding boxes of individual targets in the image, and then use a segmentation method to complete the semantic segmentation of the foreground and background. Finally, combining the boundaries of the bounding boxes of individual targets and the results of the foreground-background semantic segmentation to extract each rod-shaped object in the image.
[0004] However, since the shape of the rod-shaped object is slender and the aspect ratio and size of its bounding box are unbalanced, the accuracy of object extraction using the above-mentioned bounding box detection method is poor. Summary of the Invention
[0005] Embodiments of the present application provide an object extraction method and device to overcome the problem of poor accuracy in object extraction.
[0006] In a first aspect, an object extraction method provided by an embodiment of the present application includes:
[0007] Input a first picture including at least one object into a semantic segmentation feature network, so that the semantic segmentation feature network outputs a first feature map, where the first feature map is used to indicate the overall features of the first picture;
[0008] Process the first feature map according to a semantic segmentation classification network to obtain a semantic segmentation result map, where the semantic segmentation result map includes a layer of an object mask map and a layer of a background mask map;
[0009] Process the first feature map and the layer of the object mask map according to an interest point generation network to obtain at least one interest point, where one interest point corresponds to one object in the first picture;
[0010] Process the first feature map and the at least one interest point according to an interest point control network and an instance selection network to obtain a target mask map corresponding to each of the at least one interest point, where the target mask map is used to indicate the position of the object corresponding to the interest point in the first picture;
[0011] Determine the positions of each object in the first picture according to the target mask maps corresponding to each of the at least one interest point, so as to extract the at least one object.
[0012] In a possible design, the network for generating a network based on points of interest processes the layers of the first feature map and the object mask map to obtain at least one point of interest, including:
[0013] Input the first feature map into the network for generating points of interest, so that the network for generating points of interest outputs a second feature map, where the second feature map is used to indicate the probability of each pixel point in the first picture being a point of interest;
[0014] According to the layer of the object mask map and a first preset threshold, set the values of the pixel points that are not objects in the second feature map and the pixel points whose probability of being a point of interest is less than the first preset threshold to 0, to obtain a processed second feature map;
[0015] Traverse each pixel point in the processed second feature map. If the value of the currently traversed pixel point is greater than the values of the remaining pixel points within the range of M×N around it, then determine the currently traversed pixel point as a point of interest, where M is an integer greater than or equal to 1, and N is an integer greater than or equal to 1.
[0016] In a possible design, the network for controlling points of interest and the instance selection network process the first feature map and the at least one point of interest to obtain a target mask map corresponding to each of the at least one point of interest, including:
[0017] Process the first feature map and the at least one point of interest according to the network for controlling points of interest to obtain a variance feature and a mean feature corresponding to each of the points of interest;
[0018] Process the first feature map, the variance feature, and the mean feature according to the instance selection network to obtain a target mask map corresponding to each of the at least one point of interest.
[0019] In a possible design, the network for controlling points of interest processes the first feature map and the at least one point of interest to obtain a variance feature and a mean feature corresponding to each of the points of interest, including:
[0020] Input the first feature map into the network for controlling points of interest, so that the network for controlling points of interest outputs a third feature map, where the third feature map is used to indicate the data distribution feature of the points of interest;
[0021] For any one first point of interest among the N points of interest, process the third feature map and the first point of interest according to a first algorithm to obtain a fourth feature map, where the fourth feature map includes the feature corresponding to the coordinates of the first point of interest in the third feature map;
[0022] Input the fourth feature map into a first fully-connected convolutional layer to obtain a first control value of the first interest point, and input the fourth feature map into a second fully-connected convolutional layer to obtain a second control value of the second interest point, where the first control value is used to indicate the variance feature of the interest point, and the second control value is used to indicate the mean feature of the interest point.
[0023] In a possible design, the step of processing the first feature map, the variance feature, and the mean feature according to the instance selection network to obtain a target mask map corresponding to each of the at least one interest point includes:
[0024] Input the first feature map into the instance selection network so that the instance selection network outputs a fifth feature map, where the fifth feature map is used to indicate the semantic information feature of the at least one object in the first image;
[0025] Perform a normalization operation on the fifth feature map, and for any first interest point among the N interest points, perform a multiplication operation on the normalized fifth feature map and the first control value of the first interest point, and an addition operation with the second control value of the first interest point to obtain a sixth feature map, where the data distribution of the sixth feature map is the same as the data distribution of the first interest point;
[0026] Input the sixth feature map into at least one convolutional layer module to obtain a first target mask map corresponding to the first interest point.
[0027] In a possible design, the step of respectively determining the positions of the respective objects in the first picture according to the target mask maps corresponding to the at least one interest point to extract the at least one object includes:
[0028] Determine the overlapping area between the respective target mask maps according to the target mask maps corresponding to the at least one interest point;
[0029] Delete a second target mask map to obtain at least one remaining target mask map, where the second target mask map is one of multiple target mask maps with an overlapping area greater than a second preset threshold, and one target mask map corresponds to one object;
[0030] According to the at least one remaining target mask map, respectively determine the positions of the respective objects in the first picture to extract the at least one object.
[0031] In a possible design, before inputting the first picture including at least one object into the semantic segmentation feature network, the method further includes:
[0032] Perform scaling processing and normalization processing on the first picture.
[0033] In a possible design, the method further includes:
[0034] Construct an algorithm network, where the algorithm network includes at least one of the following networks: a semantic segmentation feature network, a semantic segmentation classification network, an interest point generation network, an interest point control network, and an instance selection network;
[0035] Train the algorithm network according to multiple groups of sample data to obtain a trained algorithm network. Each group of sample data includes a sample picture, a semantic label, an interest point label, and a mask label. The sample picture includes at least one sample object. The semantic label is used to indicate the position of the at least one sample object in the sample picture. The interest point label is used to indicate the position of the sample interest point corresponding to each sample object in the sample picture. The mask label is used to indicate the position of the sample target mask map corresponding to each sample interest point in the sample picture.
[0036] In a possible design, the training of the algorithm network according to multiple groups of sample data includes:
[0037] Input the sample picture into the algorithm network so that the algorithm network outputs a semantic result map, an interest point result map, and a mask result map;
[0038] Calculate the cross-entropy loss function according to the semantic result map, the interest point result map, the mask result map and the semantic label, the interest point label, the mask label respectively, and train the algorithm network using the backpropagation algorithm.
[0039] In a possible design, the object is a rod-shaped object.
[0040] In a second aspect, an object extraction device provided by an embodiment of the present application includes:
[0041] An input module, configured to input a first picture including at least one object into a semantic segmentation feature network so that the semantic segmentation feature network outputs a first feature map, where the first feature map is used to indicate the overall feature of the first picture;
[0042] A processing module, configured to process the first feature map according to a semantic segmentation classification network to obtain a semantic segmentation result map, where the semantic segmentation result map includes a layer of an object mask map and a layer of a background mask map;
[0043] The processing module is further configured to process the layers of the first feature map and the object mask map according to the interest points generated by the network to obtain at least one interest point, where one interest point corresponds to an object in the first picture;
[0044] The processing module is further configured to control the network and the instance selection network to process the first feature map and the at least one interest point according to the interest points to obtain a target mask map corresponding to each of the at least one interest points, where the target mask map is used to indicate the position of the object corresponding to the interest point in the first picture;
[0045] The processing module is further configured to respectively determine the positions of the respective objects in the first picture according to the target mask maps corresponding to the at least one interest points, so as to extract the at least one object.
[0046] In a possible design, the processing module is specifically configured to:
[0047] Input the first feature map into the interest point generation network, so that the interest point generation network outputs a second feature map, where the second feature map is used to indicate the probability of each pixel point in the first picture being an interest point;
[0048] According to the layer of the object mask map and a first preset threshold, set the values of the pixel points that are not objects in the second feature map and the pixel points whose probability of being an interest point is less than the first preset threshold to 0, to obtain a processed second feature map;
[0049] Traverse each pixel point in the processed second feature map. If the value of the currently traversed pixel point is greater than the values of the remaining pixel points within the range of M×N around it, then determine the currently traversed pixel point as an interest point, where M is an integer greater than or equal to 1, and N is an integer greater than or equal to 1.
[0050] In a possible design, the processing module is specifically configured to:
[0051] Process the first feature map and the at least one interest point according to the interest point control network to obtain a variance feature and a mean feature respectively corresponding to each of the interest points;
[0052] Process the first feature map, the variance feature, and the mean feature according to the instance selection network to obtain a target mask map corresponding to each of the at least one interest points.
[0053] In a possible design, the processing module is specifically configured to:
[0054] Input the first feature map into the interest point control network so that the interest point control network outputs a third feature map, where the third feature map is used to indicate the data distribution characteristics of the interest points;
[0055] For any first interest point among the N interest points, process the third feature map and the first interest point according to a first algorithm to obtain a fourth feature map, where the fourth feature map includes the features corresponding to the coordinates of the first interest point in the third feature map;
[0056] Input the fourth feature map into a first fully-connected convolutional layer to obtain a first control value of the first interest point, and input the fourth feature map into a second fully-connected convolutional layer to obtain a second control value of the second interest point, where the first control value is used to indicate the variance characteristics of the interest point, and the second control value is used to indicate the mean characteristics of the interest point.
[0057] In a possible design, the processing module is specifically configured to:
[0058] Input the first feature map into an instance selection network so that the instance selection network outputs a fifth feature map, where the fifth feature map is used to indicate the semantic information characteristics of the at least one object in the first image;
[0059] Perform a normalization operation on the fifth feature map, and for any first interest point among the N interest points, perform a multiplication operation on the normalized fifth feature map and the first control value of the first interest point, and an addition operation with the second control value of the first interest point to obtain a sixth feature map, where the data distribution of the sixth feature map is the same as the data distribution of the first interest point;
[0060] Input the sixth feature map into at least one convolutional layer module to obtain a first target mask map corresponding to the first interest point.
[0061] In a possible design, the processing module is specifically configured to:
[0062] Determine the overlapping area between the respective target mask maps according to the target mask maps corresponding to the at least one interest point;
[0063] Delete a second target mask map to obtain the remaining at least one target mask map, where the second target mask map is one of multiple target mask maps with an overlapping area greater than a second preset threshold, and one target mask map corresponds to one object;
[0064] Determine the positions of the respective objects in the first picture according to the remaining at least one target mask map to extract the at least one object.
[0065] In a possible design, the processing module is further configured to:
[0066] Before inputting the first picture including at least one object into the semantic segmentation feature network, perform scaling processing and normalization processing on the first picture.
[0067] In a possible design, the device further includes: a training module;
[0068] The training module is configured to construct an algorithm network, where the algorithm network includes at least one of the following networks: a semantic segmentation feature network, a semantic segmentation classification network, an interest point generation network, an interest point control network, an instance selection network;
[0069] Train the algorithm network according to multiple groups of sample data to obtain a trained algorithm network, where each group of sample data includes a sample picture, a semantic label, an interest point label, and a mask label. The sample picture includes at least one sample object. The semantic label is used to indicate the position of the at least one sample object in the sample picture. The interest point label is used to indicate the position of the sample interest point corresponding to each sample object in the sample picture. The mask label is used to indicate the position of the sample target mask map corresponding to each sample interest point in the sample picture.
[0070] In a possible design, the training module is specifically configured to:
[0071] Input the sample picture into the algorithm network so that the algorithm network outputs a semantic result map, an interest point result map, and a mask result map;
[0072] Calculate a cross-entropy loss function according to the semantic result map, the interest point result map, the mask result map and the semantic label, the interest point label, the mask label respectively, and train the algorithm network by using the backpropagation algorithm.
[0073] In a possible design, the object is a rod-shaped object.
[0074] In a third aspect, an object extraction device provided by an embodiment of the present application includes:
[0075] A memory for storing a program;
[0076] A processor for executing the program stored in the memory. When the program is executed, the processor is configured to execute the method according to any one of the above first aspect and various possible designs of the first aspect.
[0077] Fourthly, an embodiment of the present application provides a computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the method described in any one of the above first aspect and various possible designs of the first aspect.
[0078] An embodiment of the present application provides an object extraction method and apparatus. The method includes: inputting a first picture including at least one object into a semantic segmentation feature network, so that the semantic segmentation feature network outputs a first feature map, where the first feature map is used to indicate the overall features of the first picture. Processing the first feature map according to a semantic segmentation classification network to obtain a semantic segmentation result map, where the semantic segmentation result map includes a layer of an object mask map and a layer of a background mask map. Processing the first feature map and the layer of the object mask map according to an interest point generation network to obtain at least one interest point, where one interest point corresponds to one object in the first picture. Processing the first feature map and at least one interest point according to an interest point control network and an instance selection network to obtain a target mask map corresponding to each of the at least one interest point, where the target mask map is used to indicate the position of the object corresponding to the interest point in the first picture. Determining the positions of the respective objects in the first picture according to the target mask maps corresponding to the respective at least one interest point, so as to extract at least one object. In this embodiment, by generating at least one interest point, and controlling the selection and generation of the corresponding mask map according to the features of each interest point, the positions of the respective objects in the first picture can be obtained according to the respective target mask maps. Since the interest points are not affected by the shape ratio, size, etc. of the objects and are easily obtained through network training, it effectively avoids the poor extraction effect caused by detecting bounding boxes, and effectively ensures the effect and accuracy of extracting the rod-shaped objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for describing the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0080] Figure 1 It is a schematic diagram of the network structure of the convolutional neural network provided by the embodiment of the present application;
[0081] Figure 2 It is a flowchart of the object extraction method provided by the embodiment of the present application;
[0082] Figure 3 It is a flowchart of the object selection method provided by the embodiment of the present application;
[0083] Figure 4Schematic diagram of the first picture provided by the embodiment of the present application;
[0084] Figure 5 Schematic diagram of the semantic segmentation result map provided by the embodiment of the present application;
[0085] Figure 6 Schematic diagram of the implementation of determining the interest point provided by the embodiment of the present application;
[0086] Figure 7 Schematic diagram of the interest point provided by the embodiment of the present application;
[0087] Figure 8 Schematic diagram of the implementation of determining the data distribution characteristics of the interest point provided by the embodiment of the present application;
[0088] Figure 9 Schematic diagram of the interest point and the corresponding target mask map provided by the embodiment of the present application;
[0089] Figure 10 Flow chart of training the algorithm network provided by the embodiment of the present application;
[0090] Figure 11 Flow schematic diagram of the object extraction method provided by the embodiment of the present application;
[0091] Figure 12 Schematic diagram of the structure of the object extraction device provided by the embodiment of the present application;
[0092] Figure 13 Hardware structure schematic diagram of the object extraction device provided by the embodiment of the present application. Detailed implementation manners
[0093] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0094] First, in combination with Figure 1 introduce the convolutional neural network involved in the present application, Figure 1 Schematic diagram of the network structure of the convolutional neural network provided by the embodiment of the present application:
[0095] Convolutional Neural Network: A convolutional neural network (CNN) is a type of feedforward neural network that contains convolutional calculations and has a deep structure. It is one of the representative algorithms in deep learning. Convolutional neural networks have the ability of representational learning and can perform translation-invariant classification on input information according to their hierarchical structure. Therefore, they are also known as "translation-invariant artificial neural networks". Among them, the schematic diagram of the network structure of CNN is as shown in Figure 1 shown. CNN includes a convolutional layer and a pooling layer. The convolutional layer and the pooling layer will be introduced separately below:
[0096] Convolutional (conv) layer: The function of the convolutional layer is to extract features from the input static image. It contains multiple convolutional kernels inside. Each element that makes up the convolutional kernel corresponds to a weight coefficient and a bias term, similar to a neuron in a feedforward neural network. Each neuron in the convolutional layer is connected to multiple neurons in a region close to its position in the previous layer. The size of the region depends on the size of the convolutional kernel. When the convolutional kernel works, it will regularly scan the input features, perform matrix element multiplication and summation on the input features within the receptive field, and add the bias term.
[0097] Pooling layer: After feature extraction in the convolutional layer, the output feature map will be passed to the pooling layer for feature selection and information filtering. The pooling layer contains a pre-set pooling function, and its function is to replace the result of a single point in the feature map with the feature map statistic of its adjacent region. The pooling layer selects the pooling region in the same way as the convolutional kernel scans the feature map, which is controlled by the pooling size, stride, and padding.
[0098] Through the processing of the above convolutional layer and pooling layer, an image feature vector of the input image can be obtained. Among them, the image feature vector can include at least one channel.
[0099] During the processing of the convolutional neural network, by performing convolution on a certain range of feature maps, the patterns combined by multiple features can be extracted into one feature to obtain the next feature map. Then continue, perform convolution on the feature map, and the features continue to be combined to obtain a more complex feature map. Also because of the existence of the pooling layer, the most intense features within a certain range will be continuously extracted, and the size of the tensor will be reduced, so that the feature combinations in a large range can also be captured.
[0100] Generally speaking, a channel refers to the color channel of an image, and a feature map refers to the output result of a convolutional filter. But in fact, the essence of a channel and a feature map is the same, and both can be used to represent the data of a certain feature distribution on the previous input.
[0101] For example, when the input image is a color image with RGB channels, the input image can include 3 channels, namely the Red, Green, and Blue channels. One of the channels is for detecting a certain feature, and the magnitude of a value at a certain location in the channel reflects the strength of the current feature. For example, in a blue channel, if it is 256 levels, then if a pixel value is 255, it means a large blue degree. The meaning of the values in the feature map is similar to that of the channel, where the magnitude of the value reflects the strength of the current feature.
[0102] The background technology related to the present application will be described in detail below:
[0103] During the process of map building, it is usually necessary to extract rod-shaped objects in the image information, such as street lamp poles, electric poles, etc.
[0104] Currently, there are the following two possible implementation methods when extracting rod-shaped objects in the image information:
[0105] In a possible implementation method, the boundary of the bounding box of a single target in the image can be detected first, and then a segmentation method can be used to complete the semantic segmentation of the foreground and background. Finally, by combining the boundary of the bounding box of the single target and the result of the foreground-background semantic segmentation, each rod-shaped object in the image can be extracted.
[0106] However, in the above-described implementation method of detecting first and then segmenting, since the shape of the rod-shaped object is slender, the detection effect of the bounding box is poor. Specifically, currently, the detection of the bounding box usually uses an anchor-based detection method. Among them, the anchor-based detection method needs to pre-define the size and ratio of the anchor. For example, the ratio can be set to 1:1, 1:2, 1:3, and the size can be set to 0.1, 0.2, 0.3, a total of 9 shapes. If the set shapes are fewer, the detection effect will be poor. If the set shapes are more, the processing speed will be slow. It can be understood that because of the slender shape of the rod-shaped object, the ratio of the rod-shaped object has shapes such as 1:1 to 1:10, 1:20, etc., and the size ranges from 0.1 to 1.0 of the proportion of the image height. Therefore, using the anchor-based detection method will result in a poor detection effect of the bounding box, which poses a great challenge, and further leads to low accuracy of the extracted object.
[0107] In addition, a detection method based on the regression of the object center point to the boundary can also be used to detect the bounding box. However, this method first needs to ensure that the center point is inside the object, and the probability that the center point of the rod-shaped object is not inside the rod is relatively large. Therefore, using this method will still result in a poor detection effect of the bounding box, and further lead to a poor recognition effect of the rod-shaped object and low accuracy of the extracted rod-shaped object.
[0108] In another possible implementation, the image can be first subjected to semantic segmentation of the target to obtain the segmentation results of the foreground and the background, and then post-processing methods such as clustering can be used to separate the target into individual instance results, so as to extract each rod-shaped object in the image.
[0109] However, for the implementation method of first segmentation and then post-processing introduced above, a lot of manual experience parameter adjustments are required in the post-processing process, which will result in poor robustness of the processing. For example, in one post-processing method: after the semantic segmentation result, according to experience, the semantic segmentation result can be subjected to connected component analysis. Among them, for the disconnection points, according to the experience that the rod-shaped objects are generally vertical or horizontal, the areas within a certain range in the vertical or horizontal direction can be connected. For the connection points, according to the experience that the rod-shaped objects are generally vertical or horizontal, the result can be projected in the vertical or horizontal direction to obtain a histogram distribution, a certain threshold is set, and the connection points are separated. However, in the actual implementation process, because there are different shooting angles and different situations when shooting image data, there are always some rod-shaped objects that are not vertical or horizontal. The parameter settings that worked very well in one scenario will have a relatively poor effect in the next scenario, resulting in poor robustness. At this time, it is necessary to manually readjust the parameters according to the actual situation.
[0110] In view of the problems introduced above, the present application provides an object extraction method, which extracts the interest points corresponding to each object through an algorithm network and determines the mask graph corresponding to each interest point, where the mask graph corresponding to each interest point can indicate the position of each object in the image, so as to effectively extract the rod-shaped object from the image without performing post-processing operations such as bounding box detection and clustering. By pre-training the algorithm network and obtaining the interest points and the mask graph corresponding to the interest points according to the trained algorithm network, the recognition effect of the rod-shaped object can be effectively guaranteed.
[0111] The object extraction method provided by the present application will be introduced in detail below in combination with specific embodiments. It should be noted that the execution subject of each embodiment in the present application can be a component with data processing functions such as a server or a processor. The implementation of the execution subject in this embodiment is not limited, as long as it can perform data processing.
[0112] First, in combination with Figure 2 it is described as follows. Figure 2 FIG. is a flowchart of the object extraction method provided by an embodiment of the present application.
[0113] As Figure 2 shown, the method includes:
[0114] S201. Input a first picture including at least one object into a semantic segmentation feature network so that the semantic segmentation feature network outputs a first feature map, where the first feature map is used to indicate the overall features of the first picture.
[0115] The first picture can be used as the input picture of the algorithm network. Since the first picture includes at least one object, the object in the first picture can be extracted in this embodiment. The object can be, for example, the rod-shaped object introduced above, or in other possible implementation manners, the object can also be an object with a preset shape, and the preset shape can be, for example, a rectangle, a triangle, etc.
[0116] The semantic segmentation feature network and the semantic segmentation classification network in this embodiment are used to perform semantic segmentation on the first picture. Semantic segmentation can assign a semantic category to each pixel in the first picture to obtain a pixelated dense classification. In one possible implementation manner, the semantic segmentation feature network can include, for example, an encoder and a decoder. The encoder can be, for example, a Residual Network (ResNet) encoder, and the decoder can be, for example, a UNet decoder.
[0117] The semantic segmentation feature network processes the first picture and can output a first feature map corresponding to the first picture. The first feature map is used to indicate the overall features of the first picture. Based on the above introduction, it can be determined that the values included in the first feature map can be used to indicate the strength of the features.
[0118] S202. Process the first feature map according to the semantic segmentation classification network to obtain a semantic segmentation result map, including the layer of the object mask map and the layer of the background mask map in the semantic segmentation result map.
[0119] The semantic segmentation classification network in this embodiment is used to perform semantic segmentation processing on the first feature map to achieve semantic segmentation of the object and the background. After the semantic segmentation classification network processes the first feature map, a semantic segmentation result map can be obtained. The semantic segmentation result map includes two layers, namely the layer of the object mask map and the layer of the background mask map.
[0120] In the layer of the object mask map, the pixel points where the object is located are not 0, and the values of the remaining pixel points except the pixel points where the object is located are all 0, so as to achieve semantic segmentation of the object and the background.
[0121] In addition, in the layer of the background mask map, the pixel points where the background is located are not 0, and the values of the remaining pixel points except the pixel points where the background is located are all 0, so as to achieve semantic segmentation of the object and the background.
[0122] It can be understood that in the relevant image operations of the mask graph, the image operations have no effect on the pixel points with a value of 0 in the mask graph.
[0123] The object mask graph in this embodiment can effectively distinguish the objects and the background in the first picture, but cannot distinguish each object. For example, if there are two rod-shaped objects in the first picture, then the pixel values at the positions of these two rod-shaped objects in the object mask graph are not 0, but the two rod-shaped objects cannot be distinguished. Therefore, subsequent operations are still required to separately extract each object in the first picture.
[0124] S203. Process the first feature map and the object mask graph according to the interest point generation network to obtain at least one interest point, where one interest point corresponds to one object in the first picture.
[0125] The interest point generation network in this embodiment is used to generate the interest points of the object. Among them, the interest point can be any point in the object. In a possible implementation manner, one interest point can be generated for one object, or multiple interest points can also be generated for one object, where one interest point corresponds to one object. That is to say, the relationship between the object and the interest points generated by the interest point generation network in this embodiment is one-to-many.
[0126] In a possible implementation manner, the first feature map can be processed according to the interest point generation network to obtain the feature map related to the interest point, and then the feature map related to the interest point and the object mask graph are processed to screen out the positions of the objects from the feature map, so as to generate at least one interest point corresponding to each object respectively.
[0127] In another possible implementation manner, at least one interest point can also be randomly selected within the range of the object mask graph, as long as one interest point corresponds to one object and the interest point is a point within the range of the object mask graph.
[0128] It should be noted that the generation of the interest points is not affected by the shape ratio and size of the object. Therefore, in this embodiment, by generating at least one interest point and subsequently extracting the object based on the interest point, it can effectively avoid the poor effect of extracting the object caused by the unbalanced aspect ratio and size of the bounding box when performing bounding box detection.
[0129] S204. Process the first feature map and at least one interest point according to the interest point control network and the instance selection network to obtain the target mask graph corresponding to each at least one interest point, where the target mask graph is used to indicate the position of the object corresponding to the interest point in the first picture.
[0130] Among them, the point of interest control network is used to generate the control values of each point of interest. In a possible implementation, the point of interest control network can generate the variance feature and mean feature of each point of interest according to the first feature map.
[0131] In addition, the instance selection network in this embodiment is used to generate the mask map corresponding to each point of interest. In a possible implementation, taking any one of at least one point of interest as an example, the instance selection network can process according to the variance feature and mean feature of the point of interest, so as to obtain the target mask map corresponding to the point of interest.
[0132] The target mask map is similar to the object mask map introduced above. The difference is that the target mask map in this embodiment is used to indicate the position of a single object corresponding to the point of interest in the first picture. That is to say, compared with the object mask map, the non-zero pixel points in the target mask map are the pixel points of the position where the single object is located, so as to respectively indicate the positions of each object in the first picture.
[0133] S205. According to the target mask map corresponding to each of at least one point of interest, respectively determine the positions of each object in the first picture to extract at least one object.
[0134] Among them, each point of interest corresponds to its own target mask map, and each point of interest corresponds to its own object. Therefore, each target mask map also corresponds to an object.
[0135] In a possible implementation, an object may correspond to multiple target mask maps. Then, one of the multiple target mask maps of an object can be selected as the target mask map of the object. For example, one can be randomly selected from the multiple target mask maps, or the first obtained target mask map can be selected as the target mask map of the object according to the generation order of the target mask map.
[0136] After obtaining the target mask map corresponding to each object, the positions of each object in the first picture can be determined according to the non-zero pixel points in the target mask map, so as to extract at least one first object from the first object.
[0137] The object extraction method provided by the embodiment of the present application includes: inputting a first picture including at least one object into a semantic segmentation feature network, so that the semantic segmentation feature network outputs a first feature map, where the first feature map is used to indicate the overall feature of the first picture. Processing the first feature map according to a semantic segmentation classification network to obtain a semantic segmentation result map, where the semantic segmentation result map includes a layer of an object mask map and a layer of a background mask map. Processing the first feature map and the layer of the object mask map according to an interest point generation network to obtain at least one interest point, where one interest point corresponds to one object in the first picture. Processing the first feature map and at least one interest point according to an interest point control network and an instance selection network to obtain a target mask map corresponding to each of the at least one interest point, where the target mask map is used to indicate the position of the object corresponding to the interest point in the first picture. Determining the positions of each object in the first picture according to the target mask maps corresponding to each of the at least one interest point, so as to extract at least one object. In this embodiment, by generating at least one interest point, and controlling the selection to generate a corresponding mask map according to the features of each interest point, the positions of each object in the first picture can be obtained according to each target mask map. Since the interest points are not affected by the shape ratio, size, etc. of the object and are easily obtained through network training, the poor extraction effect caused by detecting bounding boxes is effectively avoided, and the effect and accuracy of extracting the rod-shaped object are effectively guaranteed.
[0138] Based on the above embodiments, the following combines Figures 3 to 8 to further introduce the object selection method provided by the present application in detail with an embodiment. Figure 3 is a flowchart of the object selection method provided by the embodiment of the present application. Figure 4 is a schematic diagram of the first picture provided by the embodiment of the present application. Figure 5 is a schematic diagram of the semantic segmentation result map provided by the embodiment of the present application. Figure 6 is a schematic diagram of the implementation of determining interest points provided by the embodiment of the present application. Figure 7 is a schematic diagram of the interest points provided by the embodiment of the present application. Figure 8 is a schematic diagram of the implementation of determining the data distribution characteristics of interest points provided by the embodiment of the present application. Figure 9 is a schematic diagram of the interest points and the corresponding target mask maps provided by the embodiment of the present application.
[0139] As Figure 3 shown, the method includes:
[0140] S301. Perform scaling processing and normalization processing on the first picture.
[0141] Among them, the implementation manner of the first picture can be, for example, as Figure 4 shown, where Figure 4The first picture in the example may include two rod-shaped objects, namely the rod-shaped objects indicated by 401 and 402 respectively.
[0142] In this embodiment, before performing relevant processing on the first picture, the first picture can be preprocessed first.
[0143] Among them, for the scaling process of the first picture, the first picture can be scaled to a specified size, and for the normalization process of the first picture, the values of each pixel point in the first picture can be mapped to the range of 0 to 1 for processing.
[0144] If the first picture is not normalized, due to the relatively scattered distribution of sample features, it may lead to slow learning speed or even difficulty in learning for the neural network. Therefore, in this embodiment, by performing scaling and normalization on the first picture, the standardization of the input picture in the subsequent processing can be effectively ensured, thereby improving the processing efficiency, and it can ensure that each feature is equally treated by the classifier, effectively ensuring the accuracy of the output result.
[0145] In a possible implementation, data standardization can be performed based on the mean and variance of the original data. For example, the standardization process can satisfy the following formula one:
[0146]
[0147] Among them, x is the original data before processing, μ is the mean of all data, σ is the variance of all data, and x * is the data after normalization processing.
[0148] In another possible implementation, for example, normalization can also be performed according to the maximum value and minimum value of the data. In this embodiment, the specific implementation of the normalization process is not limited, and it can be selected according to actual needs as long as the values of each pixel point can be mapped to the range of 0 to 1.
[0149] S302. Input the first picture including at least one object into the semantic segmentation feature network, so that the semantic segmentation feature network outputs a first feature map, where the first feature map is used to indicate the overall features of the first picture.
[0150] Among them, the implementation of S302 is similar to that of S201, and will not be elaborated here.
[0151] In this embodiment, the first picture input into the semantic segmentation feature network is the first picture after the above-mentioned scaling and normalization processing.
[0152] S303. Process the first feature map according to the semantic segmentation classification network to obtain a semantic segmentation result map, where the semantic segmentation result map includes a layer of the object mask map and a layer of the background mask map.
[0153] Among them, the implementation manner of S303 is similar to that of S202.
[0154] In a possible implementation manner, the semantic segmentation classification network in this embodiment may adopt a 1×1 convolution kernel, so as to output a semantic segmentation result map to obtain two layers, namely the background mask map and the object mask map.
[0155] Among them, the implementation manner of the semantic segmentation result map can be, for example, as Figure 5 shown, Figure 5 in which the white area is the layer of the object mask map, and the black area is the layer of the background mask map.
[0156] S304. Input the first feature map into the interest point generation network, so that the interest point generation network outputs a second feature map, where the second feature map is used to indicate the probability of each pixel point in the first picture being an interest point.
[0157] In this embodiment, the obtained first feature map can be input into the interest point generation network. In a possible implementation manner, the convolutional layer part of the interest point generation network can be composed of 3×3 and 1×1 convolution kernels. The interest point generation network processes the first feature map to output a second feature map, where the second feature map is used to indicate the probability of each pixel point in the first picture being an interest point.
[0158] It can be understood that the value in the second feature map can be a value in the range of 0 to 1. The larger the value, the greater the probability that the pixel point is an interest point. Correspondingly, the smaller the value, the smaller the probability that the pixel point is an interest point.
[0159] S305. According to the layer of the object mask map and the first preset threshold, set the values of the pixel points that are not objects in the second feature map and the pixel points whose probability of being an interest point is less than the first preset threshold to 0, to obtain a processed second feature map.
[0160] In this embodiment, at least one interest point needs to be selected in the first picture, where the interest point is a point in the object. Then, according to the layer of the object mask map, the values of the pixel points that are not objects in the second feature map can be set to 0, that is, the values of the pixel points in the background part are set to 0.
[0161] In addition, the second feature map in this embodiment can indicate the probability of each pixel point being an interest point. Therefore, in this embodiment, the pixel points can be filtered according to a first preset threshold, and the values of the pixel points with a probability of being an interest point less than the first preset threshold are set to 0, so as to filter out the pixel points with a probability less than the first preset threshold and obtain the processed second feature map.
[0162] Among them, the values of the pixel points that are not objects and the pixel points with a probability of being an interest point less than the first preset threshold in the processed second feature map are both 0. Then, the pixel points that are not 0 in the second feature map are the pixel points where the object is located and the probability is not less than the first preset threshold. After that, filtering out the interest points from the processed second feature map can effectively ensure the effectiveness and accuracy of the obtained interest points.
[0163] In a possible implementation manner, the first preset threshold can be set to 0.5, for example. In the actual implementation process, the first preset threshold can be selected according to the requirements for recall rate or accuracy and actual needs. For example, if a high recall rate is required, that is, it is required to prioritize identifying all rod-shaped objects and the requirement for error is not high, the first preset threshold can be set to 0.4; or if a high accuracy is required, that is, the tolerance for error is relatively low, the first preset threshold can be set to 0.6. The specific setting method of the first preset threshold can be selected according to actual needs.
[0164] S306. Traverse each pixel point in the processed second feature map. If the value of the currently traversed pixel point is greater than the values of the remaining pixel points within the range of M×N around it, then the currently traversed pixel point is determined as an interest point, where M is an integer greater than or equal to 1 and N is an integer greater than or equal to 1.
[0165] Among them, one interest point corresponds to one object in the first picture.
[0166] In this embodiment, in the processed second feature map, the pixel points corresponding to the positions where the object is located are not 0, and the probability of these pixel points being interest points is greater than the preset threshold. At this time, at least one interest point needs to be selected from the pixel points that are not 0.
[0167] In a possible implementation manner, it is possible to traverse and judge all the pixels of the processed second feature map. For each pixel point traversed, it is judged whether the value of the pixel point is greater than the values of the remaining pixel points within the range of M×N around the pixel point. If so, the currently traversed pixel point is determined as an interest point; otherwise, the currently traversed pixel point is determined not to be a pixel point.
[0168] Among them, the M×N range around the currently traversed pixel point can be the M×N range centered on the currently traversed pixel point. In a possible implementation, the values of M and N can be equal, and the values of M and N can be odd numbers.
[0169] Next, taking both M and N as 3, combined with Figure 6 a possible implementation of determining the interest point will be introduced.
[0170] Refer to Figure 6 , for example Figure 6 a part of the processed second feature map shown in. Assume that the currently traversed pixel point is the pixel point shown in 601, and the value of this pixel point is 0.7. For example, the coordinate position of the current pixel is (5, 5), and the 3×3 range around this pixel point is the range indicated by 602. Compare the value of pixel point 601 with the values of each pixel point within the range 602. Refer to Figure 6 It can be determined that its pixel point value is larger than the pixel point values of the image coordinate positions (4, 4), (4, 5), (4, 6), (5, 4), (5, 6), (6, 4), (6, 5), (6, 6) within the surrounding 3x3 range. That is to say, the value of pixel point 601 is the largest, then this pixel point 601 can be determined as the interest point. Traverse each pixel point in Figure 6 , so as to obtain Figure 6 the multiple interest points shown in, which are Figure 6 the pixel points indicated by each shaded part in.
[0171] Next, combined with Figure 7 an obtained interest point will be introduced. As Figure 7 shown, the interest point can be Figure 7 the point at 701 in. Among them, the interest point is the point in the rod-shaped object. During the implementation process, the number and position of the interest points can be selected according to actual needs, and this embodiment does not limit this.
[0172] S307: Input the first feature map into the interest point control network so that the interest point control network outputs a third feature map, where the third feature map is used to indicate the data distribution characteristics of the interest points.
[0173] After obtaining at least one interest point as described above, this embodiment needs to generate a mask map for each interest point. Then, the data distribution characteristics of each interest point can be learned first in order to generate the mask map corresponding to each interest point.
[0174] Among them, the first feature map can be input into the interest point control network. The interest point control network in this embodiment processes the first feature map and can output a third feature map. The third feature map can indicate the data distribution characteristics of the interest points. Among them, the interest point control network can include structures such as 3×3 and 1×1 convolutional layers, RoIAlign layer, and fully connected convolutional layer, for example.
[0175] S308. For any one of the N interest points, which is the first interest point, process the third feature map and the first interest point according to the first algorithm to obtain a fourth feature map. Among them, the fourth feature map includes the features corresponding to the coordinates of the first interest point in the third feature map.
[0176] In this embodiment, each interest point corresponds to its own coordinates, that is, the corresponding (x, y). Here, take any one of the N interest points, which is the first interest point, as an example for introduction. The implementation methods of the other interest points are similar.
[0177] The following can be combined with Figure 8 to understand the implementation method of obtaining the data distribution characteristics of the interest points. Refer to Figure 8 , the first feature map can be input into the interest point control network to obtain a third feature map. After that, for the first interest point, process the third feature map and the first interest point according to the first algorithm to obtain a fourth feature map.
[0178] Refer to Figure 8 It can be understood that the fourth feature map is a part of the third feature map. The fourth feature map includes the features corresponding to the coordinates of the first interest point in the third feature map. For example, if the current coordinates of the first interest point are (x, y), then according to the coordinates of the first interest point, the features corresponding to the coordinates (x, y) on the z-axis can be extracted to obtain the fourth feature map. Among them, each interest point corresponds to its own fourth feature map.
[0179] In a possible implementation manner, the first algorithm can be, for example, the RoIAlign algorithm. Among them, the RoIAlign algorithm is an algorithm for performing regional feature aggregation.
[0180] S309. Input the fourth feature map into the first fully connected convolutional layer to obtain the first control value of the first interest point, and input the fourth feature map into the second fully connected convolutional layer to obtain the second control value of the second interest point. Among them, the first control value is used to indicate the variance characteristics of the interest point, and the second control value is used to indicate the mean characteristics of the interest point.
[0181] After obtaining the fourth feature map corresponding to the first interest point, the fourth feature map can be input into the fully connected convolutional layer to obtain the control value corresponding to the first interest point.
[0182] In a possible implementation, the features corresponding to the positions in the third feature map can be mapped from the first point of interest, and the data distribution represented by the features at these positions can be learned, so as to learn the variance feature and the mean feature information respectively. For example, the fourth feature map can be input into the first fully-connected convolutional layer to obtain the first control value of the first point of interest, where the first control value is used to indicate the variance feature of the point of interest, and, the fourth feature map can be input into the second fully-connected convolutional layer to obtain the second control value of the second point of interest, and the second control value is used to indicate the mean feature of the point of interest.
[0183] After obtaining the variance feature and the mean feature of the first point of interest, that is, obtaining the data distribution feature of the first point of interest, the above operations are performed on each of at least one point of interest, so as to obtain the variance feature and the mean feature corresponding to each point of interest respectively.
[0184] S310. Input the first feature map into the instance selection network, so that the instance selection network outputs the fifth feature map, where the fifth feature map is used to indicate the semantic information feature of at least one object in the first image.
[0185] In this embodiment, the convolutional layer part of the instance selection network can be composed of 3×3 and 1×1 convolutional kernels, and the instance selection network can output the fifth feature map used to indicate the semantic information feature of at least one object in the first image.
[0186] S311. Perform a normalization operation on the fifth feature map, and perform a multiplication operation on the normalized fifth feature map and the first control value of the first point of interest, and an addition operation with the second control value of the first point of interest, so as to obtain the sixth feature map, where the data distribution of the sixth feature map is the same as the data distribution of the first point of interest.
[0187] In the above processing process, this embodiment obtains the mean feature and the mean variance corresponding to each point of interest respectively. In order to obtain the mask map corresponding to each point of interest, the data distribution in the fifth feature map used to indicate the semantic information feature can be changed to the data distribution feature corresponding to the point of interest.
[0188] In a possible implementation, a normalization operation can be performed on the fifth feature map, and the processing of the normalization operation can be the process of processing through the mean and variance introduced above. First, subtract the mean, and then divide the result of subtracting the mean by the variance. By performing a normalization operation on the fifth feature map, the data distribution in the fifth feature map can be eliminated first.
[0189] After that, perform a multiplication operation on the fifth feature map after the normalization operation and the variance feature of the first interest point, and perform an addition operation on the result of the multiplication operation and the mean feature of the first interest point. After that, a sixth feature map is obtained, thus realizing the change of the data distribution of the fifth feature map to the data distribution characteristics of the interest point. Among them, the data distribution of the sixth feature map is the same as the data distribution of the first interest point.
[0190] In this embodiment, the first interest point is taken as an example for introduction, and the sixth feature map corresponding to the first interest point is obtained. In the actual implementation process, for each interest point, its corresponding sixth feature map will be obtained.
[0191] S312. Input the sixth feature map into at least one convolutional layer module to obtain the first target mask map corresponding to the first interest point.
[0192] Taking the first interest point as an example, after obtaining the sixth feature map corresponding to the first interest point, the sixth feature map can be input into multiple convolutional layer modules. Because the data distribution of the sixth feature map is the same as the data distribution of the first interest point, and the sixth feature map is obtained from the fifth feature map used to indicate the semantic information features of at least one object in the first image, the mask result map corresponding to the first interest point can be generated.
[0193] Among them, each interest point corresponds to its own object. Then generating the mask result maps corresponding to each interest point actually means obtaining the mask result maps corresponding to each object.
[0194] In a possible implementation manner, the implementation manner of the mask result maps corresponding to each interest point can be, for example, as Figure 9 shown, see Figure 9 , assuming that 5 interest points are currently determined, one interest point corresponds to one picture, and the points shown as 901, 902, 903, 904, and 905 in the 5 pictures are the 5 interest points. Each interest point corresponds to its own mask result map, which is the black and white picture on the left side of the 5 pictures.
[0195] In the mask result map corresponding to the interest point, the white area is the target mask map, and the black area is the background mask map. See Figure 9 It can be determined that there is only one object in each target mask map, thus realizing the separation of each object.
[0196] S313. Determine the overlapping area between each target mask map according to the target mask maps corresponding to at least one interest point.
[0197] S314. Delete the second target mask map to obtain the remaining at least one target mask map, where the second target mask map is one of multiple target mask maps with a duplicate area greater than a second preset threshold, and one target mask map corresponds to one object.
[0198] Participate Figure 9 It can be determined that in this embodiment, multiple interest points may be generated for the same object, and thus multiple mask maps are generated for the same object. For example Figure 9 There are only 2 rod-shaped objects in the first picture in, but finally 5 mask maps are generated.
[0199] Therefore, in this embodiment, the redundant target mask maps are deleted, and only one target mask map among the multiple target mask maps with a duplicate area greater than the preset threshold is retained as the target mask map of one of the objects.
[0200] For example, the duplicate area between each target mask map can be determined. See Figure 9 , assuming that the duplicate area between target mask map 906 and target mask map 907 can be determined, and there is no duplicate area between target mask map 906 and target mask map 907, then no deletion is performed; then the duplicate area between target mask map 906 and target mask map 909 can be determined. After judgment, it can be determined that the duplicate area between target mask map 906 and target mask map 909 is greater than the second preset threshold, then target mask map 909 can be deleted, or target mask map 906 can be deleted. The above operations are performed for each target mask map, and finally the target mask maps corresponding to the two rod-shaped objects can be obtained.
[0201] In this embodiment, for multiple target mask maps with a duplicate area greater than the preset threshold, there is no limitation on which specific target mask map to delete. For example, one can be randomly deleted, or the one with an earlier generation time can be deleted.
[0202] In one possible implementation, the target mask maps corresponding to the two rod-shaped objects finally obtained can be two separate target mask maps; or, the target mask maps corresponding to the two rod-shaped objects finally obtained in this embodiment can also be combined target mask maps. In this target mask map, the target mask maps of the two rod-shaped objects are included respectively, and there is a corresponding relationship between the target mask map and each rod-shaped object to distinguish the two rod-shaped objects.
[0203] S315. According to the remaining at least one target mask map, determine the positions of each object in the first picture respectively to extract at least one object.
[0204] In this embodiment, after deleting the duplicate target mask images, at least one remaining target mask image is obtained. The non-zero part in the target mask image is the position of the object in the first picture, so that the positions of the respective objects in the first picture can be determined. After determining the positions of the respective objects, at least one object category two can be extracted.
[0205] The object extraction method provided by the embodiment of the present application includes: by performing scaling processing and normalization processing on the first picture, the standardization of the first picture input into the algorithm network can be effectively ensured, thereby effectively improving the processing efficiency, and by processing the second feature map through the layers of the object mask image and the first preset threshold, the pixel points of the background and the pixel points with relatively small probabilities are screened out, and the pixel point with the largest value is selected by region among the remaining pixel points, so as to obtain at least one interest point, which can effectively ensure the accuracy and effectiveness of the obtained interest point. Then, for each interest point, a respective corresponding mask image is generated, so as to obtain at least one mask image corresponding to each object respectively, and according to the overlapping area between the mask images, the overlapping mask images are deleted, thereby realizing the determination of the positions of the respective objects in the first image respectively, so as to realize the extraction of each rod-shaped object in the first image respectively. In the above process, a series of processing is performed using the algorithm network, which can effectively ensure the accuracy of the extracted object results.
[0206] On the basis of the above embodiment, in the present application, before processing the first picture based on the above-introduced algorithm networks, the algorithm networks can also be trained to ensure the accuracy of the results output by the algorithm networks. The implementation manner of training the algorithm networks in the present application will be introduced below in combination with specific embodiments. Figure 10 It is a flowchart of training the algorithm network provided by the embodiment of the present application.
[0207] As Figure 10 shown, the method includes:
[0208] S1001. Construct an algorithm network, where the algorithm network includes at least one of the following networks: semantic segmentation feature network, semantic segmentation classification network, interest point generation network, interest point control network, instance selection network.
[0209] S1002. Train the algorithm network according to multiple groups of sample data to obtain the trained algorithm network.
[0210] In this embodiment, the algorithm network can include each network introduced in the above embodiment, and in this embodiment, multiple groups of sample data can be used to train the algorithm network.
[0211] Among them, each group of sample data may include a sample picture, a semantic label, an interest point label, and a mask label. Among them, the sample picture includes at least one sample object. The semantic label is used to indicate the position of at least one sample object in the sample picture. The interest point label is used to indicate the position of the sample interest point corresponding to each sample object in the sample picture. The mask label is used to indicate the position of the sample target mask map corresponding to each sample interest point in the sample picture.
[0212] In a possible implementation manner, the semantic label may be a set of multiple coordinate points, and the position of the sample object in the sample picture is indicated by the coordinates; and, the interest point label may be a coordinate point randomly determined from the set of multiple coordinate points of the semantic label. It can be understood that the interest point is a point in the object, so a seat interest point can be randomly selected from the set of coordinate points corresponding to the object; and, the mask label may be a set of multiple coordinate points. Since the sample mask map is used to indicate the position of the sample object, the mask label in this embodiment may be a set of coordinate points of the sample object corresponding to the interest point label.
[0213] It can be understood that the sample data in this embodiment is pre-determined and can ensure the accuracy of the data therein. Therefore, training the algorithm network according to the sample data can improve the accuracy of the output result of the algorithm network after training.
[0214] In a possible implementation manner, the sample picture may be input into the algorithm network so that the algorithm network outputs a semantic result map, an interest point result map, and a mask result map.
[0215] Among them, the implementation manner in which the algorithm network outputs a semantic result map, an interest point result mask result map according to the sample picture is similar to the processing flow of each algorithm network introduced in the above embodiment, and will not be elaborated here.
[0216] Next, the cross-entropy loss function can be calculated according to the semantic result map, the interest point result map, the mask result map and the semantic label, the interest point label, the mask label respectively, and the algorithm network can be trained using the backpropagation algorithm.
[0217] The object extraction method provided by the embodiments of the present application can effectively ensure the accuracy of the output result of the algorithm network by training the algorithm network according to multiple groups of sample data before the algorithm network processes the first picture, and further ensure the accuracy of the object finally extracted from the first picture.
[0218] On the basis of the above embodiment, the following flow schematic diagram gives a systematic introduction to the implementation process of the object extraction method provided by the present application. Figure 11 This is a flow schematic diagram of the object extraction method provided by the embodiments of the present application.
[0219] As Figure 11 shown below:
[0220] First, input the first image into the semantic segmentation feature network for feature extraction operations. The semantic segmentation feature network can use a ResNet encoder and a UNet decoder to output the first feature map.
[0221] Among them, before inputting the first image into the semantic segmentation feature network, the first image can also be subjected to scaling processing and normalization processing.
[0222] After that, input the first feature map into the semantic segmentation classification network. The semantic segmentation classification network can use a 1×1 convolutional kernel to output the semantic segmentation result map. The semantic segmentation result map includes two layers, namely the layer of the object mask map and the layer of the background mask map, initially realizing the semantic segmentation of the object and the background.
[0223] Next, the first feature map can be input into the interest point generation network to generate the second feature map. Among them, the convolutional layer part of the interest point generation network can be composed of 3×3 and 1×1 convolutions. The second feature map in this embodiment can be a candidate point feature map. Subsequently, according to the layer of the object mask map and the first preset threshold, the values of the non-object regions and the regions smaller than the first preset threshold in the second feature map can be set to 0 to obtain the processed second feature map. Then, perform regional maximum point selection on the processed second feature map to select the points with the largest values in each region to obtain N interest points.
[0224] In addition, the first feature map can be input into the interest point control network to generate the third feature map. The third feature map in this embodiment can be a control feature map. The interest point control network can be composed of structures such as 3×3 and 1×1 convolutional layers, RoIAlign layers, and fully connected convolutional layers.
[0225] Subsequently, traverse each interest point. For each interest point, input the third feature map and the coordinates of the interest point into the RoIAlign algorithm to extract the corresponding fourth feature map. Then, input the fourth feature map into the fully connected convolutional layer to calculate the first control value and the second control value. The meaning of the first control value can be the variance feature of the interest point, and the meaning of the second control value can be the mean feature of the interest point.
[0226] Further, the first feature map can be input into an instance selection network to generate a fifth feature map. The convolutional layer part of the instance selection network can be composed of 3×3 and 1×1 convolutions, etc. Subsequently, after normalizing the fifth feature map, a multiplication operation is performed with the first control value and an addition operation is performed with the second control value, that is, the data distribution of the fifth feature map is changed to the data distribution learned by the points of interest, obtaining a sixth feature map. Finally, the sixth feature map is input into multiple convolutional layer modules to generate the target mask map corresponding to the points of interest.
[0227] For each point of interest, the operations of obtaining the control value and generating the target mask map are performed as described above, so as to generate the target mask map corresponding to each point of interest. Then, the overlapping area of each target mask map is determined, and the target mask map and the corresponding point of interest with an overlapping area exceeding the second preset threshold are discarded, obtaining a binary classification mask result. Finally, the mask map results corresponding to the remaining points of interest are combined to obtain the final rod-shaped object segmentation result, thereby obtaining the final object extraction result.
[0228] It can be understood that for long and narrow objects such as rod-shaped objects, due to the imbalance in the aspect ratio and size of their bounding boxes, the effect of using the detection method to complete the instance segmentation is poor. The points of interest within the rod-shaped object area are not affected by the shape ratio and size of the object, etc., and are easily obtained through network training. Therefore, the object selection method provided in the embodiments of the present application generates at least one point of interest, and controls the selection and generation of the corresponding mask map according to the characteristics of each point of interest, and combines the mask maps generated for different points of interest to complete the segmentation and extraction of each rod-shaped object, thereby effectively ensuring the effect and accuracy of the extracted rod-shaped object.
[0229] Figure 12 It is a schematic structural diagram of an object extraction device provided in an embodiment of the present application. As Figure 12 shown, the device 120 includes: an input module 1201 and a processing module 1202.
[0230] The input module 1201 is configured to input a first picture including at least one object into a semantic segmentation feature network, so that the semantic segmentation feature network outputs a first feature map, where the first feature map is used to indicate the overall features of the first picture;
[0231] The processing module 1202 is configured to process the first feature map according to a semantic segmentation classification network to obtain a semantic segmentation result map, where the semantic segmentation result map includes a layer of an object mask map and a layer of a background mask map;
[0232] The processing module 1202 is further configured to process the first feature map and the layer of the object mask map according to a point of interest generation network to obtain at least one point of interest, where one point of interest corresponds to one object in the first picture;
[0233] The processing module 1202 is further configured to process the first feature map and the at least one point of interest according to the point-of-interest control network and the instance selection network, to obtain a target mask map corresponding to each of the at least one point of interest, where the target mask map is used to indicate the position of the object corresponding to the point of interest in the first picture;
[0234] The processing module 1202 is further configured to respectively determine the positions of the respective objects in the first picture according to the target mask maps corresponding to the at least one point of interest, so as to extract the at least one object.
[0235] In a possible design, the processing module 1202 is specifically configured to:
[0236] Input the first feature map into a point-of-interest generation network, so that the point-of-interest generation network outputs a second feature map, where the second feature map is used to indicate the probability of each pixel point in the first picture being a point of interest;
[0237] According to the layer of the object mask map and a first preset threshold, set the values of the pixel points that are not objects and the pixel points whose probability of being a point of interest is less than the first preset threshold in the second feature map to 0, to obtain a processed second feature map;
[0238] Traverse each pixel point in the processed second feature map. If the value of the currently traversed pixel point is greater than the values of the remaining pixel points within the range of M×N around it, then determine the currently traversed pixel point as a point of interest, where M is an integer greater than or equal to 1, and N is an integer greater than or equal to 1.
[0239] In a possible design, the processing module 1202 is specifically configured to:
[0240] Process the first feature map and the at least one point of interest according to the point-of-interest control network, to obtain a variance feature and a mean feature respectively corresponding to each of the points of interest;
[0241] Process the first feature map, the variance feature, and the mean feature according to the instance selection network, to obtain a target mask map corresponding to each of the at least one point of interest.
[0242] In a possible design, the processing module 1202 is specifically configured to:
[0243] Input the first feature map into a point-of-interest control network, so that the point-of-interest control network outputs a third feature map, where the third feature map is used to indicate the data distribution feature of the points of interest;
[0244] For any first interest point among the N interest points, the third feature map and the first interest point are processed according to a first algorithm to obtain a fourth feature map, where the fourth feature map includes the features corresponding to the coordinates of the first interest point in the third feature map;
[0245] The fourth feature map is input into a first fully connected convolutional layer to obtain a first control value of the first interest point, and the fourth feature map is input into a second fully connected convolutional layer to obtain a second control value of the second interest point, where the first control value is used to indicate the variance feature of the interest point, and the second control value is used to indicate the mean feature of the interest point.
[0246] In a possible design, the processing module 1202 is specifically configured to:
[0247] Input the first feature map into an instance selection network so that the instance selection network outputs a fifth feature map, where the fifth feature map is used to indicate the semantic information features of the at least one object in the first image;
[0248] Perform a normalization operation on the fifth feature map, and for any first interest point among the N interest points, perform a multiplication operation on the normalized fifth feature map and the first control value of the first interest point, and an addition operation with the second control value of the first interest point to obtain a sixth feature map, where the data distribution of the sixth feature map is the same as the data distribution of the first interest point;
[0249] Input the sixth feature map into at least one convolutional layer module to obtain a first target mask map corresponding to the first interest point.
[0250] In a possible design, the processing module 1202 is specifically configured to:
[0251] Determine the overlapping area between the respective target mask maps according to the target mask maps corresponding to the at least one interest point;
[0252] Delete a second target mask map to obtain the remaining at least one target mask map, where the second target mask map is one of the multiple target mask maps with an overlapping area greater than a second preset threshold, and one target mask map corresponds to one object;
[0253] Determine the positions of the respective objects in the first picture according to the remaining at least one target mask map to extract the at least one object.
[0254] In a possible design, the processing module 1202 is further configured to:
[0255] Before inputting the first picture including at least one object into the semantic segmentation feature network, perform scaling processing and normalization processing on the first picture.
[0256] In a possible design, the device further includes: a training module 1203;
[0257] The training module 1203 is used to construct an algorithm network, where the algorithm network includes at least one of the following networks: a semantic segmentation feature network, a semantic segmentation classification network, an interest point generation network, an interest point control network, and an instance selection network;
[0258] Train the algorithm network according to multiple groups of sample data to obtain a trained algorithm network, where each group of sample data includes a sample picture, a semantic label, an interest point label, and a mask label. Among them, the sample picture includes at least one sample object, the semantic label is used to indicate the position of the at least one sample object in the sample picture, the interest point label is used to indicate the position of the sample interest point corresponding to each sample object in the sample picture, and the mask label is used to indicate the position of the sample target mask map corresponding to each sample interest point in the sample picture.
[0259] In a possible design, the training module 1203 is specifically used for:
[0260] Input the sample picture into the algorithm network so that the algorithm network outputs a semantic result map, an interest point result map, and a mask result map;
[0261] Calculate the cross-entropy loss function according to the semantic result map, the interest point result map, the mask result map and the semantic label, the interest point label, the mask label respectively, and train the algorithm network by using the backpropagation algorithm.
[0262] In a possible design, the object is a rod-shaped object.
[0263] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar. Details are not described herein again.
[0264] Figure 13 This is a schematic hardware structure diagram of the object extraction device provided in the embodiment of the present application. As Figure 13 shown, the object extraction device 130 in this embodiment includes: a processor 1301 and a memory 1302; where
[0265] The memory 1302 is used to store computer execution instructions;
[0266] A processor 1301 for executing computer-executable instructions stored in a memory to implement the respective steps performed by the object extraction method in the above embodiments. For details, reference may be made to the relevant descriptions in the foregoing method embodiments.
[0267] Optionally, the memory 1302 can be either independent or integrated with the processor 1301.
[0268] When the memory 1302 is independently provided, the object extraction device further includes a bus 1303 for connecting the memory 1302 and the processor 1301.
[0269] An embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the object extraction method performed by the object extraction device as described above.
[0270] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in an electrical, mechanical or other form.
[0271] The integrated modules implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in the various embodiments of the present application.
[0272] It should be understood that the above processor can be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), and can also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0273] The memory may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc.
[0274] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the accompanying drawings of this application are not limited to only one bus or one type of bus.
[0275] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0276] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disks, or optical discs and other media that can store program codes.
[0277] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An object extraction method, characterized in that, Including: Inputting a first picture including at least one object into a semantic segmentation feature network, so that the semantic segmentation feature network outputs a first feature map, where the first feature map is used to indicate the overall features of the first picture; Processing the first feature map according to a semantic segmentation classification network to obtain a semantic segmentation result map, where the semantic segmentation result map includes a layer of an object mask map and a layer of a background mask map; Processing the first feature map and the layer of the object mask map according to an interest point generation network to obtain at least one interest point, where one interest point corresponds to one object in the first picture; Processing the first feature map and the at least one interest point according to an interest point control network to obtain a variance feature and a mean feature respectively corresponding to each of the interest points; processing the first feature map, the variance feature, and the mean feature according to an instance selection network to obtain a target mask map respectively corresponding to each of the at least one interest points, where the target mask map is used to indicate the position of the object corresponding to the interest point in the first picture; Determining the positions of the respective objects in the first picture according to the target mask maps respectively corresponding to the at least one interest points, so as to extract the at least one object.
2. The method according to claim 1, wherein The processing the first feature map and the layer of the object mask map according to the interest point generation network to obtain at least one interest point includes: Inputting the first feature map into the interest point generation network, so that the interest point generation network outputs a second feature map, where the second feature map is used to indicate the probability of each pixel point in the first picture being an interest point; According to the layer of the object mask map and a first preset threshold, setting the values of the pixel points that are not objects and the pixel points with a probability of being an interest point less than the first preset threshold in the second feature map to 0, to obtain a processed second feature map; Traversing each pixel point in the processed second feature map, if the value of the currently traversed pixel point is greater than the values of the remaining pixel points within the range of M×N around it, then determining the currently traversed pixel point as an interest point, where M is an integer greater than or equal to 1, and N is an integer greater than or equal to 1.
3. The method according to claim 1, characterized in that The processing the first feature map and the at least one interest point according to the interest point control network to obtain a variance feature and a mean feature respectively corresponding to each of the interest points includes: Inputting the first feature map into the interest point control network, so that the interest point control network outputs a third feature map, where the third feature map is used to indicate the data distribution feature of the interest points; For any first interest point among the N interest points, processing the third feature map and the first interest point according to a first algorithm to obtain a fourth feature map, where the fourth feature map includes the features corresponding to the coordinates of the first interest point in the third feature map; Input the fourth feature map into the first fully connected convolutional layer to obtain the first control value of the first interest point, and input the fourth feature map into the second fully connected convolutional layer to obtain the second control value of the first interest point, where the first control value is used to indicate the variance feature of the interest point, and the second control value is used to indicate the mean feature of the interest point.
4. The method according to claim 3, wherein The processing of the first feature map, the variance feature, and the mean feature by the instance selection network to obtain the target mask map corresponding to each of the at least one interest point includes: Input the first feature map into the instance selection network so that the instance selection network outputs a fifth feature map, where the fifth feature map is used to indicate the semantic information feature of the at least one object in the first picture; Perform a normalization operation on the fifth feature map, and for any first interest point among the N interest points, perform a multiplication operation on the normalized fifth feature map and the first control value of the first interest point, and an addition operation with the second control value of the first interest point to obtain a sixth feature map, where the data distribution of the sixth feature map is the same as the data distribution of the first interest point; Input the sixth feature map into at least one convolutional layer module to obtain the first target mask map corresponding to the first interest point.
5. The method according to any one of claims 1 to 4, characterized in that, The determining the positions of the respective objects in the first picture according to the target mask maps corresponding to the at least one interest point to extract the at least one object includes: Determine the overlapping area between the respective target mask maps according to the target mask maps corresponding to the at least one interest point; Delete the second target mask map to obtain the remaining at least one target mask map, where the second target mask map is one of multiple target mask maps with an overlapping area greater than a second preset threshold, and one target mask map corresponds to one object; Determine the positions of the respective objects in the first picture according to the remaining at least one target mask map to extract the at least one object.
6. The method according to claim 1, characterized in that, Before inputting the first picture including at least one object into the semantic segmentation feature network, the method further includes: Perform a scaling process and a normalization process on the first picture.
7. The method according to claim 1, characterized in that, The method further includes: Construct an algorithm network, where the algorithm network includes at least one of the following networks: a semantic segmentation feature network, a semantic segmentation classification network, an interest point generation network, an interest point control network, an instance selection network; Train the algorithm network according to multiple groups of sample data to obtain a trained algorithm network, where each group of sample data includes a sample picture, a semantic label, an interest point label, and a mask label, where the sample picture includes at least one sample object, the semantic label is used to indicate the position of the at least one sample object in the sample picture, the interest point label is used to indicate the position of the sample interest point corresponding to each sample object in the sample picture, and the mask label is used to indicate the position of the sample target mask map corresponding to each sample interest point in the sample picture.
8. The method according to claim 7, characterized in that Training the algorithm network based on multiple sets of sample data includes: Inputting the sample pictures into the algorithm network so that the algorithm network outputs a semantic result map, an interest point result map, and a mask result map; Calculating the cross-entropy loss function according to the semantic result map, the interest point result map, the mask result map and the semantic label, the interest point label, the mask label respectively, and training the algorithm network by using the backpropagation algorithm.
9. The method according to any one of claims 1-4, 6, characterized in that, The object is a rod-shaped object.
10. An object extraction device, characterized in that, It includes: An input module for inputting a first picture including at least one object into a semantic segmentation feature network so that the semantic segmentation feature network outputs a first feature map, where the first feature map is used to indicate the overall feature of the first picture; A processing module for processing the first feature map according to a semantic segmentation classification network to obtain a semantic segmentation result map, where the semantic segmentation result map includes a layer of an object mask map and a layer of a background mask map; The processing module is further configured to process the first feature map and the layer of the object mask map according to an interest point generation network to obtain at least one interest point, where one interest point corresponds to one object in the first picture; The processing module is further configured to process the first feature map and the at least one interest point according to an interest point control network to obtain the variance feature and the mean feature respectively corresponding to each of the interest points; processing the first feature map, the variance feature, and the mean feature according to an instance selection network to obtain a target mask map respectively corresponding to each of the at least one interest points, where the target mask map is used to indicate the position of the object corresponding to the interest point in the first picture; The processing module is further configured to determine the positions of the respective objects in the first picture according to the target mask maps respectively corresponding to the at least one interest points, so as to extract the at least one object.
11. An object extraction device, characterized in that, It includes: A memory for storing a program; A processor for executing the program stored in the memory, and when the program is executed, the processor is configured to execute the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, It includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Monitoring video target detection method based on background elimination
CN109993091A
Instance segmentation method and device based on feature attention and sub-upsampling
CN110532955A