A method and system for video instance segmentation of street scenes
By constructing a multi-receptive field downsampling module and an anchor frame calibration module, and designing a street scene video instance segmentation model, the problems of insufficient edge feature extraction and lack of high-level feature details in the prior art are solved, and higher detection and segmentation accuracy and speed are achieved.
Patent Information
- Application Number
- CN202210073712.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-01-21
AI Technical Summary
In the prior art, the use of single receptive field sampling of multiple types of aspect ratio anchor boxes leads to insufficient edge feature extraction and lack of detailed information on the spatial location of the high-level feature of the feature pyramid.
A multi-receptive field sampling module is constructed, a spatial position information compensation feature pyramid and anchor frame calibration module are designed, combined with an anchor frame calibration detector, a street scene video instance segmentation model is constructed, and an instance segmentation is performed by collecting street scene data sets.
The adequacy of edge feature extraction and compensation of spatial location details of feature pyramids at high levels are improved, and detection and segmentation accuracy and speed are improved.
Smart Images

Figure CN115272835B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of driverless, and particularly to a method and system for video instance segmentation of street scenes. Background Technique
[0002] Environmental perception is one of the key issues in the research of driverless technology, providing an important basis for path planning, decision-making, and control execution of vehicles in traffic scenarios. Vehicles need to obtain and process environmental information in real time during driving. Currently, common environmental perception methods are mainly divided into radar, multi-sensor information fusion, and vision methods according to the types of sensors for obtaining environmental information. Radar devices are costly and can only identify depth information without being able to obtain texture and color; the method of multi-sensor information fusion is even more costly and has a high technical difficulty; while the vision method has the advantages of low cost and the ability to obtain texture and color information, so it has broader research and application value.
[0003] However, there are technical problems in the prior art that single receptive field sampling is adopted for multi-type aspect ratio anchor boxes, resulting in insufficient extraction of edge features and lack of detailed information on the spatial position of high-level features in the feature pyramid. Summary of the Invention
[0004] This application provides a method and system for video instance segmentation of street scenes, which solves the technical problems in the prior art that single receptive field sampling is adopted for multi-type aspect ratio anchor boxes, resulting in insufficient extraction of edge features and lack of detailed information on the spatial position of high-level features in the feature pyramid. It achieves the technical effect of solving the problem of insufficient extraction of edge features and compensating for the detailed information on the spatial position of high-level features in the feature pyramid through the video instance segmentation model of street scenes.
[0005] In view of the above problems, this application provides a method and system for video instance segmentation of street scenes.
[0006] In a first aspect, this application provides a method for video instance segmentation of street scenes, the method comprising: constructing a multi-receptive field downsampling module; designing a spatial position information compensation feature pyramid based on the multi-receptive field downsampling module; constructing an anchor box calibration module; designing an anchor box calibration detector based on the anchor box calibration module; constructing a video instance segmentation model of street scenes based on the spatial position information compensation feature pyramid and the anchor box calibration detector; obtaining a street scene data set, and extracting street scene instances based on the street scene data set; using the video instance segmentation model of street scenes to segment the street scene instances.
[0007] On the other hand, the present application provides a street scene video instance segmentation system, which includes: a first construction unit for constructing a multi-receptive field downsampling module; a first execution unit for designing a spatial position information compensation feature pyramid based on the multi-receptive field downsampling module; a second construction unit for constructing an anchor box calibration module; a second execution unit for designing an anchor box calibration detector based on the anchor box calibration module; a third construction unit for constructing a street scene video instance segmentation model based on the spatial position information compensation feature pyramid and the anchor box calibration detector; a third execution unit for obtaining a street scene data set and extracting street scene instances based on the street scene data set; and a fourth execution unit for segmenting the street scene instances using the street scene video instance segmentation model.
[0008] In a third aspect, the present invention provides a street scene video instance segmentation system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in the first aspect are implemented.
[0009] One or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0010] 1. By adopting the technical solution of constructing a multi-receptive field downsampling module; designing a spatial position information compensation feature pyramid using the constructed multi-receptive field downsampling module; constructing an anchor box calibration module; designing an anchor box calibration detector using the constructed anchor box calibration module; constructing a street scene video instance segmentation model according to the spatial position information compensation feature pyramid and the anchor box calibration detector; collecting a street scene data set and extracting street scene instances based on the street scene data set; and further segmenting the street scene instances using the street scene video instance segmentation model, the present application provides a street scene video instance segmentation method and system, achieving the technical effect of solving the problem of insufficient extraction of edge features and compensating for the spatial position details of high-level features of the compensation feature pyramid through the street scene video instance segmentation model.
[0011] 2. By adopting the method of using Baseline (backbone network) as a control and setting up groups of Baseline+SL-FPN, Baseline+AFC, and Baseline+AFC+SL-FPN for experiments, the average precision of detection and segmentation and the segmentation speed are compared, achieving the technical effect that the comprehensive results of detection, segmentation precision, and detection time of the street scene video instance segmentation model (Baseline+AFC+SL-FPN group) constructed based on the spatial position information compensation feature pyramid and the anchor box calibration detector are the best.
[0012] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the specific embodiments of this application are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a flowchart of a method for segmenting street scene video instances according to an embodiment of this application;
[0014] Figure 2 It is an architecture diagram of a multi-receptive field downsampling module for a method for segmenting street scene video instances according to an embodiment of this application;
[0015] Figure 3 It is an architecture diagram of a spatial position information compensation feature pyramid for a method for segmenting street scene video instances according to an embodiment of this application;
[0016] Figure 4 It is an architecture diagram of an anchor box calibration for a method for segmenting street scene video instances according to an embodiment of this application;
[0017] Figure 5 It is an anchor box calibration prediction diagram for a method for segmenting street scene video instances according to an embodiment of this application;
[0018] Figure 6 It is a prototype mask branch diagram for a method for segmenting street scene video instances according to an embodiment of this application;
[0019] Figure 7 It is an overall network structure diagram of a street scene video instance segmentation model for a method for segmenting street scene video instances according to an embodiment of this application;
[0020] Figure 8 It is a structural diagram of a street scene video instance segmentation system according to an embodiment of this application;
[0021] Figure 9 It is a structural diagram of an exemplary electronic device according to an embodiment of this application.
[0022] Explanation of the reference numerals: first building unit 11, first execution unit 12, second building unit 13, second execution unit 14, third building unit 15, third execution unit 16, fourth execution unit 17, electronic device 300, memory 301, processor 302, communication interface 303, bus architecture 304. DETAILED DESCRIPTION
[0023] The present application provides a street scene video instance segmentation method and system, which solves the technical problem in the prior art that multiple types of aspect ratio anchor frames use a single receptive field sampling, resulting in insufficient edge feature extraction and lack of feature pyramid high-level feature spatial position detail information. The street scene video instance segmentation model is used to solve the technical effect of insufficient edge feature extraction and compensating for feature pyramid high-level feature spatial position detail information.
[0024] At present, common environmental perception methods are mainly divided into radar, multi-sensor information fusion and visual methods according to the type of sensor that obtains environmental information. Radar equipment is expensive and can only identify depth information but cannot obtain texture and color; multi-sensor information fusion methods are more expensive and technically difficult; while visual methods have the advantages of low cost and the ability to obtain texture and color information, so they have broader research and application value. In the existing technology, there are technical problems that the single receptive field sampling of multiple types of aspect ratio anchor frames leads to insufficient edge feature extraction and lack of detailed information on the spatial position of high-level features in the feature pyramid.
[0025] In response to the above technical problems, the overall idea of the technical solution provided by this application is as follows:
[0026] The present application provides a street scene video instance segmentation method, the method comprising: constructing a multi-receptive field downsampling module; using the constructed multi-receptive field downsampling module to design a spatial position information compensation feature pyramid; constructing an anchor frame calibration module; using the constructed anchor frame calibration module to design an anchor frame calibration detector; constructing a street scene video instance segmentation model according to the spatial position information compensation feature pyramid and the anchor frame calibration detector; collecting a street scene data set, and extracting street scene instances based on the street scene data set; and further using the street scene video instance segmentation model to segment the street scene instances.
[0027] After introducing the basic principles of the present application, various non-limiting implementation methods of the present application will be specifically described below in conjunction with the drawings in the specification.
[0028] Embodiment 1
[0029] like Figure 1 As shown, an embodiment of the present application provides a method for segmenting street scene video instances, wherein the method comprises:
[0030] Step S100: Construct a multi-receptive field downsampling module;
[0031] Specifically, the receptive field is the size of the area on the input image mapped by the pixel points on the feature map output by each layer of the convolutional neural network. In existing video instance segmentation technologies, single receptive field sampling is often used, resulting in insufficient extraction of edge features. A multi-receptive field downsampling module (MRFS) is constructed. The multi-receptive field downsampling module can perform convolution operations on the target image using convolutional kernels with multiple different receptive fields. After performing convolution operations on the target image using the multi-receptive field downsampling module, the feature maps obtained by downsampling with different receptive fields need to be added, activated, and normalized to obtain the downsampling result.
[0032] Step S200: Design a spatial position information compensation feature pyramid based on the multi-receptive field downsampling module;
[0033] Furthermore, as Figure 3 shown, for the design of the spatial position information compensation feature pyramid based on the multi-receptive field downsampling module, step S200 of the embodiment of the present application includes:
[0034] Step S210: Obtain the features of layers C2 - C5 in the backbone network and compress the number of channels;
[0035] Step S220: Use the compressed C5 layer as the P5 layer;
[0036] Step S230: The P5 layer uses bilinear interpolation upsampling to fuse with the C4 feature map to obtain the P4 layer;
[0037] Step S240: The P4 layer uses bilinear interpolation upsampling and then fuses with the C3 feature map to obtain the P3 layer;
[0038] Step S250: The C2 layer uses the multi-receptive field downsampling module to downsample the feature map, adds it to the P3 layer feature map after convolution processing, and then obtains the F3 layer after another convolution processing;
[0039] Step S260: After the F3 layer undergoes downsampling operations using the multi-receptive field downsampling module, it is added to the P4 after one convolution processing to obtain the F4 layer;
[0040] Step S270: The F4 layer undergoes downsampling operations using the multi-receptive field downsampling module and is added to the P5 after one convolution processing to obtain the F5 layer;
[0041] Step S280: The F5 layer undergoes one convolution to obtain the F6 layer, and the F6 layer undergoes one convolution to obtain the F7 layer.
[0042] Specifically, the feature pyramid is a very important part currently used in object detection, semantic segmentation, action recognition, etc., and has very good performance in improving the model performance. The preferred process for constructing the spatial position information compensated feature pyramid is as follows. Use Ci to mark the feature maps of each layer in the image backbone network, and obtain the features of the C2-C5 layers in the backbone network. Among them, the backbone network is preferably ResNet50. ResNet50 first performs a convolution operation on the input, then contains 4 residual blocks (ResidualBlock), and finally performs a fully connected operation for the classification task. After obtaining the features of the C2-C5 layers, the number of channels is compressed to 256 through a 1×1 convolution kernel. Use Pi to mark each feature map in the image top-down structure. For example, P5 represents the feature map corresponding to the size of C5. The C5 layer after compressing the number of channels is used as the P5 layer. Since the P5 layer needs to be fused with the P4 layer, it is necessary to perform bilinear interpolation upsampling to be consistent with the C4 feature map before fusion to obtain the P4 layer. Similarly, the P4 layer is fused with the C3 feature map after bilinear interpolation upsampling to obtain the P3 layer. The C2 layer uses the multi-receptive field downsampling module to downsample the feature map, then adds it to the P3 layer feature map after one 3×3 convolution operation, and then passes through one 3×3 convolution operation to obtain the F3 layer. The F3 layer undergoes a downsampling operation through the multi-receptive field downsampling module and adds it to the P4 layer after one 3×3 convolution operation to obtain the F4 layer. The F4 layer undergoes a downsampling operation through the multi-receptive field downsampling module and adds it to the P5 layer after one 3×3 convolution operation to obtain the F5 layer. The F5 layer undergoes one convolution to obtain the F6 layer, and the F6 layer undergoes one convolution to obtain the F7 layer.
[0043] The process of obtaining the F3, F4, and F5 layers described above can be expressed by the following formula:
[0044] F i out = Conv 3×3 (MRFS(F i-1 ) + Conv 3×3 (F i ))
[0045] Among them, Conv 3×3 (·) represents a convolution operation with a 3×3 convolution kernel; MRFS(·) is the multi-receptive field downsampling module; F i-1 is the feature map of the previous layer of the shallow layer; F i is the feature map of the current layer; F i outis the output feature map. That is, after the feature map of the current layer undergoes a 3×3 convolution operation and is added to the shallow feature map of the previous layer downsampled by the multi-receptive field downsampling module, and then undergoes another 3×3 convolution operation, the output feature map of the current layer can be obtained. Designing a feature pyramid for compensating spatial position information based on the multi-receptive field downsampling module can effectively improve the problem of lack of detailed spatial position information in the high-level features of the feature pyramid in existing video instance segmentation technologies.
[0046] Step S300: Construct an anchor box calibration module;
[0047] Step S400: Design an anchor box calibration detector based on the anchor box calibration module;
[0048] Specifically, the construction process of the anchor box calibration module is as follows: First, a set of anchor boxes with different scales and positions are preset. As Figure 4 shown, the feature map is convolved using a convolution kernel that matches the aspect ratio of the anchor box to obtain a predetermined number of new feature maps. Further, the output predetermined number of new feature maps are concatenated, and the number of channels of the concatenated feature map is reduced back to the number of channels before concatenation through a 1×1 convolution kernel. An anchor box calibration detector is designed by using multiple anchor box calibration modules. The detection process of the anchor box calibration detector includes, but is not limited to, as Figure 5 shown: The feature map is sent to the Bbox, Class, and Mask detection branches after extracting features through a 3×3 convolution operation, and detection results including bounding boxes, classes, and mask coefficients are obtained based on the Bbox, Class, and Mask detection branches. The Bbox, Class, and Mask detection branches are three anchor box calibration modules. The output channels of the three anchor box calibration modules are different. The output channel number of the Bbox branch is 4a; the output channel number of the Class branch is ca; the output channel number of the Mask branch is ka; where a is the number of anchor boxes generated for each pixel point and takes a value of 3; c is the number of classes; k is the mask coefficient.
[0049] Step S500: Construct a street scene video instance segmentation model based on the spatial position information compensation feature pyramid and the anchor box calibration detector;
[0050] Furthermore, as Figures 6 - 7 shown, Embodiment S500 of the present application further includes:
[0051] Step S510: Use ResNet as the backbone network to extract features;
[0052] Step S520: Perform multi-scale feature fusion using the spatial position information compensation feature pyramid and compensate the detailed spatial position information of the high-level features;
[0053] Step S530: Calibrate the detector based on the anchor boxes to obtain bounding boxes, classes, and mask coefficients;
[0054] Step S540: Generate a prototype mask using the prototype mask branch;
[0055] Step S550: Linearly combine the prototype mask and the mask coefficients to obtain an instance mask.
[0056] Specifically, the overall network structure of the street scene video instance segmentation model is as Figure 7 shown, including: using ResNet as the backbone network to extract features from the target image, compensating the feature pyramid with spatial location information, performing multi-scale feature fusion, and compensating the spatial location detail information of the high-level features. Use the anchor box calibrated detector to obtain the output bounding boxes, classes (such as people, cars, skateboards, dogs, etc.), and mask coefficients. The prototype mask branch is as Figure 6 shown. After generating the prototype mask using the prototype mask branch, linearly combine the prototype mask and the mask coefficients to obtain an instance mask. The instance can be detected and segmented through the instance mask. For example, the construction process of a street scene video instance segmentation model includes, but is not limited to, using ResNet50 as the backbone network and selecting cosine annealing to decay the learning rate. The experimental platform is a desktop computer with the Ubuntu18.04 operating system, 16G of memory, an AMD Ryzen 7 3700X CPU, and an NVIDIA GTX3070 GPU. The deep learning framework is Pytorch1.7, and the CUDA11.0 and cuDNN8.0.5 acceleration toolkits are used to accelerate network training. Based on the Pytorch deep learning framework, a total of 80,000 iterations are performed during training, the batch size is 6, and the initial learning rate is 0.00005. A street scene video instance segmentation model is saved every 10,000 iterations, and the last street scene video instance segmentation model is used to verify the model accuracy.
[0057] Step S600: Obtain a street scene dataset and extract street scene instances based on the street scene dataset;
[0058] Step S700: Segment the street scene instances using the street scene video instance segmentation model.
[0059] Specifically, the street scene dataset is composed of a large number of daily scene videos on the street. Street scene instances are extracted from the street scene dataset, and the street scene instances include people, skateboards, sedans, dogs, trains, motorcycles, trucks, etc. The street scene dataset is divided into a training set and a validation set, and the ratio of the number of instances in the training set to the validation set is a preset ratio, preferably 9:1. The street scene video instance segmentation model is used to segment the street scene instances, and the validation set is used to verify the street scene video instance segmentation model. Through the street scene video instance segmentation model, the problems of insufficient edge feature extraction and lack of detailed information on the spatial position of high-level features in the feature pyramid can be improved.
[0060] Further, as Figure 2 shown, step S100 of the embodiment of the present application further includes:
[0061] Step S110: The multi-receptive field downsampling module includes a first branch, a second branch, a third branch, and a fourth branch;
[0062] Step S120: The first branch, the second branch, the third branch, and the fourth branch are respectively used to extract features from the image, and a first feature map, a second feature map, a third feature map, and a fourth feature map are obtained;
[0063] Step S130: After adding, activating, and normalizing the first feature map, the second feature map, the third feature map, and the fourth feature map, an output result is obtained.
[0064] Further, for obtaining the output result, step S130 of the embodiment of the present application includes:
[0065] Step S131: The formula for obtaining the output result is as follows:
[0066] P cc = Conv i×j (Conv i×j (F in ))
[0067] P pc = Conv i×j (Pool i×j (F in ))
[0068]
[0069] Among them, P cc represents two cascaded convolutions; P pc represents pooling first and then convolution; Conv i×j (·) represents a convolution operation with a convolution kernel of i×j; Pool i×j(·) represents a pooling operation with a pooling kernel of i×j; F in represents the input feature map; MRFS represents the final output; Relu is the activation function; BN is the normalization.
[0070] Specifically, the multi-receptive field downsampling module adopts four branches, among which three branches are convolutional operation samplings, and the difference lies in the different sizes of the convolutional kernels, that is, different receptive fields. The first branch first uses a 1×1 convolutional kernel to perform a convolutional operation on the feature map, and its main function is to achieve cross-channel information combination and increase non-linear features. Then, a convolutional operation with a 3×3 convolutional kernel and a stride of 2 is used to downsample the image to obtain the first feature map. The second branch first uses a 5×3 convolutional kernel with a stride of (2,1) to sample the length (w) direction of the feature map, and then uses a 3×5 convolutional kernel with a stride of (1,2) to sample the width (h) direction of the feature map to obtain the second feature map. The third branch first uses a 3×1 convolutional kernel with a stride of (2,1) to sample the length (w) direction of the feature map, and then uses a 1×3 convolutional kernel with a stride of (1,2) to sample the width (h) direction of the feature map to obtain the third feature map. The fourth branch first uses an average pooling operation with a kernel of 3×3 and a stride of 2, and then uses a 1×1 convolution to increase the non-linear features of the feature map to obtain the fourth feature map. Finally, the feature maps obtained from the four branches are added, activated, and normalized to obtain the output result. The formula for obtaining the output result is as follows:
[0071] P cc =Conv i×j (Conv i×j (F in ))
[0072] P pc =Conv i×j (Pool i×j (F in ))
[0073]
[0074] Among them, P cc represents two cascaded convolutions; P pc represents pooling first and then convolution; Conv i×j (·) represents a convolutional operation with a convolutional kernel of i×j; Pool i×j (·) represents a pooling operation with a pooling kernel of i×j; F in represents the input feature map; MRFS represents the final output; Relu is the activation function; BN is the normalization. It can be seen from that the final output is the sum of three P cc plus Ppc After that, it is normalized after being activated by the Relu function, where P cc is the concatenation of two convolutions on the input feature map, P pc is to first perform pooling on the input feature map and then perform convolution.
[0075] Furthermore, as Figure 4 shown, for constructing the anchor box calibration module, step S300 of the embodiment of the present application further includes:
[0076] Step S310: Convolve the feature map with a convolution kernel matching the aspect ratio of the anchor box to obtain a predetermined number of output feature maps;
[0077] Step S320: Concatenate the output feature maps, and use a convolution kernel of the first preset size to reduce the number of channels of the concatenated feature map back to the number of channels before concatenation;
[0078] Step S330: The anchor box calibration module can be expressed as:
[0079] F out = Conv 1×1 (Cat(Conv 5×3 (F in ) + Conv 3×5 (F in ) + Conv 3×3 (F in )))
[0080] where Conv i×j (·) represents a convolution operation with a convolution kernel of i×j; Cat(·) represents concatenating the feature maps; F in represents the input feature map; F out represents the output feature map.
[0081] Furthermore, as Figure 5 shown, for designing the anchor box calibration detector based on the anchor box calibration module, step S330 includes:
[0082] Step S331: Design the anchor box calibration detector using the Bbox, Class, and Mask detection branches. Among them, the Bbox, Class, and Mask detection branches are three anchor box calibration modules. The output channel number of the Bbox branch is 4a; the output channel number of the Class branch is ca; the output channel number of the Mask branch is ka; where a is the number of anchor boxes generated per pixel, taking the value of 3; c is the number of categories; k is the mask coefficient.
[0083] Specifically, as Figure 4As shown, after the feature map is convolved with convolutional kernels that match the aspect ratios of the anchor boxes, a predetermined number of output feature maps are obtained. The predetermined number of output feature maps is preferably three, that is, three new feature maps are obtained after convolution with three convolutional kernels. Further, the three output new feature maps are concatenated. The first preset size convolutional kernel is a 1×1 convolutional kernel, and the number of channels of the feature map after concatenation is reduced back to the number of channels before concatenation through a 1×1 convolutional kernel. The anchor box calibration module can be expressed as:
[0084] F out =Conv 1×1 (Cat(Conv 5×3 (F in )+Conv 3×5 (F in )+Conv 3×3 (F in )))
[0085] Among them, Conv i×j (·) represents the convolution operation with a convolutional kernel of i×j; Cat(·) represents concatenating the feature maps; F in represents the input feature map; F out represents the output feature map. Through the above formula, the three input feature maps are respectively convolved with convolutional kernels that match the aspect ratios of the anchor boxes. The sizes of the convolutional kernels are 5×3, 3×5, and 3×3 respectively. After convolution, they are concatenated through the Cat(·) command, and then the number of channels is restored through a 1×1 convolutional kernel.
[0086] Furthermore, an anchor box calibration detector is designed using three detection branches: Bbox, Class, and Mask. Among them, the Bbox, Class, and Mask detection branches are three anchor box calibration modules. The output channel number of the Bbox branch is 4a; the output channel number of the Class branch is ca; the output channel number of the Mask branch is ka; where a is the number of anchor boxes generated per pixel and takes the value of 3; c is the number of categories; k is the mask coefficient. The Bbox detection branch is used to detect the bounding box, the Class detection branch is used to detect the category, and the Mask detection branch is used to detect the mask coefficient. Since a is the number of anchor boxes generated per pixel and takes the value of 3, the output channel number of the Bbox branch can be obtained as 12, the output channel number of the Class branch is 3c, and the output channel number of the Mask branch is 3k.
[0087] In summary, the street scene video instance segmentation method and system provided by the embodiments of the present application have the following technical effects:
[0088] 1. Due to the adoption of the construction of a multi-receptive field downsampling module; using the constructed multi-receptive field downsampling module to design a spatial position information compensation feature pyramid; constructing an anchor box calibration module; using the constructed anchor box calibration module to design an anchor box calibration detector; according to the spatial position information compensation feature pyramid and the anchor box calibration detector, constructing a street scene video instance segmentation model; collecting a street scene data set, and extracting street scene instances based on the street scene data set; and further using the street scene video instance segmentation model to segment the street scene instances, the present application provides a street scene video instance segmentation method and system, achieving the technical effect of solving the problem of insufficient edge feature extraction and compensating the spatial position detail information of the high-level features of the feature pyramid through the street scene video instance segmentation model.
[0089] Embodiment 2
[0090] Taking the Youtube-VIS public data set as the data source, extracting the required data from the Youtube-VIS public data set, constructing a street scene data set, and verifying the effect of a street scene video instance segmentation method. The specific process is as follows:
[0091] Extract the required street scene instances from the label file provided by the Youtube-VIS data set; and divide the street scene data set into a training set and a validation set, and the ratio of the number of instances in the training set to the validation set is 9:1.
[0092] Detailed data of the data set used in the experiment: The instances are divided into seven categories including people, skateboards, cars, dogs, trains, motorcycles, and trucks. Among them, the training set has 329 video clips, with a total number of frames of 7212, containing 603 instances, and the validation set has 53 video clips, with a total number of frames of 1097, containing 88 instances. The specific data is shown in Table 1, and the ratio of the number of instances in the training set to the validation set is 9:1.
[0093] Table 1 Number of instances in the street scene data set
[0094]
[0095] Experimental parameters: ResNet50 is used for the backbone network; the cosine annealing learning rate decay is selected. The experimental platform is a desktop computer with Ubuntu 18.04 operating system, 16G of memory, AMD Ryzen 7 3700X CPU, and NVIDIA GTX 3070 GPU. The deep learning framework is Pytorch 1.7, and the CUDA 11.0 and cuDNN 8.0.5 acceleration toolkits are used to accelerate network training. Based on the Pytorch deep learning framework, the training is iterated 80,000 times in total, the batch size is 6, and the initial learning rate is 0.00005. A model is saved every 10,000 iterations, and the last model is used to verify the model accuracy.
[0096] Table 2 Model verification results
[0097]
[0098] As shown in Table 2, compared with the Baseline results, it can be seen that the Spatial Location Information Compensated Feature Pyramid (SL-FPN) improves the average accuracy of network detection and segmentation by 7.76% and 4.75% respectively; the Anchor Frame Calibration (AFC) improves the average accuracy of network detection and segmentation by 8.63% and 5.09% respectively; when the Spatial Location Information Compensated Feature Pyramid and the Anchor Frame Calibration are used together, the average detection and segmentation reach 38.73% and 37.03% respectively, and compared with the Baseline, the average accuracy of detection and segmentation is improved by 9.26% and 6.46% respectively, and the speed reaches 26 frames per second for segmentation.
[0099] Embodiment III
[0100] Based on the same inventive concept as a street scene video instance segmentation method in the foregoing embodiments, as Figure 8 shown, the embodiment of the present application provides a street scene video instance segmentation system, wherein the system includes:
[0101] The first construction unit 11 is used to construct a multi-receptive field downsampling module;
[0102] The first execution unit 12 is used to design a spatial location information compensated feature pyramid based on the multi-receptive field downsampling module;
[0103] The second construction unit 13 is used to construct an anchor frame calibration module;
[0104] The second execution unit 14 is used to design an anchor frame calibration detector based on the anchor frame calibration module;
[0105] The third construction unit 15, which is used to construct a street scene video instance segmentation model based on the spatial position information compensation feature pyramid and the anchor box calibration detector;
[0106] The third execution unit 16, which is used to obtain a street scene data set and extract street scene instances based on the street scene data set;
[0107] The fourth execution unit 17, which is used to segment the street scene instances using the street scene video instance segmentation model.
[0108] Furthermore, the system includes:
[0109] The first acquisition unit, which is used to respectively perform feature extraction on an image using the first branch, the second branch, the third branch, and the fourth branch to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map;
[0110] The second acquisition unit, which is used to add, activate, and normalize the first feature map, the second feature map, the third feature map, and the fourth feature map to obtain an output result.
[0111] Furthermore, the system includes:
[0112] The fifth execution unit, which is used to obtain the features of the C2-C5 layers in the backbone network and compress the number of channels;
[0113] The sixth execution unit, which is used to use the compressed C5 layer as the P5 layer;
[0114] The third acquisition unit, which is used to fuse the P5 layer with the C4 feature map using bilinear interpolation upsampling to obtain the P4 layer;
[0115] The fourth acquisition unit, which is used to fuse the P4 layer with the C3 feature map after using bilinear interpolation upsampling to obtain the P3 layer;
[0116] The fifth acquisition unit, which is used to downsample the feature map of the C2 layer using the multi-receptive field downsampling module, add it to the feature map of the P3 layer after convolution processing, and then obtain the F3 layer after another convolution processing;
[0117] The sixth acquisition unit, which is used to downsample the F3 layer using the multi-receptive field downsampling module and add it to the P4 layer after one convolution processing to obtain the F4 layer;
[0118] A seventh acquisition unit, which is used to add the downsampling operation of the F4 layer through the multi-receptive field downsampling module and the P5 after one convolution process to obtain the F5 layer;
[0119] An eighth acquisition unit, which is used to obtain the F6 layer by convolving the F5 layer once, and obtain the F7 layer by convolving the F6 layer once.
[0120] Further, the system includes:
[0121] A ninth acquisition unit, which is used to convolve the feature map with a convolution kernel matching the aspect ratio of the anchor box to obtain a predetermined number of output feature maps;
[0122] A seventh execution unit, which is used to splice the output feature maps and use a convolution kernel of a first preset size to reduce the number of channels of the spliced feature map back to the number of channels before splicing.
[0123] Further, the system includes:
[0124] An eighth execution unit, which is used to design the anchor box calibration detector using the Bbox, Class, and Mask detection branches, where the Bbox, Class, and Mask detection branches are three anchor box calibration modules, the output channel number of the Bbox branch is 4a; the output channel number of the Class branch is ca; the output channel number of the Mask branch is ka; where a is the number of anchor boxes generated per pixel and takes a value of 3; c is the number of categories; k is the mask coefficient.
[0125] Further, the system includes:
[0126] A ninth execution unit, which is used to extract features using ResNet as the backbone network;
[0127] A tenth execution unit, which is used to perform multi-scale feature fusion using the spatial position information compensation feature pyramid and compensate for the spatial position detail information of the high-level features;
[0128] A tenth acquisition unit, which is used to obtain the bounding box, category, and mask coefficient based on the anchor box calibration detector;
[0129] A first generation unit, which is used to generate a prototype mask using the prototype mask branch;
[0130] An eleventh acquisition unit, which is used to linearly combine the prototype mask and the mask coefficient to obtain an instance mask.
[0131] Exemplary electronic device
[0132] The following will refer to Figure 9 to describe the electronic device according to the embodiments of the present application.
[0133] Based on the same inventive concept as the method for video instance segmentation of a street scene in the foregoing embodiments, the embodiments of the present application further provide a system for video instance segmentation of a street scene, including: a processor, the processor is coupled to a memory, and the memory is used to store a program. When the program is executed by the processor, the system is enabled to execute the method according to any one of the first aspect.
[0134] The electronic device 300 includes: a processor 302, a communication interface 303, and a memory 301. Optionally, the electronic device 300 may further include a bus architecture 304. Among them, the communication interface 303, the processor 302, and the memory 301 may be interconnected through the bus architecture 304; the bus architecture 304 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus architecture 304 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0135] The processor 302 may be a CPU, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application solution.
[0136] The communication interface 303 uses any device such as a transceiver for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), wired access network, etc.
[0137] The memory 301 can be a ROM or other types of static storage devices that can store static information and instructions, a RAM or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor through the bus architecture 304. The memory can also be integrated with the processor.
[0138] Among them, the memory 301 is used to store computer execution instructions for implementing the solution of this application, and is controlled by the processor 302 for execution. The processor 302 is used to execute the computer execution instructions stored in the memory 301, so as to implement a street scene video instance segmentation method provided in the above embodiments of this application.
[0139] Optionally, the computer execution instructions in the embodiments of this application can also be referred to as application program codes, and the embodiments of this application do not make specific limitations thereto.
[0140] The embodiments of this application provide a street scene video instance segmentation method, where the method includes: constructing a multi-receptive field downsampling module; using the constructed multi-receptive field downsampling module to design a spatial position information compensation feature pyramid; constructing an anchor box calibration module; using the constructed anchor box calibration module to design an anchor box calibration detector; constructing a street scene video instance segmentation model according to the spatial position information compensation feature pyramid and the anchor box calibration detector; collecting a street scene data set, and extracting street scene instances based on the street scene data set; and further using the street scene video instance segmentation model to segment the street scene instances.
[0141] Those of ordinary skill in the art can understand that the various numerical numbers such as the first and second involved in this application are only for the convenience of description and are not used to limit the scope of the embodiments of this application, nor do they represent the order of precedence. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one" means one or more. At least two means two or more. "At least one", "any one" or their similar expressions refer to any combination of these items, including any combination of single items (pieces) or plural items (pieces). For example, at least one (piece, type) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0142] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0143] In the embodiments of the present application, the various illustrative logical units and circuits described can be implemented or operate the described functions by a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other similar configuration.
[0144] The steps of the methods or algorithms described in the embodiments of the present application can be directly embedded in hardware, a software unit executed by a processor, or a combination of the two. The software unit can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be disposed in an ASIC, and the ASIC can be disposed in a terminal. Optionally, the processor and the storage medium can also be disposed in different components of the terminal. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, thereby providing instructions executed on the computer or other programmable device for implementing the steps for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 the steps of the functions specified in one block or multiple blocks.
[0145] Although the present application has been described in conjunction with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application and are considered to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.
Claims
1. A method for video instance segmentation of street scenes, characterized in that, The method includes: Constructing a multi-receptive field downsampling module; Designing a spatial position information compensation feature pyramid based on the multi-receptive field downsampling module, including: Obtaining the features of the C2-C5 layers in the backbone network and compressing the number of channels; Taking the C5 layer after compressing the number of channels as the P5 layer; The P5 layer uses bilinear interpolation upsampling to fuse with the C4 feature map to obtain the P4 layer; The P4 layer uses bilinear interpolation upsampling and then fuses with the C3 feature map to obtain the P3 layer; The C2 layer uses the multi-receptive field downsampling module to downsample the feature map, adds it to the P3 layer feature map after convolution processing, and then obtains the F3 layer after another convolution processing; The F3 layer undergoes a downsampling operation through the multi-receptive field downsampling module and adds it to the P4 after one convolution processing to obtain the F4 layer; The F4 layer undergoes a downsampling operation through the multi-receptive field downsampling module and adds it to the P5 after one convolution processing to obtain the F5 layer; The F5 layer undergoes one convolution to obtain the F6 layer, and the F6 layer undergoes one convolution to obtain the F7 layer; Constructing an anchor box calibration module, including: Applying a convolution kernel matching the aspect ratio of the anchor box to perform convolution on the feature map to obtain a predetermined number of output feature maps; Concatenating the output feature maps and using a first preset size convolution kernel to reduce the number of channels of the concatenated feature map back to the number of channels before concatenation; The anchor box calibration module can be expressed as: F out = Conv 1×1 (Cat(Conv 5×3 (F in ) + Conv 3×5 (F in ) + Conv 3×3 (F in ))) Among them, Conv i×j (·) represents a convolution operation with a convolution kernel of i×j; Cat(·) represents concatenating feature maps; F in represents the input feature map; F out represents the output feature map; Designing an anchor box calibration detector based on the anchor box calibration module, including: Using the Bbox, Class, and Mask detection branches to design the anchor box calibration detector, where the Bbox, Class, and Mask detection branches are three anchor box calibration modules, the output channel number of the Bbox branch is 4a; the output channel number of the Class branch is ca; the output channel number of the Mask branch is ka; where a is the number of anchor boxes generated per pixel and takes the value of 3; c is the number of categories; k is the mask coefficient; Constructing a street scene video instance segmentation model based on the spatial position information compensation feature pyramid and the anchor box calibration detector; Obtaining a street scene dataset and extracting street scene instances based on the street scene dataset; Using the street scene video instance segmentation model to segment the street scene instances.
2. The method according to claim 1, characterized in that, The method further includes: The multi-receptive field downsampling module includes a first branch, a second branch, a third branch, and a fourth branch; Using the first branch, the second branch, the third branch, and the fourth branch to extract features from the image respectively to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map; Adding, activating, and normalizing the first feature map, the second feature map, the third feature map, and the fourth feature map to obtain an output result.
3. The method according to claim 2, wherein For obtaining the output result, the method further includes: The formula for obtaining the output result is as follows: P cc = Conv i×j (Conv i×j (F in )) P pc = Conv i×j (Pool i×j (F in )) Among them, P cc is expressed as two cascaded convolutions; P pc is expressed as pooling first and then convolution; Conv i×j () represents a convolution operation with a convolution kernel of i×j; Pool i×j () represents a pooling operation with a pooling kernel of i×j; F in represents the input feature map; MRFS represents the final output; Relu is the activation function; BN is the normalization.
4. The method according to claim 1, wherein The method further includes: Using ResNet as the backbone network to extract features; Using the spatial position information compensation feature pyramid for multi-scale feature fusion and compensating the spatial position detail information of the high-level features. Based on the anchor box calibration detector, obtain the bounding box, category, and mask coefficient; Use the prototype mask branch to generate a prototype mask; Linearly combine the prototype mask with the mask coefficient to obtain an instance mask.
5. A street scene video instance segmentation system, characterized in that, The system includes: A first construction unit for constructing a multi-receptive field downsampling module; A first execution unit for designing a spatial position information compensation feature pyramid based on the multi-receptive field downsampling module; including: Obtain the features of the C2-C5 layers in the backbone network and compress the number of channels; Use the C5 layer with compressed channel number as the P5 layer; The P5 layer uses bilinear interpolation upsampling to fuse with the C4 feature map to obtain the P4 layer; The P4 layer uses bilinear interpolation upsampling and then fuses with the C3 feature map to obtain the P3 layer; The C2 layer uses the multi-receptive field downsampling module to downsample the feature map, add it to the feature map of the P3 layer after convolution processing, and then obtain the F3 layer after another convolution processing; The F3 layer is downsampled by the multi-receptive field downsampling module and added to the P4 after one convolution processing to obtain the F4 layer; The F4 layer is downsampled by the multi-receptive field downsampling module and added to the P5 after one convolution processing to obtain the F5 layer; The F5 layer undergoes one convolution to obtain the F6 layer, and the F6 layer undergoes one convolution to obtain the F7 layer; A second construction unit for constructing an anchor box calibration module; including: Apply a convolution kernel matching the aspect ratio of the anchor box to convolve the feature map to obtain a predetermined number of output feature maps; Concatenate the output feature maps and use a first preset size convolution kernel to reduce the number of channels of the concatenated feature map back to the number of channels before concatenation; The anchor box calibration module can be expressed as: F out = Conv 1×1 (Cat(Conv 5×3 (F in ) + Conv 3×5 (F in ) + Conv 3×3 (F in ))) Among them, Conv i×j (·) represents a convolution operation with a convolution kernel of i×j; Cat(·) represents concatenating feature maps; F in represents the input feature map; F out represents the output feature map; A second execution unit for designing an anchor box calibration detector based on the anchor box calibration module; including: Use the Bbox, Class, and Mask detection branches to design the anchor box calibration detector, where the Bbox, Class, and Mask detection branches are three anchor box calibration modules, the output channel number of the Bbox branch is 4a; the output channel number of the Class branch is ca; the output channel number of the Mask branch is ka; where a is the number of anchor boxes generated per pixel and takes a value of 3; c is the number of categories; k is the mask coefficient; A third construction unit for constructing a street scene video instance segmentation model based on the spatial position information compensation feature pyramid and the anchor box calibration detector; A third execution unit for obtaining a street scene data set and extracting street scene instances based on the street scene data set; A fourth execution unit for segmenting the street scene instances using the street scene video instance segmentation model.
6. A street scene video instance segmentation system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1-4.