Traffic scene weakly supervised video instance segmentation method and system

By constructing a weakly supervised loss module based on projection loss function and color similarity loss function, and combining it with a location feature refinement module and a grouping control module, the problem of large target location information acquisition error in weakly supervised video instance segmentation network is solved, and higher accuracy traffic scene instance segmentation is achieved.

CN116805407BActive Publication Date: 2026-05-15KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
KUNMING UNIV OF SCI & TECH
Filing Date
2023-06-27
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing weakly supervised video instance segmentation networks suffer from a lack of pixel-level annotation information and insufficient supervision of bounding box information, leading to large errors in target location information acquisition and low mask segmentation quality.

Method used

We construct a weakly supervised loss module based on projection loss function and color similarity loss function, and combine it with a location feature refinement module, a multi-scale localization feature pyramid and a grouping control module to build a weakly supervised video instance segmentation model for traffic scenes. This model is used to segment traffic scene instances.

Benefits of technology

It improves the accuracy of segmentation and localization, as well as the precision of mask segmentation, thereby enhancing the quality of weakly supervised video instance segmentation in traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116805407B_ABST
    Figure CN116805407B_ABST
Patent Text Reader

Abstract

The application discloses a kind of traffic scene weak supervision video instance segmentation method and system, it is related to weak supervision video instance segmentation field, wherein the method includes: constructing weak supervision loss module;Based on position feature refinement module, construct multi-scale positioning feature pyramid;Grouping control module is constructed;Based on weak supervision loss module, multi-scale positioning feature pyramid and grouping control module, traffic scene weak supervision video instance segmentation model is constructed;Based on traffic scene dataset extraction traffic scene instance;The traffic scene instance is input into the traffic scene weak supervision video instance segmentation model, and the traffic scene weak supervision video instance segmentation model is used to the traffic scene instance is segmented.Solve the current due to weak supervision video instance segmentation network pixel level annotation information is deficient, and boundary box information supervision intensity is insufficient, leading to network segmentation when obtaining target position information exists error, and the technical problem that mask segmentation quality is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of weakly supervised video instance segmentation, and more specifically, to a method and system for weakly supervised video instance segmentation in traffic scenes. Background Technology

[0002] Environmental perception technology is crucial for the safety of autonomous vehicles, providing essential judgment for path planning and decision-making in traffic scenarios. Vehicles equipped with autonomous driving systems need to accurately perceive road conditions and surrounding targets during operation, enabling the system to monitor its environment in real time and make correct decisions in various situations. Common environmental perception technologies include multi-line LiDAR sensors, millimeter-wave radar, single-line LiDAR, ultrasonic radar, multi-sensor information fusion, and visual recognition. However, radar technology requires expensive equipment and can only acquire depth information of scene targets, failing to process color and texture information. Multi-sensor information fusion requires the fusion of sampling data from multiple sensors, which is even more costly and technically challenging. Visual recognition, as a rising star in autonomous driving environmental perception technology, is low-cost and can effectively process information such as the color and texture of targets in the environment. With the rapid development of deep learning, visual recognition has broad research and application value as an environmental perception technology.

[0003] Currently, the training process of video instance segmentation methods used in visual recognition technology requires high-precision mask labels, which is too costly. Using weakly supervised video instance segmentation networks as a visual recognition technology for autonomous driving systems can solve this problem. However, existing technologies suffer from technical problems such as the lack of pixel-level annotation information and insufficient supervision of bounding box information in weakly supervised video instance segmentation networks, which leads to errors in obtaining target location information and low mask segmentation quality during network segmentation. Summary of the Invention

[0004] This application provides a method and system for weakly supervised video instance segmentation in traffic scenes. It solves the technical problems of current weakly supervised video instance segmentation networks, which suffer from insufficient pixel-level annotation information and inadequate bounding box supervision, leading to errors in target location information acquisition and low mask segmentation quality.

[0005] In view of the above problems, this application provides a method and system for weakly supervised video instance segmentation in traffic scenes.

[0006] Firstly, this application provides a method for weakly supervised video instance segmentation in traffic scenes. The method is applied to a weakly supervised video instance segmentation system for traffic scenes. The method includes: constructing a weakly supervised loss module based on a projection loss function and a color similarity loss function; constructing a location feature refinement module; constructing a multi-scale localization feature pyramid based on the location feature refinement module; constructing a grouping and adjustment module, and placing the grouping and adjustment module after a mask branch, wherein the grouping and adjustment module performs deep semantic parsing of the input features through channel grouping and weighted adjustment; constructing a weakly supervised video instance segmentation model for traffic scenes based on the weakly supervised loss module, the multi-scale localization feature pyramid, and the grouping and adjustment module; obtaining a traffic scene dataset and extracting traffic scene instances based on the traffic scene dataset; inputting the traffic scene instances into the weakly supervised video instance segmentation model for segmentation; and segmenting the traffic scene instances using the weakly supervised video instance segmentation model.

[0007] Secondly, this application also provides a weakly supervised video instance segmentation system for traffic scenes, wherein the system includes: a first construction module, which is used to construct a weakly supervised loss module based on a projection loss function and a color similarity loss function; a second construction module, which is used to construct a location feature refinement module; a third construction module, which is used to construct a multi-scale localization feature pyramid based on the location feature refinement module; a fourth construction module, which is used to construct a grouping and adjustment module and place the grouping and adjustment module after a mask branch, wherein the grouping and adjustment module performs deep semantic parsing on the input features through channel grouping and weighted adjustment; a fifth construction module, which is used to construct a weakly supervised video instance segmentation model for traffic scenes based on the weakly supervised loss module, the multi-scale localization feature pyramid, and the grouping and adjustment module; a scene instance extraction module, which is used to obtain a traffic scene dataset and extract traffic scene instances based on the traffic scene dataset; and a segmentation module, which is used to input the traffic scene instances into the weakly supervised video instance segmentation model for traffic scenes and segment the traffic scene instances through the weakly supervised video instance segmentation model.

[0008] Thirdly, this application also provides a traffic scene weakly supervised video instance segmentation device, including: a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the traffic scene weakly supervised video instance segmentation method provided in this application.

[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0010] A weakly supervised loss module is constructed based on projection loss and color similarity loss functions; a multi-scale localization feature pyramid is constructed based on a location feature refinement module; a grouping and adjustment module is constructed; and a weakly supervised video instance segmentation model for traffic scenes is constructed based on the weakly supervised loss module, the multi-scale localization feature pyramid, and the grouping and adjustment module. Traffic scene instances are extracted from a traffic scene dataset; these traffic scene instances are input into the weakly supervised video instance segmentation model, and the model segments the traffic scene instances. This achieves the technical effect of improving the accuracy of segmentation and localization, improving the precision of mask segmentation, and enhancing the overall quality of weakly supervised video instance segmentation in traffic scenes.

[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments of this disclosure will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.

[0013] Figure 1 This is a flowchart illustrating a weakly supervised video instance segmentation method for traffic scenes according to this application;

[0014] Figure 2 This is a schematic diagram of the projection loss in a weakly supervised video instance segmentation method for traffic scenes according to this application;

[0015] Figure 3 This is a schematic diagram of color similarity loss in a weakly supervised video instance segmentation method for traffic scenes according to this application;

[0016] Figure 4 This is a diagram illustrating the architecture of the multi-scale localization feature pyramid in a weakly supervised video instance segmentation method for traffic scenes proposed in this application.

[0017] Figure 5 This is an architecture diagram of the group control module in a weakly supervised video instance segmentation method for traffic scenes according to this application;

[0018] Figure 6 This is an architecture diagram for obtaining instance masks in a weakly supervised video instance segmentation method for traffic scenes according to this application;

[0019] Figure 7This is an architecture diagram of the location feature refinement module in a weakly supervised video instance segmentation method for traffic scenes according to this application;

[0020] Figure 8 This is a schematic diagram of the structure of a weakly supervised video instance segmentation system for traffic scenes according to this application;

[0021] Figure 9 This is a schematic diagram of the structure of a weakly supervised video instance segmentation device for a traffic scene according to this application.

[0022] Explanation of reference numerals in the attached drawings: First building module 11, Second building module 12, Third building module 13, Fourth building module 14, Fifth building module 15, Scene instance extraction module 16, Segmentation module 17, Electronic device 300, Bus 301, Processor 302, Transceiver 303, Bus interface 304, Memory 305, Operating system 306, Application program 307, and User interface 308. Detailed Implementation

[0023] This application provides a method and system for weakly supervised video instance segmentation in traffic scenes. It addresses the technical problems of current weakly supervised video instance segmentation networks, which suffer from insufficient pixel-level annotation information and inadequate bounding box supervision, leading to errors in target location acquisition and low mask segmentation quality. The method achieves the technical effect of segmenting traffic scene instances using a weakly supervised video instance segmentation model, improving segmentation and localization accuracy, mask segmentation precision, and overall segmentation quality.

[0024] Example 1

[0025] Please see the appendix Figure 1 This application provides a method for segmenting weakly supervised video instances in traffic scenes. The method is applied to a weakly supervised video instance segmentation system for traffic scenes, and specifically includes the following steps:

[0026] Step S100: Construct a weakly supervised loss module based on the projection loss function and the color similarity loss function;

[0027] The projection loss function includes:

[0028]

[0029]

[0030] L XY =L X +L Y

[0031] Where W and H represent the width and height of the image; x n and yn s represents the projection vectors of the maximum values ​​of the true bounding box on the x and y axes, respectively; n L represents the projection of the maximum probability that a pixel in the corresponding direction is a foreground pixel; X and L Y These represent the projection losses along the x-axis and y-axis, respectively; L XY This represents the final projection loss.

[0032] The color similarity loss function includes:

[0033] P(y e =1)=m c,d ·m e,f +(1-m c,d )·(1-m e,f )

[0034]

[0035] Among them, y e This represents an edge e in an undirected graph constructed from image pixels; m c,d Let m represent the probability that one of the endpoints of edge e is a foreground target; e,f Let P(y) represent the probability that the other endpoint of edge e is a foreground target; e =1) represents the probability that edge e is labeled 1, where in an undirected graph, an edge labeled 1 indicates that the two endpoints maintain consistent class, and vice versa; E in The set of edges within the bounding box that contain at least one pixel; N is the set of edges within the bounding box. in The number of sides; L cs This represents the loss due to color similarity.

[0036] Specifically, a projection loss function and a color similarity loss function are constructed and added to the weakly supervised loss module. See attached diagram. Figure 2 As shown, the projection loss function uses the true annotations of instance bounding boxes to supervise the environmental category of image pixels, constraining the instance distribution within the bounding boxes, improving the segmentation accuracy of the network's instance prediction mask, and solving the problem of poor segmentation accuracy caused by the small proportion of instance targets in traffic scenes. (See attached image) Figure 3 As shown, the color similarity loss function effectively constrains the pairwise affinity loss by calculating the color similarity weight matrix, thus assisting the pairwise affinity loss in supervising the instance mask and effectively distinguishing the environmental categories of instances within the bounding box. By combining the projection loss function and the color similarity loss function, effective instance segmentation of video can be achieved without mask annotation, improving mask segmentation accuracy.

[0037] The projection loss function includes:

[0038]

[0039]

[0040] L XY =L X +L Y

[0041] Where W and H represent the width and height of the image; x n and y n s represents the projection vectors of the maximum values ​​of the true bounding box on the x and y axes, respectively; n L represents the projection of the maximum probability that a pixel in the corresponding direction is a foreground pixel; X and L Y These represent the projection losses along the x-axis and y-axis, respectively; L XY This represents the final projection loss.

[0042] The color similarity loss function includes:

[0043] P(y e =1)=m c,d ·m e,f +(1-m c,d )·(1-m e,f )

[0044]

[0045] Among them, y e This represents an edge e in an undirected graph constructed from image pixels; m c,d Let m represent the probability that one of the endpoints of edge e is a foreground target; e,f Let P(y) represent the probability that the other endpoint of edge e is a foreground target; e =1) represents the probability that edge e is labeled 1, where in an undirected graph, an edge labeled 1 indicates that the two endpoints maintain consistent class, and vice versa; E in The set of edges within the bounding box that contain at least one pixel; N is the set of edges within the bounding box. in The number of sides; L cs This represents the loss due to color similarity.

[0046] Step S200: Construct a location feature refinement module;

[0047] Specifically, the location feature refinement module uses multi-level dilated convolutions to extract receptive fields of different ranges from the target feature map, obtains multi-scale feature semantic information, and uses the extracted multi-scale feature semantic information to make up for the instance semantic information missing in the high-level features.

[0048] Step S300: Based on the location feature refinement module, construct a multi-scale positioning feature pyramid;

[0049] Further details are attached. Figure 4 As shown, step S300 of this application further includes:

[0050] Step S310: Set ResNet50 as the backbone network and obtain the C1, C2, C3, C4 and C5 layers in the backbone network;

[0051] Step S320: Compress the number of channels in layers C3, C4, and C5 using convolution;

[0052] Step S330: Input the C1 layer into the location feature refinement module to obtain multi-scale information feature one, multi-scale information feature two, and multi-scale information feature three;

[0053] Step S340: Set the C5 layer after channel compression as the P5 layer;

[0054] Step S350: After performing bilinear interpolation upsampling on the P5 layer, it is fused with the C4 layer to obtain the P4 layer;

[0055] Step S360: After performing bilinear interpolation upsampling on the P4 layer, merge it with the C3 layer to obtain the P3 layer;

[0056] Step S370: Input the multi-scale information feature one, the multi-scale information feature two, and the multi-scale information feature three into the P3 layer, the P4 layer, and the P5 layer to obtain the P'3 layer, P'4 layer, and P'5 layer with multi-scale spatial location information;

[0057] Step S380: Perform convolution on layers P'3, P'4, and P'5 to obtain layers F3, F4, and F5;

[0058] Step S390: Perform two consecutive convolutions on the F5 layer to obtain the F6 and F7 layers.

[0059] Specifically, ResNet50 is selected as the backbone network. A 7×7 convolutional layer and a 3×3 max pooling layer are used to extract features from the initial image data input to the network to obtain layer C1. Layer C1 is then passed through four blocks of ResNet50 to obtain layers C2, C3, C4, and C5. Subsequently, layers C2, C3, C4, and C5 are each passed through 1×1 convolutional layers to compress the number of feature channels of each layer from the original {256, 512, 1024, 2048} to 256. Finally, the top layer feature map is upsampled by bilinear interpolation from top to bottom and added to the feature map of the next layer. The feature maps of the subsequent layers are also obtained by the same method to obtain layers P3, P4, and P5.

[0060] The location feature refinement module is used to supplement the spatial location information of the low-level feature C1 layer to the P3, P4 and P5 layers respectively. In practice, the C1 layer is input into the location feature refinement module to generate sparse feature maps of receptive fields at different scales, and then added to the P3, P4 and P5 layers respectively. 3×3 convolution is used to extract features to obtain the F3, F4 and F5 layers output by the multi-scale localization feature pyramid. Finally, the F5 layer is passed through two consecutive convolutional layers to obtain the F6 and F7 layers respectively.

[0061] Step S400: Construct a grouping control module and place the grouping control module after the mask branch, wherein the grouping control module performs deep semantic parsing of the input features through channel grouping and weighting control;

[0062] Further details are attached. Figure 5 As shown, step S400 of this application further includes:

[0063] Step S410: The group control module divides the input features into four groups of single features according to the number of channels;

[0064] Step S420: Traverse the four groups of single features and perform the same weighting control to obtain four groups of weighted control features. Then, concatenate the four groups of weighted control features by channel to obtain the final output features.

[0065] Step S410 of this application further includes:

[0066] Step S411: Traverse the single features in the four groups of single features and perform grouped convolution to obtain grouped features;

[0067] Step S412: Input the grouping features into the max pooling layer and average pooling layer in the grouping control module to obtain the max pooling features and average pooling features;

[0068] Step S413: After the maximum pooling feature and the average pooling feature are first compressed to 1 by convolution, a weight layer is generated according to the SoftMax function, and the weight layer is convolutionally expanded to the size of the original grouped features to obtain the maximum control weight and the average control weight.

[0069] Step S414: Multiply the maximum control weight and the average control weight with the single group of features respectively to obtain the maximum control feature and the average control feature, and add the maximum control feature and the average control feature with the single group of features respectively to obtain the weighted control feature.

[0070] The calculation formula of the group control module includes:

[0071]

[0072] Where ConC(·) represents the concatenation operation; M q1 M q2 M q3 and M q4 These are the channel features of the mask features after being evenly divided according to the number of channels; M G IW represents the control feature obtained after the mask feature is input into the group control module; MW represents the pooling control of each group of channel features; AW represents the maximum pooling control operation in pooling control; MP(), SM() and AP() represent maximum pooling, normalized exponential function and average pooling respectively.

[0073] Specifically, the construction process of the group control module is as follows: First, the mask features are divided into four groups of channel features according to the number of channels: M q1 M q2 M q3 M q4 Each group of channel features has 2 channels. Then, the single-channel features are processed through single-channel group convolutions to extract feature information by channel. Next, the single-channel features are batch normalized and activated, and then passed through max pooling and average pooling layers to extract salient semantic features and mean semantic features. Then, the pooling control features extracted by the two pooling layers are compressed from 2 to 1 through 1×1 convolutions, and then processed through batch normalization and activation functions before being input into a normalization exponential function to calculate two control weights. Then, a 1×1 convolution is used to expand the number of channels of the control weights to their original size before compression. Finally, the calculated control weights are multiplied by the original single-channel features to obtain the maximum control feature and the average control feature. These maximum and average control features are then added to the original single-channel features to obtain the single-group control feature. After performing the above operations on all four groups of channel features, they are concatenated by channel, allowing the mask features to complete cross-channel learning in different subspaces, resulting in the final grouped control features.

[0074] The calculation formula for the group control module includes:

[0075]

[0076] Where ConC(·) represents the concatenation operation; M q1 M q2 M q3 and M q4 These are the channel features of the mask features after being evenly divided according to the number of channels; M GIW represents the control feature obtained after the mask feature is input into the group control module; MW represents the pooling control of each group of channel features; AW represents the maximum pooling control operation in pooling control; MP(), SM() and AP() represent maximum pooling, normalized exponential function and average pooling respectively.

[0077] Step S500: Based on the weakly supervised loss module, the multi-scale localization feature pyramid, and the grouping control module, construct a weakly supervised video instance segmentation model for traffic scenes;

[0078] Furthermore, step S500 of this application also includes:

[0079] Step S510: Use the ResNet50 as the backbone network to extract features;

[0080] Step S520: The multi-scale positioning feature pyramid is used to map spatial location information and multi-scale receptive fields to high-level features using the low-level features of the backbone network, thereby supplementing the output features with multi-scale spatial location information and receptive fields.

[0081] Step S530: Use the group control module to perform weighted control on the mask feature part: group the mask features equally according to the number of channels, and use the group control module to enable each group of channel features to complete cross-channel feature learning in its subspace;

[0082] Step S540: Concatenate the mask features and relative position information obtained from the grouping control module;

[0083] Step S550: Input the mask features with relative position information into the mask dynamic convolutional layer to obtain the instance mask.

[0084] Specifically, such as Figure 6 As shown, ResNet50 is used as the backbone network to extract features. A multi-scale localization feature pyramid is used to map spatial location information and multi-scale receptive fields to high-level features using low-level features of the backbone network, supplementing the output features with multi-scale spatial location information and receptive fields. The grouping and adjustment module is used to perform weighting and adjustment operations on the mask feature part. The mask features after grouping and adjustment are concatenated with the relative position matrix by channel, and then input into the masking dynamic convolution to obtain the final instance prediction mask.

[0085] For example, the construction process of the weakly supervised video instance segmentation model for traffic scenes includes, but is not limited to, using ResNet50 as the backbone network. The experimental hardware environment consisted of a GeForce RTX 3060 GAMING XTRIO 12G GPU, an MSI CPU, and an Intel i5 10400F CPU. The software environment included Ubuntu 18.04, cuDNN-8.0.5, CUDA-11.1, Python-3.7, and the deep learning framework PyTorch-1.8. The experiment was built based on the PyTorch deep learning framework. During the training process, images of uniform size 360×640 were used as initial data inputs into the network. The experiment iterated 23,000 times in one training cycle, with a batch size of 4 input videos. The initial learning rate was set to 0.005, and cosine annealing was used to decay the learning rate.

[0086] Step S600: Obtain a traffic scene dataset and extract traffic scene instances based on the traffic scene dataset;

[0087] Step S700: Input the traffic scene instance into the traffic scene weakly supervised video instance segmentation model, and segment the traffic scene instance using the traffic scene weakly supervised video instance segmentation model.

[0088] Specifically, the traffic scene dataset consists of videos containing a large number of instance targets within traffic scenes. For example, in the traffic scene instances, the instances in the videos are divided into seven categories: trucks, motorcycles, trains, skateboards, dogs, people, and cars. The training set contains a total of 328 video segments with 7074 frames, of which 602 are instance targets. The validation set contains 53 video segments with 1095 frames, of which 86 are instance targets. A weakly supervised network is trained using a training set without mask annotations, and the training results are validated using a validation set with mask annotations. This weakly supervised video instance segmentation model for traffic scenes addresses the problems of excessive mask labeling costs during training and the lack of pixel-level mask annotations during weakly supervised network training, which leads to inaccurate instance target location information and low-quality predicted masks when the network generates prediction masks.

[0089] Furthermore, such as Figure 7 As shown, step S200 in this embodiment further includes:

[0090] Step S210: The location feature refinement module includes sub-module one, sub-module two, and sub-module three;

[0091] Step S220: Based on submodule one, submodule two, and submodule three, feature extraction is performed on the low-level features to obtain a first feature map, a second feature map, and a third feature map;

[0092] Step S230: After applying activation functions and batch normalization to the first feature map, the second feature map, and the third feature map, respectively, map them onto the high-level feature map to obtain the output result.

[0093] The formula for obtaining the output result is as follows:

[0094] F i =Conv 3×3 (SIR 6-i (C1)+P i i∈[3,5]

[0095] Where (F3,F4,F5) and (P3,P4,P5) represent feature maps (F3,F4,F5) and (P3,P4,P5) in the multi-scale localization feature pyramid, respectively; Conv 3×3 () represents a convolution operation with a 3×3 kernel; SIR 6-i () indicates the 6th to ith sub-module in the position feature refinement module.

[0096] Specifically, the location feature refinement module includes sub-module one (SIR-1), sub-module two (SIR-2), and sub-module three (SIR-3). Sub-module one (SIR-1), sub-module two (SIR-2), and sub-module three (SIR-3) are used to extract features from the low-level features to obtain a first feature map, a second feature map, and a third feature map. The first feature map, the second feature map, and the third feature map are then subjected to activation functions and batch normalization, and mapped onto the high-level feature map to obtain the output result.

[0097] The formula for obtaining the output result is as follows:

[0098] F i =Conv 3×3 (SIR 6-i (C1)+P i i∈[3,5]

[0099] Where (F3,F4,F5) and (P3,P4,P5) represent feature maps (F3,F4,F5) and (P3,P4,P5) in the multi-scale localization feature pyramid, respectively; Conv 3×3 () represents a convolution operation with a 3×3 kernel; SIR 6-i () indicates the 6th to ith sub-module in the position feature refinement module.

[0100] In summary, the weakly supervised video instance segmentation method for traffic scenes provided in this application has the following technical effects:

[0101] 1. A weakly supervised loss module is constructed based on projection loss and color similarity loss functions; a multi-scale localization feature pyramid is constructed based on a location feature refinement module; a grouping and adjustment module is constructed; based on the weakly supervised loss module, the multi-scale localization feature pyramid, and the grouping and adjustment module, a weakly supervised video instance segmentation model for traffic scenes is constructed; traffic scene instances are extracted from the traffic scene dataset; the traffic scene instances are input into the weakly supervised video instance segmentation model for traffic scene segmentation, and the traffic scene instances are segmented by the weakly supervised video instance segmentation model. This achieves the technical effect of improving the accuracy of segmentation and localization, improving the precision of mask segmentation, and enhancing the overall quality of weakly supervised video instance segmentation for traffic scenes.

[0102] 2. By using W-VIS (weakly supervised baseline network) as a control and setting up W-VIS+MSL-FPN (multi-scale localization feature pyramid) group, W-VIS+GRC (group control module) group, and W-VIS+MSL-FPN+GRC group for experiments, the average segmentation accuracy and segmentation speed of multiple groups of experiments were compared. The results showed that the traffic scene weakly supervised video instance segmentation model (W-VIS+MSL-FPN+GRC group) based on multi-scale localization feature pyramid and group control achieved the best overall results in terms of detection and segmentation accuracy and segmentation speed.

[0103] Example 2

[0104] Using the YouTube-VIS2019 public dataset as the data source, we extracted the required data from the YouTube-VIS2019 public dataset and constructed a traffic scene dataset. We then validated the weakly supervised video instance segmentation method for traffic scenes using multi-scale localization and group control, and tested its segmentation performance. The specific process is as follows:

[0105] The required traffic scene instance targets were extracted from the tag files provided by the YouTube-VIS2019 public dataset, and then divided into training and validation sets with an instance ratio of 9:1.

[0106] The instances in the videos were divided into 7 categories: trucks, motorcycles, trains, skateboards, dogs, people, and cars. The training set contained a total of 328 video clips with 7074 frames, of which 602 were instances. The validation set contained 53 video clips with 1095 frames, of which 86 were instances.

[0107] Table 1 Number of instances in the traffic scenario dataset

[0108]

[0109] Experimental parameters: ResNet50 was used as the backbone network. The hardware environment consisted of a GeForce RTX3060 GAMING X TRIO 12G GPU, an MSI GPU, and an Intel i5 10400F CPU. The software environment included Ubuntu 18.04, cuDNN-8.0.5, CUDA-11.1, Python 3.7, and the deep learning framework PyTorch-1.8. The experiment was built using the PyTorch deep learning framework. During training, images of uniform size (360×640) were used as initial data input to the network. The experiment iterated 23,000 times in one training cycle, with a batch size of 4 input videos. The initial learning rate was set to 0.005, and cosine annealing was used to decay the learning rate.

[0110] Table 2 Model Validation Results

[0111]

[0112] As shown in Table 2, the average segmentation accuracy of the weakly supervised baseline network W-VIS is 32.2%. When the original feature pyramid is replaced with a multi-scale localization feature pyramid (MSL-FPN) on top of W-VIS, the segmentation accuracy is 36.6%, which is a 4.4% improvement compared to W-VIS. When a group control module (GRC) is added after the mask branch of W-VIS, the average segmentation accuracy of the network reaches 36.2%, which is a 4% improvement compared to W-VIS. If GRC and MSL-FPN are combined and added to the W-VIS network, the segmentation accuracy reaches the best of 37.9%, which is a 5.7% improvement compared to W-VIS.

[0113] Example 3

[0114] Based on the same inventive concept as the weakly supervised video instance segmentation method for traffic scenes described in the foregoing embodiments, this invention also provides a weakly supervised video instance segmentation system for traffic scenes. Please refer to the appendix. Figure 8 The system includes:

[0115] The first construction module 11 is used to construct a weakly supervised loss module based on the projection loss function and the color similarity loss function;

[0116] The second construction module 12 is used to construct the position feature refinement module;

[0117] The third construction module 13 is used to construct a multi-scale positioning feature pyramid based on the location feature refinement module.

[0118] The fourth construction module 14 is used to construct a grouping control module and place the grouping control module after the mask branch. The grouping control module performs deep semantic parsing of the input features through channel grouping and weighting control.

[0119] The fifth construction module 15 is used to construct a weakly supervised video instance segmentation model for traffic scenes based on the weakly supervised loss module, the multi-scale localization feature pyramid, and the grouping control module.

[0120] Scene instance extraction module 16, which is used to obtain a traffic scene dataset and extract traffic scene instances based on the traffic scene dataset;

[0121] The segmentation module 17 is used to input the traffic scene instance into the traffic scene weakly supervised video instance segmentation model, and to segment the traffic scene instance through the traffic scene weakly supervised video instance segmentation model.

[0122] The projection loss function includes:

[0123]

[0124]

[0125] L XY =L X +L Y

[0126] Where W and H represent the width and height of the image; x n and y n s represents the projection vectors of the maximum values ​​of the true bounding box on the x and y axes, respectively; n L represents the projection of the maximum probability that a pixel in the corresponding direction is a foreground pixel; X and L Y These represent the projection losses along the x-axis and y-axis, respectively; L XY This represents the final projection loss.

[0127] The color similarity loss function includes:

[0128] P(y e =1)=m c,d ·m e,f +(1-m c,d )·(1-me,f )

[0129]

[0130] Among them, y e This represents an edge e in an undirected graph constructed from image pixels; m c,d Let m represent the probability that one of the endpoints of edge e is a foreground target; e,f Let P(y) represent the probability that the other endpoint of edge e is a foreground target; e =1) represents the probability that edge e is labeled 1, where in an undirected graph, an edge labeled 1 indicates that the two endpoints maintain consistent class, and vice versa; E in The set of edges within the bounding box that contain at least one pixel; N is the set of edges within the bounding box. in The number of sides; L cs This represents the loss due to color similarity.

[0131] Furthermore, the system also includes:

[0132] The first execution module is used for the location feature refinement module, which includes sub-module one, sub-module two, and sub-module three.

[0133] The feature map acquisition module is used to extract features from low-level features based on submodule one, submodule two, and submodule three, respectively, to obtain a first feature map, a second feature map, and a third feature map;

[0134] The mapping module is used to map the first feature map, the second feature map, and the third feature map onto a high-level feature map after applying activation functions and batch normalization, respectively, to obtain the output result.

[0135] The formula for obtaining the output result is as follows:

[0136] F i =Conv 3×3 (SIR 6-i (C1)+P i i∈[3,5]

[0137] Where (F3,F4,F5) and (P3,P4,P5) represent feature maps (F3,F4,F5) and (P3,P4,P5) in the multi-scale localization feature pyramid, respectively; Conv 3×3 (·) represents a convolution operation with a 3×3 kernel; SIR 6-i (·) indicates the 6th to ith sub-module in the position feature refinement module.

[0138] Furthermore, the system also includes:

[0139] A network layer acquisition module is used to set ResNet50 as a backbone network and acquire the C1, C2, C3, C4 and C5 layers in the backbone network.

[0140] A channel number compression module is used to compress the channel number of layers C3, C4, and C5 through convolution;

[0141] A multi-scale information feature acquisition module is used to input the C1 layer into the location feature refinement module to acquire multi-scale information feature one, multi-scale information feature two, and multi-scale information feature three.

[0142] The second execution module is used to set the C5 layer after channel number compression as the P5 layer;

[0143] The third execution module is used to perform bilinear interpolation upsampling on the P5 layer and then fuse it with the C4 layer to obtain the P4 layer.

[0144] The fourth execution module is used to perform bilinear interpolation upsampling on the P4 layer and then fuse it with the C3 layer to obtain the P3 layer;

[0145] The fifth execution module is used to input the multi-scale information feature one, the multi-scale information feature two, and the multi-scale information feature three into the P3 layer, the P4 layer, and the P5 layer to obtain the P'3 layer, the P'4 layer, and the P'5 layer with multi-scale spatial location information.

[0146] The sixth execution module is used to perform convolution on layers P'3, P'4, and P'5 to obtain layers F3, F4, and F5.

[0147] A continuous convolution module is used to perform two consecutive convolutions on the F5 layer to obtain the F6 and F7 layers.

[0148] Furthermore, the system also includes:

[0149] A single-group feature acquisition module is used by the group control module to divide the input features into four groups of single-group features according to the number of channels.

[0150] The weighting and control module is used to traverse the four groups of single features and perform the same weighting and control to obtain four groups of weighted and controlled features. The four groups of weighted and controlled features are then concatenated by channel to obtain the final output feature.

[0151] The weighting and control module includes:

[0152] A grouped convolutional unit is used to traverse the single features in the four groups of single features and perform grouped convolution to obtain grouped features;

[0153] A pooling feature acquisition unit is used to input the grouping features into the max pooling layer and average pooling layer in the grouping control module to obtain max pooling features and average pooling features.

[0154] The weight acquisition unit is used to first compress the number of channels of the maximum pooling feature and the average pooling feature to 1 through convolution, then generate a weight layer according to the SoftMax function, and then expand the weight layer to the size of the original grouped features through convolution to obtain the maximum control weight and the average control weight.

[0155] The weighted regulation feature acquisition unit is used to multiply the maximum regulation weight and the average regulation weight with the single set of features respectively to obtain the maximum regulation feature and the average regulation feature, and then add the maximum regulation feature and the average regulation feature with the single set of features respectively to obtain the weighted regulation feature.

[0156] The calculation formula of the group control module includes:

[0157]

[0158] Where ConC(·) represents the concatenation operation; M q1 M q2 M q3 and M q4 These are the channel features of the mask features after being evenly divided according to the number of channels; M G IW represents the control feature obtained after the mask feature is input into the group control module; MW represents the pooling control performed on the channel features of each group; AW represents the maximum pooling control operation in pooling control; MP(·), SM(·) and AP(·) represent maximum pooling, normalized exponential function and average pooling respectively.

[0159] Example 4

[0160] like Figure 9 As shown, based on the same inventive concept as the traffic scene weak supervision video instance segmentation method in the foregoing embodiments, this application embodiment also provides a traffic scene weak supervision video instance segmentation device, including: a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. When the program is executed by the processor, the system performs the method described in any of the first aspects.

[0161] In this application, the traffic scene weak supervision video instance segmentation device further includes: when a computer program stored in memory 305 and capable of running on processor 302 is executed by processor 302, it implements various processes of the above-described method embodiment for controlling output data.

[0162] Transceiver 303 is used to receive and send data under the control of processor 302.

[0163] In this application, a bus architecture (represented by bus 301) is used. Bus 301 may include any number of interconnected buses and bridges. Bus 301 connects various circuits, including one or more processors represented by processor 3020 and memory represented by memory 305.

[0164] Bus 301 represents one or more of several types of bus architectures, including memory buses and memory controllers, peripheral buses, accelerated graphics ports, processors, or local buses using any bus architecture from various bus architectures. As an example and not a limitation, such architectures include: industry-standard architecture buses, microchannel architecture buses, extended buses, video electronics standards associations, and peripheral interconnect buses.

[0165] Processor 302 can be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processors described above include: general-purpose processors, central processing units, network processors, digital signal processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLs), programmable logic arrays, microcontroller units or other programmable logic devices, discrete gates, transistor logic devices, and discrete hardware components. They can implement or execute the methods, steps, and logic block diagrams disclosed in this application. For example, the processor can be a single-core processor or a multi-core processor, and the processor can be integrated on a single chip or located on multiple different chips.

[0166] Processor 302 can be a microprocessor or any conventional processor. The method steps disclosed in this application can be directly executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in readable storage media known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, registers, etc. The readable storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0167] Bus 301 can also connect various other circuits, such as peripheral devices, voltage regulators, or power management circuits. Bus interface 304 provides an interface between bus 301 and transceiver 303, all of which are well known in the art. Therefore, this application will not describe them further.

[0168] Transceiver 303 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. For example, transceiver 303 receives external data from other devices, and transceiver 303 is used to transmit data processed by processor 302 to other devices. Depending on the nature of the computer device, a user interface 308 may also be provided, such as a touchscreen, physical keyboard, monitor, mouse, speaker, microphone, trackball, joystick, or stylus.

[0169] It should be understood that, in this application, memory 305 may further include memory remotely configured relative to processor 302, and such remotely configured memory can be connected to a server via a network. One or more portions of the aforementioned network may be an ad hoc network, intranet, extranet, virtual private network, local area network, wireless local area network, wide area network, wireless wide area network, metropolitan area network, the Internet, public switched telephone network, conventional telephone network, cellular telephone network, wireless network, wireless fidelity network, and combinations of two or more of the aforementioned networks. For example, cellular telephone networks and wireless networks may be Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Global System for Microwave Access (GSMA), General Packet Radio Service (GPRS), Wideband Code Division Multiple Access (WDMA), Long Term Evolution (LTE), LTE Frequency Division Duplex (FDMA), LTE Time Division Duplex (TDM), Advanced Long Term Evolution (ALE), Universal Mobile Communications (GMMA), Enhanced Mobile Broadband (EMB), Massive Machine-Type Communications (MMTC), Ultra Reliable Low Latency Communications (ULSC).

[0170] It should be understood that the memory 305 in this application may be volatile memory or non-volatile memory, or may include both volatile memory and non-volatile memory. Among them, non-volatile memory includes: read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, or flash memory.

[0171] Volatile memory includes random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SRAM), double data rate synchronous dynamic random access memory (DRAM), enhanced synchronous dynamic random access memory (ERRAM), synchronous linked dynamic random access memory (SRAM), and direct memory bus (DMB) RAM. The memory 305 of the weakly supervised video instance segmentation device for a traffic scene described in this application includes, but is not limited to, the above-described and any other suitable types of memory.

[0172] In this application, memory 305 stores the following elements of operating system 306 and application program 307: executable modules, data structures, or subsets thereof, or extended sets thereof.

[0173] Specifically, the operating system 306 includes various device programs, such as the framework layer, kernel library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 307 includes various applications, such as media players and browsers, used to implement various application functions. Programs implementing the methods of this application can be included in the application program 307. The application program 307 includes applets, objects, components, logic, data structures, and other computer device executable instructions that perform specific tasks or implement specific abstract data types.

[0174] In addition, this application also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the various processes of the above-described method embodiment for controlling output data and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0175] This application provides a method for weakly supervised video instance segmentation in traffic scenes. The method is applied to a weakly supervised video instance segmentation system for traffic scenes. The method includes: constructing a weakly supervised loss module based on a projection loss function and a color similarity loss function; constructing a multi-scale localization feature pyramid based on a location feature refinement module; constructing a grouping and adjustment module; constructing a weakly supervised video instance segmentation model for traffic scenes based on the weakly supervised loss module, the multi-scale localization feature pyramid, and the grouping and adjustment module; extracting traffic scene instances from a traffic scene dataset; inputting the traffic scene instances into the weakly supervised video instance segmentation model for segmentation; and segmenting the traffic scene instances using the weakly supervised video instance segmentation model. This method solves the technical problems of current weakly supervised video instance segmentation networks, which suffer from insufficient pixel-level annotation information and inadequate bounding box information supervision, leading to errors in obtaining target location information and low mask segmentation quality during network segmentation. It achieves the technical effect of improving segmentation and localization accuracy, mask segmentation precision, and overall quality of weakly supervised video instance segmentation in traffic scenes by using the weakly supervised video instance segmentation model.

[0176] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for weakly supervised video instance segmentation in traffic scenes, characterized in that, The method includes: A weakly supervised loss module is constructed based on the projection loss function and the color similarity loss function; Construct a location feature refinement module; Based on the location feature refinement module, a multi-scale positioning feature pyramid is constructed. A grouping control module is constructed and placed after the mask branch. The grouping control module performs deep semantic parsing of the input features through channel grouping and weighting control. Based on the weakly supervised loss module, the multi-scale localization feature pyramid, and the grouping control module, a weakly supervised video instance segmentation model for traffic scenes is constructed. Obtain a traffic scene dataset and extract traffic scene instances based on the traffic scene dataset; The traffic scene instance is input into the traffic scene weakly supervised video instance segmentation model, and the traffic scene instance is segmented by the traffic scene weakly supervised video instance segmentation model; The group control module includes: The group control module divides the input features into four groups of single features according to the number of channels; The same weighting and control are performed on the four groups of single features to obtain four groups of weighted and controlled features. The four groups of weighted and controlled features are then concatenated by channel to obtain the final output feature. The process of performing the same weighting adjustment on the four groups of single features includes: The grouped features are obtained by iterating through the individual features in the four groups of individual features and performing grouped convolution. The grouping features are input into the maximum pooling layer and the average pooling layer within the grouping control module to obtain the maximum pooling features and the average pooling features. The maximum pooling feature and the average pooling feature are first compressed to 1 channel number through convolution, and then a weight layer is generated according to the SoftMax function. The weight layer is then convolved and expanded to the size of the original grouped features to obtain the maximum control weight and the average control weight. The maximum control weight and the average control weight are multiplied by the single set of features to obtain the maximum control feature and the average control feature, respectively. The maximum control feature and the average control feature are then added to the single set of features to obtain the weighted control feature.

2. The method as described in claim 1, characterized in that, The projection loss function includes: ; ; ; Where W and H represent the width and height of the image; and These represent the projection vectors of the maximum values ​​of the true bounding box on the x-axis and y-axis, respectively; This represents the maximum projection of the probability that a pixel in the corresponding direction is a foreground pixel; and These represent the projection losses along the x-axis and y-axis, respectively. This represents the final projection loss.

3. The method as described in claim 1, characterized in that, The color similarity loss function includes: ; ; in, This represents an edge e in an undirected graph constructed from image pixels; Let e ​​represent the probability that one of the endpoints of edge e is a foreground target; Let e ​​represent the probability that the other endpoint of edge e is a foreground target; This represents the probability that edge e is labeled 1. In an undirected graph, an edge labeled 1 indicates that the two endpoints maintain the same class, while the other label is 0. The set of edges in the bounding box that contain at least one pixel; Within the box The number of sides; This represents the loss due to color similarity.

4. The method as described in claim 1, characterized in that, Construct a location feature refinement module, including: The location feature refinement module includes sub-module one, sub-module two, and sub-module three; Based on submodule one, submodule two, and submodule three, feature extraction is performed on the low-level features to obtain a first feature map, a second feature map, and a third feature map; The first feature map, the second feature map, and the third feature map are activated and batch normalized respectively, and then mapped onto the high-level feature map to obtain the output result.

5. The method as described in claim 1, characterized in that, Based on the location feature refinement module, a multi-scale positioning feature pyramid is constructed, including: Set ResNet50 as the backbone network and obtain the C1, C2, C3, C4 and C5 layers in the backbone network; The number of channels in layers C3, C4, and C5 is compressed using convolution; The C1 layer is input into the location feature refinement module to obtain multi-scale information feature one, multi-scale information feature two, and multi-scale information feature three. Set the C5 layer with compressed channel count as the P5 layer; The P5 layer is upsampled by bilinear interpolation and then fused with the C4 layer to obtain the P4 layer. The P4 layer is upsampled by bilinear interpolation and then fused with the C3 layer to obtain the P3 layer. The first multi-scale information feature, the second multi-scale information feature, and the third multi-scale information feature are input into the P3 layer, the P4 layer, and the P5 layer to obtain the P'3 layer, the P'4 layer, and the P'5 layer with multi-scale spatial location information. Convolutions are performed on layers P'3, P'4, and P'5 to obtain layers F3, F4, and F5; The F5 layer is convolved twice consecutively to obtain the F6 and F7 layers.

6. The method as described in claim 4, characterized in that, The formula for obtaining the output result is as follows: ; in,( , , )and( , , ) represent feature maps (F3, F4, F5) and (P3, P4, P5) in the multi-scale localization feature pyramid, respectively; This represents a convolution operation with a 3×3 kernel; This indicates the (6-i)th submodule in the position feature refinement module.

7. The method as described in claim 1, characterized in that, The calculation formula of the group control module includes: ; in, This represents a splicing operation; , , and These are the channel features of each group after the mask features are evenly divided according to the number of channels; This represents the control feature obtained after the mask feature is input into the group control module; This indicates pooling control applied to the characteristics of each group of channels; It is the maximum pooling control operation in pooling control; This indicates the average pooling weighting operation in pooling control; , and These represent max pooling, normalized exponential function, and average pooling, respectively.

8. A weakly supervised video instance segmentation system for traffic scenes, characterized in that, The system is used to perform the method according to any one of claims 1 to 7, the system comprising: The first construction module is used to construct a weakly supervised loss module based on the projection loss function and the color similarity loss function; The second construction module is used to construct the location feature refinement module; The third construction module is used to construct a multi-scale positioning feature pyramid based on the location feature refinement module. The fourth construction module is used to construct the grouping control module and place the grouping control module after the mask branch. The grouping control module performs deep semantic parsing of the input features through channel grouping and weighting control. The fifth construction module is used to construct a weakly supervised video instance segmentation model for traffic scenes based on the weakly supervised loss module, the multi-scale localization feature pyramid, and the grouping control module. A scene instance extraction module is used to obtain a traffic scene dataset and extract traffic scene instances based on the traffic scene dataset. The segmentation module is used to input the traffic scene instance into the traffic scene weakly supervised video instance segmentation model, and to segment the traffic scene instance through the traffic scene weakly supervised video instance segmentation model.

9. A video instance segmentation device for weakly supervised traffic scenes, characterized in that, The device includes a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method according to any one of claims 1 to 7.