A Polarization-Based Self-Attention Modulated Contextual Video Instance Multi-Scale Segmentation Method
By employing polarization self-attention modulation in a contextual mobile robot visual segmentation system, single-level and cascaded modulation models are constructed, addressing the problem of insufficient attention to high-level fine-grained features of targets in existing technologies. This improves segmentation accuracy and mask quality, and enhances the ability to locate target instances.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- KUNMING UNIV OF SCI & TECH
- Filing Date
- 2022-01-21
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies in context-based mobile robot vision segmentation systems suffer from low segmentation accuracy due to insufficient attention to high-level fine-grained features of the target and inaccurate localization of low-level spatial information.
By employing a polarization self-attention modulation method, and constructing single-level and cascaded modulation models, we enhance the adaptability to varying target shapes and the ability to focus on important regional features, thereby improving the model's segmentation performance and mask quality, and perfecting the model's ability to locate target instances.
By adjusting the different embedding positions and the number of dissimilar embeddings of the polarization self-attention mechanism, single-level and cascaded control models were designed, which enhanced the adaptability to the varied shapes of the target and the ability to focus on important regional features, improved the segmentation performance and mask quality of the model, and improved the model's ability to locate target instances.
Smart Images

Figure CN115049693B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video segmentation technology, and in particular to a polarization-self-attention modulated scenario-based multi-scale video instance segmentation method. Background Technology
[0002] Visual segmentation systems for context-sensitive mobile robots are a significant challenge in the development of intelligent robots. When a robot moves within a specific scenario, its visual imaging system is affected by its own speed, shooting angle, and distance from the target. Furthermore, the randomness of topological deformation and scale scaling of the captured target object during movement leads to blurred edges and inconsistent size and shape of the same target in subsequent imaging, greatly interfering with subsequent segmentation. Therefore, accurately and stably segmenting individual targets in video sequences under uncertain conditions is a pressing problem for context-sensitive robot vision systems.
[0003] However, when intelligent robots move autonomously in specific scenarios, the moving targets they capture generally exhibit randomness in terms of topological deformation and scale scaling. Existing technologies suffer from low segmentation accuracy due to insufficient effective attention to high-level fine-grained features of the targets and the inability to accurately locate low-level spatial information. Summary of the Invention
[0004] This application provides a contextual video instance multi-scale segmentation method with polarization self-attention modulation, which solves the technical problem of low segmentation accuracy caused by the low effective attention to high-level fine-grained features of the target and the inability to accurately locate low-level spatial information in existing technologies. By adjusting different embedding positions and dissimilar embedding numbers of the polarization self-attention mechanism, single-level and cascaded modulation models are designed respectively, which enhance the adaptability to the changing shape of the target and the ability to focus on important regional features, improve the segmentation performance and mask quality of the model, and improve the model's ability to locate target instances.
[0005] In view of the above problems, the present invention provides a contextual video instance multi-scale segmentation method with polarization self-attention modulation.
[0006] In a first aspect, this application provides a context-based video instance multi-scale segmentation method with polarization self-attention modulation. The method includes: obtaining a first video instance segmentation dataset; extracting label files from the first video instance segmentation dataset to construct a second video instance segmentation dataset suitable for contextual scenarios; dividing the second video instance segmentation dataset into a training set and a test set according to a predetermined ratio; configuring experimental parameters and building a model environment based on the training set and the test set; modulating a ResNet50-PSA modulation model embedding single-level and cascaded polarization self-attention mechanisms in a residual network based on the experimental parameters and the model environment; constructing a multi-scale spatial localization branch model that aggregates multi-granularity spatial information of target instances; and constructing a context-based video instance segmentation model based on the ResNet50-PSA modulation model and the multi-scale spatial localization branch model.
[0007] On the other hand, this application also provides a context-based video instance multi-scale segmentation system with polarization self-attention modulation. The system includes: a first acquisition unit for acquiring a first video instance segmentation dataset; a first construction unit for extracting label files from the first video instance segmentation dataset to construct a second video instance segmentation dataset suitable for contextual scenarios; a first partitioning unit for partitioning the second video instance segmentation dataset into a training set and a test set according to a predetermined ratio; a first processing unit for configuring experimental parameters and building a model environment based on the training set and the test set; a first modulation unit for modulating a ResNet50-PSA modulation model embedding single-level and cascaded polarization self-attention mechanisms in a residual network based on the experimental parameters and the model environment; a second construction unit for constructing a multi-scale spatial localization branch model that aggregates multi-granularity spatial information of target instances; and a third construction unit for constructing a contextual video instance segmentation model based on the ResNet50-PSA modulation model and the multi-scale spatial localization branch model.
[0008] Thirdly, this application provides an electronic device including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. The transceiver, the memory, and the processor are connected via the bus, and the computer program, when executed by the processor, implements the steps of any of the methods described above.
[0009] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.
[0010] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0011] This technical solution involves extracting label files from the first video instance segmentation dataset to construct a second video instance segmentation dataset suitable for contextual scenarios. This second dataset is then divided into training and testing sets according to a predetermined ratio to configure experimental parameters and build the model environment. A ResNet50-PSA control model with single-level and cascaded polarization self-attention mechanisms is modulated within the residual network. A multi-scale spatial localization branch model that aggregates multi-granularity spatial information of target instances is then constructed. Finally, a contextual video instance segmentation model is built based on the ResNet50-PSA control model and the multi-scale spatial localization branch model. This achieves the technical effect of enhancing adaptability to varying target shapes and focusing on important regional features by designing single-level and cascaded control models through adjusting different embedding positions and dissimilar embedding numbers of the polarization self-attention mechanism. This improves the model's segmentation performance and mask quality, and enhances its ability to locate target instances.
[0012] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating a scenario-based video instance multi-scale segmentation method with polarization self-attention modulation according to this application.
[0014] Figure 2 This is a schematic diagram of a contextual video instance multi-scale segmentation model with polarization self-attention modulation, according to an embodiment of this application.
[0015] Figure 3 This is a schematic diagram of the ResNet50-PSA regulation model structure in an embodiment of this application;
[0016] Figure 4 This is a schematic diagram of the spatial positioning branch model in an embodiment of this application;
[0017] Figure 5 This is a schematic diagram of the structure of a polarization self-attention modulated contextual video instance multi-scale segmentation system according to this application;
[0018] Figure 6 This is a schematic diagram of the structure of an exemplary electronic device of this application.
[0019] Explanation of reference numerals in the attached drawings: First acquisition unit 11, first construction unit 12, first partitioning unit 13, first processing unit 14, first control unit 15, second construction unit 16, third construction unit 17, bus 1110, processor 1120, transceiver 1130, bus interface 1140, memory 1150, operating system 1151, application program 1152, and user interface 1160. Detailed Implementation
[0020] As will be apparent to those skilled in the art from the description of this application, this application can be implemented as a method, apparatus, electronic device, and computer-readable storage medium. Therefore, this application can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. Furthermore, in some embodiments, this application can also be implemented as a computer program product contained in one or more computer-readable storage media, which includes computer program code.
[0021] The aforementioned computer-readable storage medium may be any combination of one or more computer-readable storage media. Computer-readable storage media include: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media include: portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, optical disc read-only memory, optical storage devices, magnetic storage devices, or any combination thereof. In this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0022] The acquisition, storage, use, and processing of data in this application all comply with relevant national laws and regulations.
[0023] This application describes the provided methods, apparatus, and electronic devices using flowcharts and / or block diagrams.
[0024] It should be understood that each block of a flowchart and / or block diagram, as well as combinations of blocks in a flowchart and / or block diagram, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine that, when executed by a computer or other programmable data processing apparatus, creates means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0025] These computer-readable program instructions may also be stored in a computer-readable storage medium that enables a computer or other programmable data processing device to function in a particular manner. In this way, the instructions stored in the computer-readable storage medium produce an instruction apparatus product that includes the functions / operations specified in the blocks of a flowchart and / or block diagram.
[0026] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer-implemented process, such that the instructions that execute on the computer or other programmable data processing apparatus provide a process for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0027] This application will now be described with reference to the accompanying drawings.
[0028] Example 1
[0029] like Figure 1 As shown, this application provides a polarization self-attention modulated contextual video instance multi-scale segmentation method, the method comprising:
[0030] Step S100: Obtain the first video instance segmentation dataset;
[0031] Specifically, the first video instance segmentation dataset is the original dataset obtained through the visual segmentation system of a contextual mobile robot. Video instance segmentation refers to the task of detecting, segmenting, and tracking objects of interest in a video. Video instance segmentation not only involves detecting and segmenting objects in a single frame, but also finding the correspondence between each object across multiple frames, i.e., associating and tracking them.
[0032] Step S200: Extract the label files from the first video instance segmentation dataset to construct a second video instance segmentation dataset suitable for contextual scenarios;
[0033] Specifically, instance segmentation provides different labels for individual instances of objects belonging to the same class. Therefore, instance segmentation can be defined as a technique that simultaneously solves the problems of object detection and semantic segmentation. A second video instance segmentation dataset suitable for the contextualized scenario is extracted using the label files in the first video instance segmentation dataset. For example, the second video instance segmentation dataset suitable for the contextualized scenario described in this application includes 1005 instances and 657 videos covering 12 categories, including pandas, apes, monkeys, bears, giraffes, leopards, foxes, deer, zebras, tigers, elephants, and sea lions.
[0034] Step S300: Divide the second video instance segmentation dataset into a training set and a test set according to a predetermined ratio;
[0035] Specifically, the second video instance segmentation dataset is divided into a training set and a test set according to a predetermined ratio. For example, the ratio of the number of videos to the number of instances in the training set and the test set of the second video instance segmentation dataset can be 5:1, where the training set contains 549 videos and 829 instances, and the test set contains 108 videos and 188 instances. The details are shown in Table 1: Instance Distribution of the Second Video Instance Segmentation Dataset.
[0036] Table 1
[0037]
[0038] Step S400: Configure experimental parameters and build the model environment based on the training set and test set;
[0039] Furthermore, step S400 of this application also includes:
[0040] Step S410: Obtain experimental configuration information;
[0041] Step S420: Analyze the foreground target instances in the second video instance segmentation dataset to obtain the motion characteristics of the foreground target instances;
[0042] Step S430: Based on the experimental configuration information and the motion characteristics, build the experimental environment and configure the experimental parameters for the basic model.
[0043] Specifically, based on the training and testing sets, a network training strategy suitable for the second video instance segmentation dataset is determined. The experimental configuration information includes computer configuration information, such as computing performance. Motion characteristic analysis is performed on the foreground object instances in the second video instance segmentation dataset. Since the video itself is sequence-level data, motion foreground object analysis aims to extract changing regions from the background image in the sequence image. This is crucial for post-processing such as object tracking, object classification, and behavior understanding. The analysis mainly includes background modeling, frame difference, and optical flow methods to obtain the motion characteristics of the foreground object instances in the video sequence. Based on the computer configuration and the motion characteristics of the foreground object instances in the video sequence, experiments are conducted to build and configure the parameters of the basic model, determining a network training strategy suitable for the second video instance segmentation dataset.
[0044] For example, the input image scale is uniformly set to 550×550. During the training phase, the number of iterations is set to 100,000 steps, with an initial learning rate of 5e-4. From steps 0 to 3400, the learning rate is gradually increased from 0 to the pre-set size using a warmup learning rate method. After steps 3400, cosine annealing (Cosine Annealing LR) is used to decay the learning rate, improving the model's loss convergence speed. The batch size is set to 8, and model weights are saved every 5000 steps, for a total of 20 models. The optimal model is selected by comprehensively comparing the accuracy and inference speed of the last four trained models. The specific hyperparameter settings are shown in Table 2: Experimental Parameters.
[0045] Table 2
[0046]
[0047] During training, model weights are saved every 5000 steps, for a total of 20 models. The accuracy and inference speed of the last four trained models are compared to select the optimal model. The optimal model is then used to test the test set to evaluate the final performance of the model.
[0048] Step S500: Based on the experimental parameters and the model environment, modulate the ResNet50-PSA control model with embedded single-stage and cascaded polarization self-attention mechanisms in the residual network;
[0049] Specifically, based on the experimental parameters and the model environment, a ResNet50-PSA control model with embedded single-stage and cascaded polarized self-attention mechanisms was modulated within the residual network. ResNet is an abbreviation for residual network, and ResNet50 is its typical network, containing 49 convolutional layers and one fully connected layer. Polarized self-attention (PSA) is used to solve pixel-level regression tasks and can maintain good performance on fine-grained pixel-level tasks. By employing a nonlinear function of fine-grained regression output distribution and polarized filtering, it can reduce the information loss caused by dimensionality reduction.
[0050] Step S600: Construct a multi-scale spatial localization branch model that aggregates multi-granularity spatial information of target instances;
[0051] Furthermore, step S600 of this application also includes:
[0052] Step S610: Construct a feature pyramid based on the ResNet50-PSA control model;
[0053] Step S620: Based on the current output layer and neighboring low-level feature maps of the feature pyramid horizontal mapping, perform multi-scale feature interaction with the multi-scale spatial localization branch model in a bidirectional feature integration manner.
[0054] Specifically, a multi-scale spatial localization branch model is constructed that aggregates multi-granularity spatial information of target instances. The multi-scale space is obtained by convolving the original image with a two-dimensional Gaussian function. By continuously changing the parameters, continuously varying images are obtained. Compared with the original image, the information in these images gradually decreases, and detailed information is gradually smoothed out, but the number of pixels remains unchanged, that is, the resolution remains unchanged. The multi-scale space has the same resolution at different scales, which can be understood as the image size being the same at different scales.
[0055] like Figure 2 As shown, a feature pyramid is constructed based on the ResNet50-PSA control model. The feature pyramid is a feature extractor designed based on the concept of a feature pyramid, aiming to improve accuracy and speed. It consists of two parts: a bottom-up approach and a top-down approach. The bottom-up approach uses traditional convolutional networks for feature extraction. As convolution deepens, spatial resolution decreases and spatial information is lost, but more high-level semantic information is detected. Images contain targets of different sizes, and different targets have different features. The feature pyramid utilizes shallow features to distinguish simple targets and deep features to distinguish complex targets.
[0056] By utilizing the current output layer, which is horizontally mapped from the feature pyramid, and its neighboring lower-level feature maps, a bottom-up bidirectional feature integration approach is used, along with the multi-scale spatial localization branch model, to perform multi-scale feature interaction. This enriches the feature texture information and deep semantic information of the current feature map, improving the model's ability to locate target instances. For example, as... Figure 4 As shown, the specific steps include:
[0057] Step 1: The lowest feature map M3 is obtained by lateral mapping of P3 through a 3×3 convolution with a stride of 1;
[0058] Step 2: Downsample M3 by performing a 3×3 convolution with a stride of 2 until it is the same size as the P4 feature map;
[0059] Step 3: Add the result of Step 2 to M3 element by element;
[0060] Step 4: The result of Step 3 is subjected to a feature integration operation through a 3×3 convolution with a stride of 2 to obtain M4, which overcomes the background information interference problem caused by the feature fusion method of direct addition.
[0061] Step 5: Downsample M4 by performing a 3×3 convolution with a stride of 2 until it is the same size as the P5 feature map;
[0062] Step 6: Add the result of Step 2 to M4 element by element;
[0063] Step 7: The result of Step 6 is processed by a 3×3 convolution with a stride of 2 to perform a feature integration operation to obtain M5. The feature integration of its own multi-granularity information is used to generate a high-quality feature map with rich content by utilizing features from neighboring layers and a constant spatial resolution.
[0064] To address the issues of target instance scaling and blurred edge texture information, a multi-granularity spatial localization branch model is constructed and interacts with a feature pyramid network at multiple scales. This enriches the edge texture information of the target instance under the high-level feature map and improves the model's ability to locate the target instance.
[0065] Step S700: Based on the ResNet50-PSA control model and the multi-scale spatial localization branch model, construct a contextual video instance segmentation model.
[0066] Specifically, by comparing the applicability of single-level and cascaded control models to the second video instance segmentation dataset, the best ResNet50-PSA control model is selected by embedding a polarization self-attention mechanism after the fourth residual block of the residual network. Based on this, a contextual video instance segmentation model is established by combining a spatial localization branch. The contextual video instance segmentation model is used to improve the segmentation quality and accuracy of video instances.
[0067] The contextual video instance segmentation model is mainly divided into two stages, such as... Figure 2 As shown, the feature extraction stage includes the ResNet-PSA control model, the feature pyramid network, and the spatial localization branch; the feature post-processing stage includes the prediction head branch, the prototype mask branch, etc. The image mask is mainly used for extracting regions of interest, masking, and extracting structural features.
[0068] In the feature extraction stage, video frames are first input into the ResNet50-PSA control model to extract semantic features. By establishing nonlinear deep feature channel dependencies and spatial location associations, the model focuses on important regions and target features. The backbone network of this control model mainly consists of four residual blocks. A top-down feature pyramid sampling method is used to introduce richer deep semantic information into feature maps of different scales. Finally, a multi-scale spatial localization branch model is constructed that aggregates multi-granularity spatial information of target instances. Through a bottom-up bidirectional feature integration method, high-resolution, relatively weak semantic information but strong spatial location information low-level feature maps are gradually fused with low-resolution but semantically strong high-level feature maps, enriching the deep abstract features and spatial localization information of the multi-scale feature maps.
[0069] In the feature post-processing stage, the multi-scale feature map of the spatial localization branch is used as input to the prediction head model, and the output is obtained after passing through three parallel branches. Specifically, the classification branch obtains the target classification confidence of the current image; the anchor box branch generates the target detection box position offset of the current image; and the mask branch predicts a mask coefficient for each anchor box. Simultaneously, the bottom-level features of the feature pyramid are used as input to the prototype mask branch, which runs parallel to the prediction head model. This prototype mask tensor, consisting of a five-layer fully convolutional network, generates a 32-channel-dimensional prototype mask tensor, corresponding to the number of mask coefficients. Finally, fast non-maximum suppression is applied to the anchor box network of the prediction head model to filter overlapping candidate boxes at the same target location, and the prototype mask is linearly combined with its corresponding template coefficients to generate an instance mask.
[0070] By adjusting different embedding positions and dissimilar embedding numbers in the polarization self-attention mechanism, single-level and cascaded control models were designed to enhance adaptability to varying target shapes and the ability to focus on important regional features, thereby improving the model's segmentation performance and mask quality. To address the issues of target instance scale scaling and relatively blurred edge texture information, a multi-granularity spatial localization branch model was constructed and interacted with a feature pyramid network to enrich the edge texture information of target instances under high-level feature maps, thus improving the model's ability to locate target instances.
[0071] Furthermore, step S500 of this application also includes:
[0072] Step S510: Calculate the output of each residual block in the original residual network;
[0073] Step S520: Calculate the residual block output with embedded polarization self-attention mechanism;
[0074] Step S530: Design a single-stage control model by combining the residual network and different numbers of polarization self-attention mechanisms;
[0075] Step S540: Design a cascaded control model by combining the residual network and the polarization self-attention mechanism at different positions.
[0076] Furthermore, the formula for calculating the output of each residual block in the original residual network is as follows:
[0077]
[0078] In the formula, I(C in H in W in Given any residual block input, C in H in W inThese represent the number of channels and the height and width of the feature map, respectively. Conv(·) is the convolution operation, n is the sequence number of the residual block, and m is the number of output channels of the corresponding convolution in the current nth residual block.
[0079] Furthermore, the calculation formula for the residual block output of the embedded polarization self-attention mechanism is as follows:
[0080]
[0081] In the formula, and These are channel self-attention and spatial self-attention, respectively, within the polarization self-attention mechanism.
[0082] Specifically, the steps for regulating the ResNet50-PSA control model with embedded single-stage and cascaded polarization self-attention mechanisms in the residual network are as follows:
[0083] Step 1: Calculate the output of each residual block in the original residual network;
[0084] Furthermore, the calculation formula for the output of each residual block in the original residual network is as follows:
[0085]
[0086] In the formula, I(C in H in W in Given any residual block input, C in H in W in These represent the number of channels and the height and width of the feature map, respectively. Conv(·) is the convolution operation, n is the sequence number of the residual block, and m is the number of output channels of the corresponding convolution in the current nth residual block.
[0087] Step 2: Calculate the residual block output of the embedded polarization self-attention mechanism;
[0088] Furthermore, the formula for calculating the residual block output of the embedded polarization self-attention mechanism is as follows:
[0089]
[0090] In the formula, and These are channel self-attention and spatial self-attention, respectively, within the polarization self-attention mechanism.
[0091] Step 3: Design a single-stage control model by combining residual networks and different numbers of polarization self-attention mechanisms;
[0092] Single-stage control models designed by combining residual networks and varying numbers of polarization self-attention mechanisms include: exemplary models, such as... Figure 3 As shown, the single-stage control model takes into account the influence of the polarization self-attention mechanism embedded in the position of each residual block. The residual blocks with embedded polarization self-attention mechanism are regarded as independent individuals. The PSA switch is controlled after residual block 1, residual block 2, residual block 3, and residual block 4 respectively, thereby controlling the embedding position of the polarization attention mechanism. The specific working situation is shown in Table 3: Single-stage control model. ON and OFF represent the switch state being open and closed, respectively. In the other switches of each embedding position, the original switch state is closed and the PSA switch state is open.
[0093] Table 3
[0094]
[0095] Step 4: Design a cascaded control model by combining residual networks and polarization self-attention mechanisms at dissimilar locations.
[0096] The cascaded control model, which combines residual networks and polarization self-attention mechanisms at dissimilar positions, is designed as follows: Considering the impact of the control interaction process between different individuals on the final segmentation effect, the cascaded control model sequentially accumulates the embedding models at dissimilar positions and controls the corresponding PSA switches, as shown in Table 4: Cascaded Control Model. Similar to the single-stage model, in each embedding position, the original switch state is closed, and the PSA switch state is open for the remaining switches (unspecified).
[0097] Table 4
[0098]
[0099] The experimental results of the single-level control model, which embeds polarization self-attention mechanisms after residual block 1, residual block 2, residual block 3, and residual block 4 respectively, are shown in Table 5: Comparison of experimental results of single-level control models.
[0100] Table 5
[0101]
[0102] The experimental results of the cascaded control model with polarization self-attention mechanism embedded after residual blocks 1 and 2, residual blocks 1 and 3, residual blocks 1 and 4, residual blocks 2 and 3, residual blocks 2 and 4, residual blocks 3 and 4, and all residual blocks are shown in Table 6: Comparison of experimental results of cascaded control models.
[0103] Table 6
[0104]
[0105] As shown in Table 5, the detection and segmentation accuracy of the single-level control model shows a progressively increasing trend as the self-attention mechanism is embedded deeper. This is because the model undergoes a semantic morphological process from simple to abstract during feature extraction. Residual block 1, as the first residual block in the backbone network, lacks deep abstract semantic features, requiring a cognitive process to identify effective information, thus resulting in relatively low accuracy improvement. In contrast, residual block 4, as the tail residual block of the backbone network, generates deep abstract semantics, which is precisely the feature structure required by the attention mechanism. Table 7 compares the experimental results of the YolactEdge benchmark and the single-level control model. It can be seen that compared to the benchmark model, the optimal single-level control model achieves a maximum segmentation accuracy of 38.36%, an improvement of 6.8%, and a detection accuracy of 37.99%. However, the speed decreases from 90 frames per second to 80 frames per second. This is because the introduction of the self-attention mechanism inevitably increases the number of model parameters, thus affecting the testing speed. However, the model still generally meets the standard for real-time segmentation.
[0106] Table 7
[0107]
[0108] Combining Tables 5 and 6, it is found that the cascaded control model generally does not perform as well as the single-stage control model in terms of segmentation accuracy. While embedding a single self-attention mechanism module improves the model's accuracy to some extent, combining it with other self-attention locations leads to a decrease in the current model's segmentation accuracy. In Table 5, the single-stage control model embedded after residual block 1 and residual block 2 achieves segmentation accuracies of 34.52% and 35.16%, respectively, while in Table 6, the cascaded control model, with both controls combined, experiences a decrease in segmentation accuracy to 32.39%. However, its overall segmentation accuracy is generally improved compared to the YolactEdge baseline model.
[0109] In summary, the single-stage control model embedding the polarization self-attention mechanism after residual block 4 achieves optimal segmentation accuracy. Both the single-stage and cascaded control models improve segmentation accuracy for the current task, demonstrating the effectiveness of our proposed method. However, the embedding of the polarization self-attention mechanism introduces a certain loss in detection accuracy. This is because the self-attention mechanism tends to capture feature correlations within the same semantic meaning by modeling the relationships between spatial pixel positions or feature channels, focusing more on pixel regression. This may lead to the misclassification of semantic features from other instances onto the current instance, causing incorrect guidance for the current detection task.
[0110] Furthermore, the steps in this application also include:
[0111] Step S810: Use the training set to train the model and apply the test set to evaluate the performance of the contextual video instance segmentation model.
[0112] Specifically, the training set is used to train the model, and the test set is applied to evaluate the performance of the contextual video instance segmentation model. A qualitative analysis comparison of the contextual video instance segmentation model before and after the improvement is shown in Table 8: Experimental results of the spatial localization branch. It can be seen that the introduction of the spatial localization branch improves the average detection and segmentation accuracy of the model by 6.34% and 3.39%, respectively.
[0113] Table 8
[0114]
[0115] In summary, the polarization self-attention modulated contextual video instance multi-scale segmentation method provided in this application has the following technical effects:
[0116] (1) By extracting the label files from the first video instance segmentation dataset, a second video instance segmentation dataset suitable for contextual scenarios is constructed. The second video instance segmentation dataset is divided into training and testing sets according to a predetermined ratio to configure experimental parameters and build the model environment. The ResNet50-PSA control model with embedded single-level and cascaded polarization self-attention mechanisms is controlled in the residual network. Then, a multi-scale spatial localization branch model that aggregates multi-granularity spatial information of target instances is constructed. Finally, a contextual video instance segmentation model is constructed based on the ResNet50-PSA control model and the multi-scale spatial localization branch model. This achieves the technical effect of designing single-level and cascaded control models by controlling different embedding positions and dissimilar embedding numbers of polarization self-attention mechanisms, enhancing adaptability to varying target shapes and focusing ability on important regional features, improving the segmentation performance and mask quality of the model, and perfecting the model's ability to locate target instances.
[0117] (2) In view of the problem of target instance scaling and relatively vague edge texture information, this invention constructs a multi-granularity spatial localization branch model and a feature pyramid network to carry out multi-scale feature interaction, which enriches the edge texture information of the target instance under the high-level feature map and improves the model's ability to locate the target instance.
[0118] Example 2
[0119] Based on the same inventive concept as the polarization self-attention modulated contextual video instance multi-scale segmentation method in the foregoing embodiments, this invention also provides a polarization self-attention modulated contextual video instance multi-scale segmentation system, such as... Figure 5 As shown, the system includes:
[0120] The first obtaining unit 11 is used to obtain a first video instance segmentation dataset;
[0121] The first construction unit 12 is used to extract the label files in the first video instance segmentation dataset and construct a second video instance segmentation dataset suitable for contextual scenarios.
[0122] The first partitioning unit 13 is used to divide the second video instance segmentation dataset into a training set and a test set according to a predetermined ratio.
[0123] The first processing unit 14 is used to configure experimental parameters and build the model environment according to the training set and the test set.
[0124] The first control unit 15 is used to control the ResNet50-PSA control model embedded with single-stage and cascaded polarization self-attention mechanisms in the residual network based on the experimental parameters and the model environment.
[0125] The second construction unit 16 is used to construct a multi-scale spatial localization branch model that aggregates multi-granularity spatial information of target instances;
[0126] The third construction unit 17 is used to construct a contextual video instance segmentation model based on the ResNet50-PSA control model and the multi-scale spatial localization branch model.
[0127] Furthermore, the system also includes:
[0128] The second obtaining unit is used to obtain experimental configuration information;
[0129] The third obtaining unit is used to analyze the foreground target instances of the second video instance segmentation dataset to obtain the motion characteristics of the foreground target instances;
[0130] The second processing unit is used to build the experimental environment and configure the experimental parameters for the basic model based on the experimental configuration information and the motion characteristics.
[0131] Furthermore, the system also includes:
[0132] The first computing unit is used to calculate the output of each residual block in the original residual network;
[0133] The second computing unit is used to compute the residual block output embedded with the polarization self-attention mechanism;
[0134] The first design unit is used to design a single-stage control model by combining the residual network and different numbers of polarization self-attention mechanisms.
[0135] The second design unit is used to design a cascaded control model by combining the residual network and the polarization self-attention mechanism at dissimilar locations.
[0136] Furthermore, the system also includes:
[0137] The fourth construction unit is used to construct a feature pyramid based on the ResNet50-PSA control model;
[0138] The first interaction unit is used to perform multi-scale feature interaction with the multi-scale spatial localization branch model based on the current output layer and the neighboring low-level feature maps of the feature pyramid horizontal mapping in a bidirectional feature integration manner.
[0139] Furthermore, the system also includes:
[0140] The third processing unit is used to use the training set to train the model and apply the test set to evaluate the performance of the contextual video instance segmentation model.
[0141] The foregoing Figure 1 The various variations and specific examples of the polarization self-attention modulated contextual video instance multi-scale segmentation method in Embodiment 1 are also applicable to the polarization self-attention modulated contextual video instance multi-scale segmentation system of this embodiment. Through the foregoing detailed description of the polarization self-attention modulated contextual video instance multi-scale segmentation method, those skilled in the art can clearly understand the implementation method of the polarization self-attention modulated contextual video instance multi-scale segmentation system of this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.
[0142] In addition, this application also provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. The transceiver, the memory, and the processor are respectively connected via the bus. When the computer program is executed by the processor, it implements the various processes of the above-described method embodiment for controlling output data and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0143] Exemplary electronic devices
[0144] For details, see Figure 6 As shown, this application also provides an electronic device, which includes a bus 1110, a processor 1120, a transceiver 1130, a bus interface 1140, a memory 1150, and a user interface 1160.
[0145] In this application, the electronic device further includes: a computer program stored in the memory 1150 and executable on the processor 1120, which, when executed by the processor 1120, implements the various processes of the method embodiment described above for controlling the output data.
[0146] Transceiver 1130 is used to receive and send data under the control of processor 1120.
[0147] In this application, a bus architecture (represented by bus 1110) is used. Bus 1110 may include any number of interconnected buses and bridges. Bus 1110 connects various circuits, including one or more processors represented by processor 1120 and memory represented by memory 1150.
[0148] Bus 1110 represents one or more of several types of bus architectures, including memory buses and memory controllers, peripheral buses, accelerated graphics ports, processors, or local buses using any bus architecture from various bus architectures. As an example and not a limitation, such architectures include: industry-standard architecture buses, microchannel architecture buses, extended buses, video electronics standards associations, and peripheral interconnect buses.
[0149] The processor 1120 can be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processors described above include: general-purpose processors, central processing units, network processors, digital signal processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLs), programmable logic arrays, microcontroller units or other programmable logic devices, discrete gates, transistor logic devices, and discrete hardware components. They can implement or execute the methods, steps, and logic block diagrams disclosed in this application. For example, the processor can be a single-core processor or a multi-core processor, and the processor can be integrated on a single chip or located on multiple different chips.
[0150] Processor 1120 can be a microprocessor or any conventional processor. The method steps disclosed in this application can be directly executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in readable storage media known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, registers, etc. The readable storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0151] Bus 1110 can also connect various other circuits, such as peripheral devices, voltage regulators, or power management circuits. Bus interface 1140 provides an interface between bus 1110 and transceiver 1130, all of which are well known in the art. Therefore, this application will not describe them further.
[0152] Transceiver 1130 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. For example, transceiver 1130 receives external data from other devices, and transceiver 1130 transmits data processed by processor 1120 to other devices. Depending on the nature of the computer device, a user interface 1160 may also be provided, such as a touchscreen, physical keyboard, monitor, mouse, speaker, microphone, trackball, joystick, or stylus.
[0153] It should be understood that, in this application, memory 1150 may further include memory remotely configured relative to processor 1120, and such remotely configured memory can be connected to a server via a network. One or more portions of the aforementioned network may be an ad hoc network, intranet, extranet, virtual private network, local area network, wireless local area network, wide area network, wireless wide area network, metropolitan area network, the Internet, public switched telephone network, conventional telephone network, cellular telephone network, wireless network, wireless fidelity network, and combinations of two or more of the aforementioned networks. For example, cellular telephone networks and wireless networks may be Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Global System for Microwave Interconnection and Access (GSMA), General Packet Radio Service (GPRS), Wideband Code Division Multiple Access (WDMA), Long Term Evolution (LTE), LTE Frequency Division Duplex (FDMA), LTE Time Division Duplex (TDMA), Advanced Long Term Evolution (ALE), Universal Mobile Communications (UMC), Enhanced Mobile Broadband (EMB), Massive Machine-Type Communications (MMTC), Ultra Reliable Low Latency Communications (ULSC).
[0154] It should be understood that the memory 1150 in this application may be volatile memory or non-volatile memory, or may include both volatile memory and non-volatile memory. Non-volatile memory includes: read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, or flash memory.
[0155] Volatile memory includes random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SRAM), double data rate synchronous dynamic random access memory (DRAM), enhanced synchronous dynamic random access memory (ERRAM), synchronous linked dynamic random access memory (SRAM), and direct memory bus (DMB) RAM. The memory 1150 of the electronic device described in this application includes, but is not limited to, the above-described and any other suitable types of memory.
[0156] In this application, memory 1150 stores the following elements of operating system 1151 and application program 1152: executable modules, data structures, or subsets thereof, or extended sets thereof.
[0157] Specifically, the operating system 1151 includes various device programs, such as a framework layer, a core library layer, and a driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 1152 includes various applications, such as a media player and a browser, used to implement various application functions. Programs implementing the methods of this application can be included in the application program 1152. The application program 1152 includes applets, objects, components, logic, data structures, and other computer device executable instructions that perform specific tasks or implement specific abstract data types.
[0158] In addition, this application also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the various processes of the above-described method embodiment for controlling output data and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0159] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A contextual video instance multi-scale segmentation method with polarization self-attention modulation, characterized in that, The method includes: Obtain the first video instance segmentation dataset; The label files in the first video instance segmentation dataset are extracted to construct a second video instance segmentation dataset suitable for contextual scenarios; The second video instance segmentation dataset is divided into a training set and a test set according to a predetermined ratio; Configure experimental parameters and build the model environment based on the training and test sets; Based on the experimental parameters and the model environment, the ResNet50-PSA control model with embedded single-stage and cascaded polarization self-attention mechanisms was modulated in the residual network. Construct a multi-scale spatial localization branch model that aggregates multi-granularity spatial information of target instances; Based on the ResNet50-PSA control model and the multi-scale spatial localization branch model, a contextual video instance segmentation model is constructed. The method further includes: Calculate the output of each residual block in the original residual network; Calculate the residual block output of the embedded polarization self-attention mechanism; A single-stage control model is designed by combining the residual network and different numbers of polarization self-attention mechanisms; A cascaded control model is designed by combining the residual network and the polarization self-attention mechanism at dissimilar locations; The formula for calculating the residual block output of the embedded polarization self-attention mechanism is as follows: ; In the formula, This refers to the nth original bottleneck structure residual block in the ResNet50 network. and These are channel self-attention and spatial self-attention mechanisms, respectively; The method further includes: Based on the ResNet50-PSA control model, a feature pyramid is constructed. Based on the current output layer and neighboring low-level feature maps of the feature pyramid's lateral mapping, multi-scale feature interaction is performed with the multi-scale spatial localization branch model using a bidirectional feature integration approach, including: Step 1: The lowest feature map M3 is obtained by lateral mapping of P3 through a 3×3 convolution with a stride of 1; Step 2: Downsample M3 by performing a 3×3 convolution with a stride of 2 until it is the same size as the P4 feature map; Step 3: Add the result of the horizontal mapping of P4 after a 3×3 convolution with a stride of 1 to the downsampling result in Step 2 element by element. Step 4: Perform feature integration operation on the result of Step 3 by lateral mapping of a 3×3 convolution with a stride of 2 to obtain M4; Step 5: Downsample M4 by performing a 3×3 convolution with a stride of 2 until it is the same size as the P5 feature map; Step 6: Add the result of the horizontal mapping of P5 after a 3×3 convolution with a stride of 1 to the downsampling result in Step 5 element by element. Step 7: Perform feature integration operation on the result of Step 6 by lateral mapping of a 3×3 convolution with a stride of 2 to obtain M5.
2. The method as described in claim 1, characterized in that, The method further includes: Obtain experimental configuration information; Analyze the foreground target instances in the second video instance segmentation dataset to obtain the motion characteristics of the foreground target instances; Based on the experimental configuration information and the motion characteristics, the experimental environment and experimental parameters are set up for the basic model.
3. The method as described in claim 1, characterized in that, The method further includes: The training set is used to train the model, and the test set is applied to evaluate the performance of the contextual video instance segmentation model.
4. A context-based video instance multi-scale segmentation system with polarization self-attention modulation, characterized in that, The system is used to perform the contextual video instance multi-scale segmentation method with polarization self-attention modulation as described in any one of claims 1 to 3, the system comprising: A first obtaining unit is configured to obtain a first video instance segmentation dataset. The first construction unit is used to extract the label files from the first video instance segmentation dataset and construct a second video instance segmentation dataset suitable for contextual scenarios. The first partitioning unit is used to divide the second video instance segmentation dataset into a training set and a test set according to a predetermined ratio. The first processing unit is configured to configure experimental parameters and build the model environment according to the training set and the test set. The first control unit is used to control the ResNet50-PSA control model embedded with single-level and cascaded polarization self-attention mechanisms in the residual network based on the experimental parameters and the model environment. The ResNet50 contains 49 convolutional layers and one fully connected layer. The second construction unit is used to construct a multi-scale spatial localization branch model that aggregates multi-granularity spatial information of target instances; The third building unit is used to construct a contextual video instance segmentation model based on the ResNet50-PSA control model and the multi-scale spatial localization branch model. By comparing the applicability of single-level and cascaded control models to the second video instance segmentation dataset, the best model for the ResNet50-PSA control model is selected by embedding a polarization self-attention mechanism after the fourth residual block of the residual network. On this basis, the contextual video instance segmentation model is established by combining the spatial localization branch.
5. A polarization-self-attention modulated contextual video instance multi-scale segmentation electronic device, comprising a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the transceiver, the memory, and the processor are connected via the bus, characterized in that, When the computer program is executed by the processor, it implements the steps of the method as described in any one of claims 1-3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-3.