A dynamic gating mechanism regulated cruise vehicle end-of-life sensitive target weakly supervised video instance segmentation method

CN118781131BActive Publication Date: 2026-09-18KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410851032.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2026-09-18
Estimated Expiration
2044-06-27

AI Technical Summary

Technical Problem

[0004]鉴于上述问题,本发明通过提供一种动态门控机制调控的巡飞器末段时敏目标弱监督视频实例分割方法,解决了巡飞器末段背景弱监督视频实例分割任务由于复杂背景、实例运动模糊性和实例尺度差异大导致目标分割精度差的问题,达到通过控制各通道权重参数对背景信息进行针对性抑制实现目标实例特征的强化,并通过特征通道间信息交互融合范式实现通道间的特征补偿增强模型对实例边界的感知能力,通过多尺度卷积加权操作帮助模型能同时关注到不同尺度的目标实例改善掩码分割质量

Benefits of technology

[0010]First, an adaptive parameter feature reconstruction unit and a feature channel interaction optimization module are embedded in the lateral propagation path of the feature pyramid. By adjusting the information weight coefficients of each channel, redundant background information is suppressed. Inter-channel information fusion is used to reduce the adverse effects of instance motion blur on feature extraction, enhancing the model's ability to perceive and recognize video instance targets. Second, a multi-resolution feature aggregator is introduced into the mask. Multi-kernel parallel convolution operations help the model focus on feature details to generate more accurate mask information, achieving a technical solution for weakly supervised video instance segmentation of time-sensitive targets at the end of a loitering pod. Furthermore, by controlling the weight parameters of each channel to specifically suppress background information and achieving feature compensation between channels through inter-channel feature interaction fusion, the model's ability to perceive instance boundaries is enhanced, improving the feature pyramid's ability to fuse dynamic instance feature information, thus improving the model's segmentation performance and mask quality. Additionally, multi-kernel parallel convolution operations help the model focus on feature details to generate more accurate mask information, improving the model's segmentation of multi-scale target instances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118781131B_ABST
    Figure CN118781131B_ABST
Patent Text Reader

Abstract

The application discloses a kind of dynamic gating mechanism regulation's cruise vehicle end section time-sensitive target weak supervision video instance segmentation method, belong to video instance segmentation technical field.The application embeds adaptive parameter feature reconstruction unit in feature pyramid horizontal transmission path to adjust each channel information weight and suppress background information in feature map;Second, feature channel interaction optimization module is embedded after the adaptive parameter feature reconstruction unit, and it is connected with the adaptive parameter feature reconstruction unit residual, the information interaction fusion mechanism between feature channels is used, feature compensation between channels is realized, to improve the perception ability of model to time-sensitive target boundary;Finally, introduce multi-resolution feature aggregator in mask branch, using multi-core parallel convolution operation, help model focus on multi-scale target feature information, to generate more accurate mask information.The application improves the precision of cruise vehicle end section time-sensitive target weak supervision video instance segmentation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video instance segmentation technology, and in particular to a method for weakly supervised video instance segmentation of time-sensitive targets in the terminal phase of a loitering rovers, controlled by a dynamic gating mechanism. Background Technology

[0002] With technological advancements, loitering pods are increasingly used in various missions, particularly in surveillance, search and rescue. During missions, loitering pods require precise target identification, location, and tracking.

[0003] Currently, visual instance segmentation methods based on loitering rovers have been extensively studied and applied to assisted driving tasks. However, static image instance segmentation lacks temporal information such as target movement speed and trajectory, making it difficult for flight systems to accurately perceive high-speed dynamic instances. Video instance segmentation can ensure fine-grained segmentation of dynamic instances while maintaining temporal consistency and spatial location correlation, helping to provide behavioral cues for dynamic instances and predict the motion trends of instances across frames. However, weakly supervised video instance segmentation tasks in the terminal phase of loitering rovers suffer from poor target segmentation accuracy due to complex backgrounds, instance motion ambiguity, and large differences in instance scale. Summary of the Invention

[0004] In view of the above problems, this invention provides a method for segmenting time-sensitive targets in weakly supervised video at the terminal stage of a loitering pod using a dynamic gating mechanism. This method solves the problem of poor target segmentation accuracy in weakly supervised video at the terminal stage of a loitering pod due to complex backgrounds, instance motion ambiguity, and large differences in instance scale. It achieves the enhancement of target instance features by controlling the weight parameters of each channel to suppress background information. Furthermore, it enhances the model's ability to perceive instance boundaries by using a feature channel information interaction and fusion paradigm. Finally, it improves the mask segmentation quality by using multi-scale convolution weighted operations to help the model simultaneously focus on target instances at different scales.

[0005] Firstly, this application provides a method for weakly supervised video instance segmentation of time-sensitive targets in the terminal phase of a loitering pod, controlled by a dynamic gating mechanism. The method includes: embedding an adaptive parameter feature reconstruction unit in the lateral transmission path of the feature pyramid, adjusting the weights of feature information in each channel through a learnable gating threshold to enhance the feature information extracted from the feature pyramid; introducing a feature channel interaction optimization module in the lateral transmission path of the feature pyramid and connecting it to the residual of the adaptive parameter feature reconstruction unit, utilizing channel interaction compensation to enhance the model's ability to perceive the boundary of time-sensitive targets and reduce the impact of target motion ambiguity on feature fusion; embedding a multi-resolution feature aggregator in the mask branch, using multi-scale convolutional weighting operations to help the model simultaneously focus on time-sensitive target instances at different scales, improving the quality of mask segmentation; and constructing a weakly supervised video instance segmentation model for time-sensitive targets in the terminal phase of a loitering pod through the adaptive parameter feature reconstruction unit, the feature channel interaction optimization module, and the multi-resolution feature aggregator.

[0006] On the other hand, this application also provides a weakly supervised video instance segmentation system for time-sensitive targets in the terminal phase of a loitering pod controlled by a dynamic gating mechanism, comprising: a first construction unit, which is used to construct an adaptive parameter feature reconstruction unit; a second construction unit, which is used to construct a feature channel interaction optimization module; a third construction unit, which is used to construct a multi-resolution feature aggregator; and a fourth construction unit, which is used to construct a weakly supervised video instance segmentation model for time-sensitive targets in the terminal phase of a loitering pod based on the adaptive parameter feature reconstruction unit, the feature channel interaction optimization module, and the multi-resolution feature aggregator.

[0007] Thirdly, this application provides an electronic device including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. The transceiver, the memory, and the processor are connected via the bus, and the computer program, when executed by the processor, implements the steps of any of the methods described above.

[0008] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0010] First, an adaptive parameter feature reconstruction unit and a feature channel interaction optimization module are embedded in the lateral propagation path of the feature pyramid. By adjusting the information weight coefficients of each channel, redundant background information is suppressed. Inter-channel information fusion is used to reduce the adverse effects of instance motion blur on feature extraction, enhancing the model's ability to perceive and recognize video instance targets. Second, a multi-resolution feature aggregator is introduced into the mask. Multi-kernel parallel convolution operations help the model focus on feature details to generate more accurate mask information, achieving a technical solution for weakly supervised video instance segmentation of time-sensitive targets at the end of a loitering pod. Furthermore, by controlling the weight parameters of each channel to specifically suppress background information and achieving feature compensation between channels through inter-channel feature interaction fusion, the model's ability to perceive instance boundaries is enhanced, improving the feature pyramid's ability to fuse dynamic instance feature information, thus improving the model's segmentation performance and mask quality. Additionally, multi-kernel parallel convolution operations help the model focus on feature details to generate more accurate mask information, improving the model's segmentation of multi-scale target instances.

[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating a method for segmenting weakly supervised video instances of time-sensitive targets in the terminal phase of a loitering rovers, controlled by a dynamic gating mechanism, according to an embodiment of the present invention.

[0013] Figure 2 This is a schematic diagram of the weakly supervised video instance segmentation network structure for time-sensitive targets in the terminal phase of a loitering rovers, according to an embodiment of the present invention.

[0014] Figure 3 This is a schematic diagram of the adaptive parameter feature reconstruction unit structure according to an embodiment of the present invention;

[0015] Figure 4 This is a schematic diagram of the feature channel interaction optimization module structure according to an embodiment of the present invention;

[0016] Figure 5 This is a schematic diagram of the multi-resolution feature aggregator structure according to an embodiment of the present invention;

[0017] Figure 6 This is a schematic diagram of the structure of a video instance segmentation system for weakly supervised time-sensitive targets in the terminal phase of a loitering rovers, controlled by a dynamic gating mechanism, according to the present invention.

[0018] Figure 7 This is a schematic diagram of the process of the present invention.

[0019] Explanation of reference numerals in the attached drawings: First building unit 601, second building unit 602, third building unit 603, fourth building unit 604, bus 1110, processor 1120, transceiver 1130, bus interface 1140, memory 1150, operating system 1151, application program 1152, and user interface 1160. Detailed Implementation

[0020] As will be apparent to those skilled in the art from the description of this application, this application can be implemented as a method, apparatus, electronic device, and computer-readable storage medium. Therefore, this application can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. Furthermore, in some embodiments, this application can also be implemented as a computer program product contained in one or more computer-readable storage media, which includes computer program code.

[0021] The aforementioned computer-readable storage medium may be any combination of one or more computer-readable storage media. Computer-readable storage media include: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media include: portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, optical disc read-only memory, optical storage devices, magnetic storage devices, or any combination thereof. In this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0022] The acquisition, storage, use, and processing of data in this application all comply with relevant national laws and regulations.

[0023] This application describes the provided methods, apparatus, and electronic devices using flowcharts and / or block diagrams.

[0024] It should be understood that each block of a flowchart and / or block diagram, as well as combinations of blocks in a flowchart and / or block diagram, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine that, when executed by a computer or other programmable data processing device, creates means for implementing the functions / operations specified in the blocks of the flowchart / block diagram.

[0025] These computer-readable program instructions may also be stored in a computer-readable storage medium that enables a computer or other programmable data processing device to function in a particular manner. In this way, the instructions stored in the computer-readable storage medium produce an instruction apparatus product that includes the functions / operations specified in the blocks of a flowchart and / or block diagram.

[0026] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer-implemented process, such that the instructions that execute on the computer or other programmable data processing apparatus provide a process for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0027] This application will now be described with reference to the accompanying drawings.

[0028] Example 1

[0029] like Figure 1-2 As shown, this application provides a method for segmenting weakly supervised video instances of time-sensitive targets in the terminal phase of a loitering rovers, controlled by a dynamic gating mechanism. The method includes:

[0030] Step S100: Embed adaptive parameter feature reconstruction units in the lateral propagation path of the feature pyramid;

[0031] Specifically, an adaptive parameter feature reconstruction unit is embedded in the lateral propagation path of the feature pyramid. By adjusting the information weight coefficients of each channel, redundant background information is suppressed, thereby enhancing the model's ability to identify video instance targets.

[0032] Step S200: Embed a feature channel interaction optimization unit in the feature pyramid lateral propagation path and connect it to the adaptive parameter feature reconstruction unit in a residual manner;

[0033] Specifically, feature channel interaction optimization units are embedded in the horizontal propagation path of the feature pyramid. The inter-channel information fusion method is used to reduce the adverse effects of instance motion blur on feature extraction and enhance the model's ability to perceive video instance targets.

[0034] Step S300: Introduce a multi-resolution feature aggregator in the mask branch;

[0035] Specifically, in the masking branch, multi-kernel parallel convolution operations are used to help the model simultaneously focus on multi-scale target feature information to generate more accurate mask information.

[0036] Step S400: Based on the adaptive parameter feature reconstruction unit, feature channel interaction optimization unit, and multi-resolution feature aggregator, build a weakly supervised video instance segmentation model for time-sensitive targets in the terminal phase of a loitering rovers.

[0037] Furthermore, such as Figure 3 As shown, the formula for step S100 is expressed as follows:

[0038]

[0039] In the formula, β and θ represent the adaptive parameters and the adaptive gating threshold, k represents the number of iterations, Gradient_Descent(·) represents the gradient descent algorithm, Softmax(·) represents the Softmax activation function, and W βk The normalized weight parameters are represented by Gate(·), which represents the gating mechanism.

[0040] Specifically, it includes:

[0041] Step 1: Use the adaptive parameter β to represent the contribution of each channel feature map to the instance feature learning, and generate the feature-related weights W by exponentially quantifying the learnable parameter β using the Softmax activation function. β The input feature map X is weighted. input Remapping to X sig ;

[0042] Step 2: Learn the dynamic threshold θ through gradient backpropagation, and adjust the feature map weights W based on the θ value. β Divide into W1 and W2;

[0043] Step 4: Place X sig Multiply the features by weight factors W1 and W2 respectively, and then multiply by the input feature X. input Multiplying them yields two weighted features: X1 and X2;

[0044] Step 5: Divide X1 and X2 into X equal parts along the channel dimension. 11 X 12 and X 21 X 22 The four elements are then cross-fused in pairs to obtain X1' and X'2.

[0045] Step 6: Finally, concatenate X1' and X'2 along the channel dimension to obtain the output X. result ;

[0046] By constructing an adaptive parameter feature reconstruction unit, a dynamic threshold gating strategy is used to adjust the channel information weights to suppress background information and enhance the feature information extracted by the feature pyramid.

[0047] Based on experimental verification of three feature cross-reconstruction methods, the results are shown in Table 1. When using method A for feature reconstruction (X... 11 With X 22 X 12 and X 21 When using the fusion method, the average segmentation accuracy of the model is improved by 1.5% compared to the baseline model, reaching 33.4%; however, when using the B method for feature reconstruction (X... 11 and X 21 X 12 and X 22 The feature reconstruction using fusion resulted in a 1.8% decrease in model accuracy compared to the baseline model, only 30.1%; however, the C-method (X... 11 and X 12 X 21 and X 22 When the features were reconstructed (through fusion), the model's average accuracy dropped significantly compared to the baseline network, reaching only 27.8%. This experimental result demonstrates the effectiveness of the feature cross-reconstruction method (Method A) of this unit.

[0048] Table 1 Comparison of three feature cross-reconstruction methods

[0049]

[0050] like Figure 4 As shown, the formula for step S200 is expressed as follows:

[0051]

[0052] In the formula, Y input Indicates input features, Y represents the splitting of the channel dimension. fusion The symbols represent the feature fusion results. Feature1(·) indicates feature extraction using method 1, and Feature2(·) indicates feature extraction using method 2. This indicates summation pixel by pixel.

[0053] Specifically, it also includes:

[0054] Step 1: Input feature tensor Y input The channel is divided into Y1 and Y2. Y is obtained by channel compression of these two. s ', where s∈{1,2};

[0055] Step 2: For the Y s ', where s∈{1,2} simultaneously perform different feature extraction operations to obtain features Y1" and Y2";

[0056] Step 3: First, new feature maps Y1" and Y2" are obtained along the channel dimension and connected to supplement channel information. Second, normalization is performed, and channel separation and fusion operations are executed to achieve information compensation between boundary feature channels.

[0057] Step 4: Finally, combine the results of the channel separation and fusion operation with the input feature Y. input The output feature map Y is obtained by performing weighted multiplication. result ;

[0058] By utilizing inter-channel information interaction compensation strategies, the model can accurately perceive dynamic instance boundary information, enhance its ability to fuse dynamic instance target feature information, and better meet the needs of cross-frame instance boundary drift for video instance segmentation tasks.

[0059] For example, as shown in Table 2, a Channel Interaction Compensation Unit (CIFC) is introduced into the lateral transmission path of the feature pyramid of the baseline network to improve the model's ability to perceive dynamic instance boundaries, thereby achieving an average segmentation accuracy of 34.6%.

[0060] Table 2 Comparison of Channel Interaction Compensation Units Introduced in the Models

[0061]

[0062] like Figure 5 As shown, the formula for step S300 is expressed as follows:

[0063] Where K∈{3,5,7}, b∈{1,2,3};

[0064] In the formula, Z input Indicates input features, λ represents a convolution operation with a kernel size of K×K, input channels of m, and output channels of n, and λ represents the result of multi-scale target feature information fusion.

[0065] Specifically, it also includes:

[0066] Step 1, Z input The input tensor undergoes three parallel K×K convolution operations, where K∈{3,5,7}, capturing multi-scale feature information Z from the input tensor. b Where b∈{1,2,3};

[0067] Step 2: Sum the features of the outputs of the three convolution operations to obtain the global attention weight map λ;

[0068] Step 3: Normalize the global attention weight map λ by row to map its weight values ​​to obtain the average attention weight map;

[0069] Step 4: Combine the average attention weight map with the input feature tensor Z.input Performing Hadamard multiplication highlights the important parts of the input features, generating a tensor Z containing multi-scale feature information. result ;

[0070] The parallel convolution scheme helps the model learn multi-scale target features separately, avoiding the aliasing of target instance feature information, and making the model applicable to video instance segmentation tasks of multi-scale targets in the terminal phase of a loitering rovers.

[0071] For example, as shown in Table 3, the introduction of a multi-resolution feature aggregator (MFF) in the mask branch of the baseline model helps the model to focus on multi-scale target feature information simultaneously, thereby improving the average segmentation accuracy of the model by 2.7%.

[0072] Table 3 Comparison of Models Incorporating Multi-resolution Feature Aggregators

[0073]

[0074] Step S400: Based on the adaptive parameter feature reconstruction unit, feature channel interaction optimization module and multi-resolution feature aggregator, construct a weakly supervised instance segmentation model for time-sensitive targets in the terminal phase of a loitering rovers.

[0075] Specifically, to address the challenge of complex backgrounds reducing the model's ability to perceive foreground features in time-sensitive target video instance segmentation tasks at the terminal stage of a loitering astronaut, an adaptive parametric feature reconstruction unit is proposed. This unit employs a dynamic threshold gating strategy to adjust channel information weights, suppressing background information and enhancing the feature information extracted by the feature pyramid. Furthermore, the rapid movement of target instances between video frames leads to large boundary displacements and drifts, interfering with feature fusion. To address this, a feature channel interaction optimization module is proposed. This module strengthens information interaction between feature channels through channel splitting, fusion, and weighting, reducing the impact of motion ambiguity on feature fusion. Finally, the significant differences in target scale within frames and the aliasing of multi-scale instance feature information pose a challenge to mask generation. To address this, a multi-resolution feature aggregator is proposed. This module uses multi-scale convolutional weighting operations to help the model simultaneously focus on target instances at different scales, improving mask segmentation quality.

[0076] In summary, the method for segmenting weakly supervised video instances of time-sensitive targets in the terminal phase of a loitering rovers, controlled by a dynamic gating mechanism, provided in this application has the following technical effects:

[0077] An adaptive parametric feature reconstruction unit and a channel feature interaction compensation unit are introduced into the feature pyramid and connected in a residual manner. This effectively suppresses redundant background information in the feature information while aggregating and compensating for inter-channel feature information to mitigate the impact of dynamic target instances on feature fusion, improving the model's boundary awareness of dynamic instances and making the instance features output by the residual network more representative. In addition, a multi-resolution feature aggregator is proposed and embedded in the mask branch, aiming to extract multi-resolution instance features from the same feature map. This allows the model to simultaneously focus on instance targets at different scales, thereby outputting complete mask information. Finally, the weakly supervised network for time-sensitive targets in the terminal phase of a loitering rovers is improved through the aforementioned adaptive parametric feature reconstruction unit, channel feature compensation unit, and multi-resolution feature aggregator.

[0078] Model segmentation accuracy.

[0079] Example 2

[0080] Based on the same inventive concept as the method for segmenting weakly supervised video instances of time-sensitive targets in the terminal phase of a loitering pod controlled by a dynamic gating mechanism in the foregoing embodiments, this invention also provides a system for segmenting weakly supervised video instances of time-sensitive targets in the terminal phase of a loitering pod controlled by a dynamic gating mechanism, such as... Figure 6 As shown, the system includes:

[0081] The first construction unit 601 is used to construct an adaptive parameter feature reconstruction unit;

[0082] The second construction unit 602 is used to construct the channel feature interaction compensation unit;

[0083] The third construction unit 603 is used to construct a multi-resolution feature aggregator;

[0084] The fourth construction unit 604 is used to construct a weakly supervised instance segmentation model for time-sensitive targets in the terminal phase of a loitering rovers through the adaptive parameter feature reconstruction unit, the channel feature interaction compensation unit, and the multi-resolution feature aggregator.

[0085] Furthermore, the system also includes:

[0086] The first design unit is used to design an adaptive parameter feature reconstruction unit for targeted suppression of background information.

[0087] The second design unit is used to design a feature channel interaction optimization module to reduce the adverse effects of motion blur on feature fusion.

[0088] The third design unit is used to design a multi-resolution feature aggregator to address the challenge of mask generation caused by the large differences in target scale within the frame and the aliasing of multi-scale instance feature information.

[0089] In addition, this application also provides a patrol drone with image acquisition function and the weakly supervised video instance segmentation network model, which can achieve the same technical effect. To avoid repetition, it will not be described in detail here.

[0090] like Figure 7 As shown, the real-time image data collected by the roving drone in this application is transmitted to a local computer and processed by the weakly supervised video instance segmentation model in this application to obtain a fine mask segmentation result, thereby providing assistance to downstream tasks.

[0091] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for segmenting weakly supervised video instances of time-sensitive targets in the terminal phase of a loitering rovers, controlled by a dynamic gating mechanism, characterized in that... The method includes: An adaptive parameter feature reconstruction unit is embedded in the horizontal transmission path of the feature pyramid. The feature information extracted by the feature pyramid is enhanced by adjusting the weight of feature information in each channel through a learnable gating threshold. The formula for the adaptive parameter feature reconstruction unit is expressed as: In the formula, β , θ This represents the adaptive parameters and adaptive gating threshold, where k represents the number of iterations. This represents the gradient descent algorithm. This represents the Softmax activation function. The weight parameters represent the normalization parameters. Indicates a gating mechanism; Specifically, it includes: Step 1: Using adaptive parameters This represents the contribution of each channel feature map to instance feature learning, and the learnable parameters are exponentially quantified using the Softmax activation function. Generate feature-related weights The input feature map is weighted. X input Remapped to X sig ; Step 2: Apply gradient backpropagation to the dynamic threshold. θ To learn, based on θ The value will be the feature map weight. W β Split into W 1 and W 2 ; Step 3, X sig respectively with weighting factors W 1 and W 2 Multiply the features and then combine them with the input features. X input Multiplying them yields two weighted features: X 1 and X 2 ; Step 4: X 1 , X 2 Each is divided into equal parts along the channel dimension. X 11 , X 12 and X 21 , X 22 The four elements are then cross-fused in pairs to obtain the desired features. and ; Step 5, finally and The output is obtained by concatenating along the channel dimension. X result ; A feature channel interaction optimization module is introduced into the horizontal transmission path of the feature pyramid and residually connected to the adaptive parameter feature reconstruction unit. Channel interaction compensation is used to help the model perceive the time-sensitive target boundary and reduce the impact of target motion ambiguity on feature fusion. The formula for the feature channel interaction optimization module is expressed as follows: In the formula, Y input Indicates input features, This indicates the splitting of the channel dimension. Y fusion Indicates the feature fusion result, This indicates that feature extraction is performed using method 1. This indicates that feature extraction is performed using two methods. This indicates summation pixel by pixel; Specifically, it also includes: Step 1: Input feature tensor Y input They are all divided into channels. Y 1 and Y 2, and obtained by channel compression of the two. ,in ; Step 2, regarding the above ,in Simultaneously perform different feature extraction operations to obtain features and ; Step 3: First, a new feature map will be obtained along the channel dimension. and The connection is made to supplement channel information, and then normalization is performed, followed by channel separation and fusion operations to achieve information compensation between boundary feature channels. Step 4: Finally, combine the results of the channel separation and fusion operation with the input features. The output feature map is obtained by performing weighted multiplication. ; By embedding a multi-resolution feature aggregator in the mask branch and using multi-scale convolution weighting operations, the model can focus on time-sensitive target instances at different scales simultaneously, thereby improving the quality of mask segmentation. A weakly supervised video instance segmentation model for time-sensitive targets in the terminal phase of a loitering pod is constructed using the adaptive parameter feature reconstruction unit, the feature channel interaction optimization module, and the multi-resolution feature aggregator.

2. The method for segmenting weakly supervised video instances of time-sensitive targets in the terminal phase of a loitering rovers, as described in claim 1, is characterized in that... The step of embedding a multi-resolution feature aggregator in the mask branch, through multi-scale convolutional weighting operations, helps the model simultaneously focus on time-sensitive target instances at different scales, thereby improving the quality of mask segmentation. In the formula, Indicates input features, Indicates that the convolution kernel is A convolution operation with m input channels and n output channels, where λ represents the result of multi-scale target feature information fusion.

3. A video instance segmentation system for weakly supervised time-sensitive targets in the terminal phase of a loitering rovers, characterized by a dynamic gating mechanism, wherein... The system includes: The first building unit is used to build an adaptive parameter feature reconstruction unit; the adaptive parameter feature reconstruction unit is embedded in the horizontal propagation path of the feature pyramid, and the feature information extracted by the feature pyramid is enhanced by adjusting the weight of feature information of each channel through a learnable gating threshold. The formula for the adaptive parameter feature reconstruction unit is expressed as: In the formula, β and θ represent the adaptive parameters and the adaptive gating threshold, respectively, and k represents the number of iterations. This represents the gradient descent algorithm. This represents the Softmax activation function. The weight parameters represent the normalization parameters. Indicates a gating mechanism; Specifically, it includes: Step 1: Use the adaptive parameter β to represent the contribution of each channel feature map to the learning of instance features, and generate feature-related weights by exponentially quantifying the learnable parameter β using the Softmax activation function. The input feature map is weighted. Remapped to ; Step 2: Learn the dynamic threshold θ using the gradient backpropagation mechanism, and weight the feature map according to the θ value. Split into and ; Step 3, respectively with weighting factors and Multiply the features and then combine them with the input features. Multiplying them yields two weighted features: and ; Step 4: , Each is divided into equal parts along the channel dimension. , and , The four elements are then cross-fused in pairs to obtain the desired features. and ; Step 5, finally and The output is obtained by concatenating along the channel dimension. ; The second building unit is used to build a feature channel interaction optimization module. The feature channel interaction optimization module is embedded in the feature pyramid horizontal propagation path and connected to the adaptive parameter feature reconstruction unit in a residual manner. The formula for the feature channel interaction optimization module is expressed as follows: In the formula, Indicates input features, This indicates the splitting of the channel dimension. Indicates the feature fusion result, This indicates that feature extraction is performed using method 1. This indicates that feature extraction is performed using two methods. This indicates summation pixel by pixel; Specifically, it also includes: Step 1: Input feature tensor They are all divided into channels. and And by compressing the two channels, we can obtain... ,in ; Step 2, regarding the above ,in Simultaneously perform different feature extraction operations to obtain features and ; Step 3: First, a new feature map will be obtained along the channel dimension. and The connection is made to supplement channel information, and then normalization is performed, followed by channel separation and fusion operations to achieve information compensation between boundary feature channels. Step 4: Finally, combine the results of the channel separation and fusion operation with the input features. The output feature map is obtained by performing weighted multiplication. ; The third building unit is used to build a multi-resolution feature aggregator. The multi-resolution feature aggregator is embedded in the mask branch. Through multi-scale convolution weighting operations, it helps the model to focus on time-sensitive target instances at different scales at the same time, thereby improving the quality of mask segmentation. The fourth construction unit is used to construct a weakly supervised instance segmentation model for time-sensitive targets in the terminal phase of a loitering rovers based on the adaptive parameter feature reconstruction unit, the feature channel interaction optimization module, and the multi-resolution feature aggregator.

4. An electronic device comprising a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the transceiver, the memory, and the processor are connected via the bus, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-2.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-2.

Citation Information

Patent Citations

  • Traffic scene weak supervision video instance segmentation method and system

    CN116805407A

  • Deep learning based robot target recognition and motion detection method, storage medium and apparatus

    US11763485B1