Method and device for instance segmentation and depth estimation of video frame image

By employing a dual-ring feature interaction decoding structure and a multi-head cross-attention mechanism, the problem of insufficient feature map adaptability during the decoding process in instance segmentation and depth estimation tasks is solved, thereby improving the efficiency and accuracy of the model and making it suitable for autonomous driving and public transportation monitoring.

CN120997270APending Publication Date: 2025-11-21DALIAN NATIONALITIES UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511323421.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing instance segmentation and depth estimation tasks cannot meet the different requirements of instance segmentation and depth estimation when the decoder restores the feature map obtained by the encoder to the original image resolution. This results in the feature map being unable to adapt to the needs of the two tasks, reducing the efficiency and accuracy of the model.

Method used

A dual-ring feature interaction decoding structure is adopted. Through a shared encoder and a bidirectional cross-task interaction module, the bidirectional information transmission of features for instance segmentation and depth estimation tasks is realized. A multi-head cross-attention mechanism is used to build a bridge for information propagation between tasks, thereby enhancing the adaptability of feature maps.

Benefits of technology

It improves the collaboration between instance segmentation and depth estimation, reduces computational complexity, increases processing speed, and improves segmentation accuracy under occlusion conditions, making it suitable for autonomous driving and public transportation monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997270A_ABST
    Figure CN120997270A_ABST
Patent Text Reader

Abstract

The invention discloses a video frame image instance segmentation and depth estimation method and device, belongs to the technical field of image segmentation, aims to solve the problem of sufficient interaction of image instance segmentation and depth estimation tasks, and is characterized in that a second feature cross-task interaction module obtains a first feature image according to a query vector of a position obtained by a second feature cross-task interaction module according to a third feature image; acquiring a first favorable information feature map transmitted to a depth estimation decoder from the features of the instance segmentation decoder; according to the third feature map and the query vector of the position obtained by the first feature cross-task interaction module according to the second feature map, a second favorable information feature map transmitted to the instance segmentation decoder from the features of the depth estimation decoder is obtained, and the instance segmentation decoder outputs an instance segmentation map fused with the second favorable information feature map; a depth estimation decoder outputs a depth estimation map fusing the first favorable information feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image segmentation technology, specifically relating to a method and apparatus for instance segmentation and depth estimation of video frame images, and particularly a dual-ring feature interactive decoding structure that combines instance segmentation and depth estimation. Background Technology

[0002] Instance segmentation and depth estimation, two key tasks in computer vision, have been widely applied in many fields. Instance segmentation aims to segment each object instance in an image at the pixel level, enabling accurate identification of object boundaries. It is widely used in autonomous driving, video surveillance, and medical image analysis. Depth estimation, on the other hand, obtains depth information of each pixel in a scene through monocular or multi-view images, thereby understanding the three-dimensional structure and spatial layout of objects. It has important applications in augmented reality (AR), robot navigation, and 3D modeling.

[0003] With the emergence of Convolutional Neural Networks (CNNs) in deep learning, researchers in both fields have proposed numerous deep learning-based frameworks for image instance segmentation and depth estimation. These two technologies have developed rapidly in recent years, but in many practical applications, they are often interdependent. For example, in autonomous driving, vehicles need to know not only what the objects in front are (instance segmentation task) but also the distance between these objects and the camera (depth estimation task) to make safe driving decisions. This shows that these two tasks are usually complementary and can share a large amount of feature information. However, most existing solutions process instance segmentation and depth estimation tasks independently, first using an instance segmentation network to obtain the pixel-level boundaries of the target, and then using a depth estimation network to infer the depth of these segmented instances. However, this separate processing method has shortcomings in computational efficiency and inference accuracy. Because the two networks run independently, the complementarity between the tasks cannot be fully utilized. For example, depth information can help to segment object boundaries more accurately, while the shape and contour information of instances can provide richer contextual information for depth estimation. Furthermore, independent processing may lead to additional computational resource consumption and performance limitations, failing to meet the requirements of real-time applications. Therefore, how to effectively combine these two tasks has become a research hotspot in the field of computer vision.

[0004] A patent application titled "A Semantic Segmentation-Assisted Unsupervised Depth Estimation Method for Binocular Vision," publication number CN113096176A, proposes a fully convolutional neural network (LCN) for predicting depth maps. The LCN consists of an encoder and a decoder, with the semantic segmentation task and disparity estimation task sharing the same encoder and decoder. The method involves taking photos using a device equipped with a binocular camera to obtain color left and right views. These color views are then input into the LCN, and the predicted disparity maps output by the LCN in this scenario are used to calculate the loss function, thereby training the LCN. Finally, a single color image is input into the trained LCN, and the predicted disparity map is output, thus obtaining the predicted depth map.

[0005] A patent application titled "A Depth Estimation Method and Apparatus Based on Pyramid Segmentation Attention," with publication number CN114565655A, proposes a depth estimation method based on pyramid segmentation attention. The method includes the following steps: acquiring and preprocessing an image; performing depth estimation calculations on the preprocessed image data through a depth estimation network; and outputting a predicted depth map. Specifically, the depth estimation of the image data includes: inputting the preprocessed image data into a pyramid segmentation attention module; downsampling the preprocessed image data and sequentially inputting it from bottom to top into a multi-layer boundary fusion module, passing information from shallow layers to deeper layers to extract edge information; performing calculations on the encoder's output data through a depth correlation module and outputting it through a decoder; and fusing the decoder's output with the outputs of the pyramid segmentation attention module and the multi-layer boundary fusion module to output a predicted depth map. This method enriches the feature space, considers information from the global region, and can obtain contextual correlations, thus improving the accuracy of depth estimation.

[0006] A patent application titled "A Monocular 3D Instance Segmentation Method Guided by Depth Information," publication number CN116258734A, proposes a method for monocular 3D instance segmentation guided by depth information. This method includes an instance information and depth information acquisition module and a 3D instance segmentation module. The instance information and depth information acquisition module employs a multi-task network to simultaneously perform 2D instance segmentation and depth estimation. During instance segmentation, depth features from the depth estimation branch are introduced to further improve the accuracy of instance segmentation. The 3D instance segmentation module uses the acquired instance and depth information to generate a 3D point cloud containing instance information. Then, an adaptive point cloud filtering unit is designed to filter out noise points in the point cloud, achieving monocular 3D instance segmentation of the target. It can accurately segment 3D instances of scene targets in application scenarios with only RGB cameras and a lack of 3D instance segmentation data labels.

[0007] The aforementioned patent application combines segmentation and depth estimation tasks. However, when the decoder upsamples the feature map obtained from the encoder to restore the original image resolution, instance segmentation requires more precise boundary information, while depth estimation requires more precise geometric information. When both tasks are optimized on the same feature map, improper feature sharing can lead to the feature map being unsuitable for either precise segmentation or depth estimation, thus reducing model efficiency and accuracy. Therefore, designing an efficient network structure that allows instance segmentation and depth estimation to interact and mutually promote each other remains a significant research challenge. To better utilize the correlation matrix between instance segmentation and depth estimation, information beneficial to depth estimation can be extracted from segmentation features, or information beneficial to instance segmentation can be discovered from depth features. This invention proposes a bidirectional cross-task interaction module to model the global correlation between the two tasks. A multi-head cross-attention mechanism is used to build a bridge for information propagation between the two tasks. Through parallel bidirectional mapping of multiple subspaces, multi-head cross-attention enables precise and robust information interaction between instance segmentation and depth estimation at different semantic-geometric granularities. Summary of the Invention

[0008] To address the issue of enabling instance segmentation and depth estimation tasks in images to fully interact and complement each other, so that instance segmentation has accurate boundary information and depth estimation has accurate geometric information.

[0009] In a first aspect, a video frame image instance segmentation and depth estimation apparatus according to some embodiments of this application includes:

[0010] A shared encoder takes frame images from the input video stream and uses them to extract the first feature map.

[0011] The instance segmentation decoder receives the first feature map from the shared encoder and obtains the second feature map.

[0012] The depth estimation decoder receives the first feature map from the shared encoder and obtains the third feature map.

[0013] The first feature cross-task interaction module obtains the first advantageous information feature map passed to the depth estimation decoder from the features of the instance segmentation decoder, based on the second feature map and the query vector of the position obtained by the second feature cross-task interaction module based on the third feature map.

[0014] The second feature cross-task interaction module obtains the second advantageous information feature map passed to the instance segmentation decoder from the features of the depth estimation decoder, based on the third feature map and the query vector of the position obtained by the first feature cross-task interaction module based on the second feature map.

[0015] in:

[0016] The instance segmentation decoder outputs an instance segmentation map that fuses the second advantageous information feature map, based on the second feature map and the second advantageous information feature map of the second feature cross-task interaction module.

[0017] The depth estimation decoder outputs a depth estimation map that fuses the first advantageous information feature map, based on the third feature map and the first advantageous information feature map of the first feature cross-task interaction module.

[0018] According to some embodiments of this application, a video frame image instance segmentation and depth estimation apparatus is provided, wherein:

[0019] The first feature cross-task interaction module obtains the first matrix based on the second feature map. ;

[0020] The second feature cross-task interaction module obtains the second matrix based on the third feature map. ;

[0021] The first feature cross-task interaction module is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the first space. ;

[0022] The second feature, the cross-task interaction module, is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the second space. ;

[0023] The first feature cross-task interaction module will use the first space global correlation weight matrix. The first matrix Multiply them to obtain the first advantageous information feature map;

[0024] The second feature cross-task interaction module will use the second space global correlation weight matrix. The second matrix Multiplying them together yields the second advantageous information feature map;

[0025] Among them, the first space global correlation weight matrix As shown in the following formula:

[0026]

[0027] Among them, the second space global correlation weight matrix As shown in the following formula:

[0028]

[0029] In the formula, Indicates the scaling factor;

[0030] in, This represents the query vector at each position in the current input sequence within the first matrix. This represents the key vector at each position in the input sequence within the first matrix. This represents the value vector at each position in the input sequence within the first matrix; This represents the query vector at each position in the current input sequence within the second matrix. This represents the key vector at each position in the input sequence within the second matrix. This represents the value vector at each position in the input sequence within the second matrix.

[0031] According to some embodiments of this application, a video frame image instance segmentation and depth estimation apparatus is provided, wherein:

[0032] Second feature map Third feature map Spatial dimensions are expressed as ,in , Indicates the feature map height. Indicates the width of the feature map. Indicates the number of channels;

[0033] Second feature map After passing through LayerNorm and linear layers, the first matrix is ​​obtained; the third feature map is obtained. After passing through LayerNorm and linear layers, the first matrix is ​​obtained;

[0034] The first feature cross-task interaction module outputs the first advantageous information feature map through a linear layer, which is used to output the first advantageous information feature map.

[0035] The second feature cross-task interaction module outputs the second advantageous information feature map through a linear layer, which is used to output the second advantageous information feature map.

[0036] The instance segmentation decoder adds the second advantageous information feature map output by the second feature cross-task interaction module to the second feature map, passes through the LayerNorm layer and the multilayer perceptron, and outputs an instance segmentation map that fuses the second advantageous information feature map;

[0037] The depth estimation decoder adds the first advantageous information feature map output by the first feature cross-task interaction module to the third feature map, passes through the LayerNorm layer and the multilayer perceptron, and outputs a depth estimation map that fuses the first advantageous information feature map.

[0038] An instance segmentation and depth estimation apparatus for a video frame image according to some embodiments of this application, wherein the instance segmentation decoder includes...

[0039] The convolutional layer extracts features from the first feature map to obtain the fourth feature map.

[0040] The channel intention module obtains a weighted feature map of the channel dimension based on the fourth feature map.

[0041] The spatial attention module uses the fourth feature map as the weighted feature map of the spatial dimension.

[0042] The fusion layer adds the weighted channel dimension feature map and the weighted spatial dimension feature map to obtain the fused feature map;

[0043] The activation layer adds a non-linear factor to the fused feature map;

[0044] Pooling layers reduce the resolution of the fused feature maps and extract the main features;

[0045] An upsampling layer is used to increase the resolution of the fused feature map.

[0046] An instance segmentation and depth estimation apparatus for a video frame image according to some embodiments of this application, wherein the depth estimation decoder includes...

[0047] The convolutional layer extracts features from the first feature map to obtain the fourth feature map.

[0048] The channel intention module obtains a weighted feature map of the channel dimension based on the fourth feature map.

[0049] The spatial attention module uses the fourth feature map as the weighted feature map of the spatial dimension.

[0050] The fusion layer adds the weighted channel dimension feature map and the weighted spatial dimension feature map to obtain the fused feature map;

[0051] The activation layer adds a non-linear factor to the fused feature map;

[0052] Pooling layers reduce the resolution of the fused feature maps and extract the main features;

[0053] An upsampling layer is used to increase the resolution of the fused feature map.

[0054] According to some embodiments of this application, a video frame image instance segmentation and depth estimation apparatus is provided, wherein the channel intention module performs global average pooling on the input fourth feature map. Global channel information is compressed into channel descriptors. ;

[0055] The channel attention module generates a channel attention weight map using a multilayer perceptron model. ;

[0056] The channel intention module multiplies the weight map element-wise with the input feature map to obtain a weighted feature map of the channel dimensions. ;

[0057] in, Indicates the number of channels. Indicates the feature map height. This indicates the width of the feature map, with the subscript 'c' representing the feature map in the channel dimension. The number of channels in the feature map is represented by . The feature map has a height and width of 1.

[0058] An instance segmentation and depth estimation apparatus for a video frame image according to some embodiments of this application, wherein the spatial attention module performs segmentation on a fourth feature map. Elements at the same position in each channel are globally pooled to obtain spatial descriptors.

[0059] The spatial attention module uses 7×7 convolutional kernels to process spatial descriptors. Perform convolution operations to generate a two-dimensional spatial attention weight map. ;

[0060] The spatial attention module will use the spatial attention weight graph. With the fourth feature map Each channel is multiplied element-wise to obtain a weighted feature map with spatial dimensions. ;

[0061] in, The 1 in the middle represents the number of channels. Indicates the feature map height. Indicates the width of the feature map. Indicates the number of channels.

[0062] An instance segmentation and depth estimation apparatus for video frame images according to some embodiments of this application, wherein the video includes external video taken while driving or traffic monitoring video.

[0063] A method for instance segmentation and depth estimation of video frame images according to some embodiments of this application includes:

[0064] The shared encoder extracts the first feature map from frame images in the input video stream;

[0065] The instance segmentation decoder receives the first feature map from the shared encoder and obtains the second feature map;

[0066] The depth estimation decoder receives the first feature map from the shared encoder and obtains the third feature map;

[0067] The first feature cross-task interaction module obtains the first advantageous information feature map passed to the depth estimation decoder from the features of the instance segmentation decoder based on the second feature map and the query vector of the position obtained by the second feature cross-task interaction module based on the third feature map.

[0068] The second feature cross-task interaction module obtains the second advantageous information feature map passed to the instance segmentation decoder from the features of the depth estimation decoder based on the third feature map and the query vector of the position obtained by the first feature cross-task interaction module based on the second feature map.

[0069] in:

[0070] The instance segmentation decoder outputs an instance segmentation map that fuses the second advantageous information feature map, based on the second feature map and the second advantageous information feature map of the second feature cross-task interaction module.

[0071] The depth estimation decoder outputs a depth estimation map that fuses the first advantageous information feature map, based on the third feature map and the first advantageous information feature map of the first feature cross-task interaction module.

[0072] According to some embodiments of this application, a method for instance segmentation and depth estimation of a video frame image is provided, wherein:

[0073] The first feature cross-task interaction module obtains the first matrix based on the second feature map. ;

[0074] The second feature cross-task interaction module obtains the second matrix based on the third feature map. ;

[0075] The first feature cross-task interaction module is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the first space. ;

[0076] The second feature, the cross-task interaction module, is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the second space. ;

[0077] The first feature cross-task interaction module will use the first space global correlation weight matrix. The first matrix Multiply them to obtain the first advantageous information feature map;

[0078] The second feature cross-task interaction module will use the second space global correlation weight matrix. The second matrix Multiplying them together yields the second advantageous information feature map;

[0079] Among them, the first space global correlation weight matrix As shown in the following formula:

[0080]

[0081] Among them, the second space global correlation weight matrix As shown in the following formula:

[0082]

[0083] In the formula, Indicates the scaling factor;

[0084] in, This represents the query vector at each position in the current input sequence within the first matrix. This represents the key vector at each position in the input sequence within the first matrix. This represents the value vector at each position in the input sequence within the first matrix; This represents the query vector at each position in the current input sequence within the second matrix. This represents the key vector at each position in the input sequence within the second matrix. This represents the value vector at each position in the input sequence within the second matrix;

[0085] in:

[0086] Second feature map Third feature map Spatial dimensions are expressed as ,in , Indicates the feature map height. Indicates the width of the feature map;

[0087] Second feature map After passing through LayerNorm and linear layers, the first matrix is ​​obtained; the third feature map is obtained. After passing through LayerNorm and linear layers, the first matrix is ​​obtained;

[0088] The first feature cross-task interaction module outputs the first advantageous information feature map through a linear layer, which is used to output the first advantageous information feature map.

[0089] The second feature cross-task interaction module outputs the second advantageous information feature map through a linear layer, which is used to output the second advantageous information feature map.

[0090] The instance segmentation decoder adds the second advantageous information feature map output by the second feature cross-task interaction module to the second feature map, passes through the LayerNorm layer and the multilayer perceptron, and outputs an instance segmentation map that fuses the second advantageous information feature map;

[0091] The depth estimation decoder adds the first advantageous information feature map output by the first feature cross-task interaction module to the third feature map, passes it through a LayerNorm layer and a multilayer perceptron, and outputs a depth estimation map that fuses the first advantageous information feature map.

[0092] Beneficial effects:

[0093] In a first aspect, the apparatus and method of the present invention introduce a bidirectional feature cross-task interaction module during the decoding process to achieve bidirectional information transfer of features between instance segmentation and depth estimation tasks, fully leveraging their complementary advantages. This structure aims to enable interaction between instance segmentation and depth estimation features during the decoding phase, utilizing features from one task to enhance the performance of the other, thereby improving the overall system performance and robustness. This method not only enhances inter-task collaboration but also effectively reduces computational complexity and increases processing speed.

[0094] Secondly, this invention is applicable to target segmentation under severe occlusion conditions. The instance segmentation loop employs a dual attention mechanism of channel and spatial attention. Channel weights are adjusted globally, and spatial information is refined locally, thereby maximizing the expressive power of the features. Channel attention enhances channel features related to the target object, reducing irrelevant information introduced by occlusion. Spatial attention guides the model to focus on unoccluded areas, enhancing the response to important regions. Thus, the dual attention mechanism can automatically adjust according to the input features, adapting to different occlusion scenarios.

[0095] In a third-party context, this invention is applicable to autonomous driving technology and can be applied to computer vision environmental perception, making it suitable for the field of autonomous driving. This invention can extract and classify instance target information such as pedestrians, vehicles, and buildings in the driving environment, as well as the depth information of the entire driving environment, providing comprehensive feature information for the network model and offering crucial safety assurance for normal driving.

[0096] In the fourth aspect, the present invention is applicable to public transportation monitoring systems. The present invention's effective identification of pedestrians, vehicles and road environment meets the needs of road traffic scenarios and provides drivers with auxiliary means for safe driving. The present invention can effectively judge vehicles that violate traffic rules and pedestrians that do not abide by traffic rules, which can improve the working efficiency of public transportation systems. Attached Figure Description

[0097] Figure 1 This is a schematic diagram of the overall network architecture;

[0098] Figure 2 This is a schematic diagram of the bidirectional feature cross-task interaction module structure;

[0099] Figure 3 This is a schematic diagram of traffic road segmentation and distance measurement in Example 1, where a is the input image and b is the output image;

[0100] Figure 4 This is a schematic diagram of image segmentation and ranging in the autonomous driving system in Example 2, where a is the input image and b is the output image. Detailed Implementation

[0101] To make the technical solutions and advantages of this application clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application; that is, the described embodiments are only a part of the embodiments of this application, and not all of them. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in different configurations.

[0102] Technical Terminology Explanation:

[0103] (1) Instance segmentation: It is defined as distinguishing different objects of different categories and different individuals of the same category in a given image based on the attributes of each pixel. The goal is to identify each object instance in the image and generate an accurate pixel-level mask for each instance;

[0104] (2) Depth estimation: Using one or more RGB images from a single viewpoint or multiple viewpoints, estimate the distance of each pixel in the image relative to the shooting source. Monocular estimation based on deep learning relies on the relationship between pixel values ​​reflecting the depth relationship;

[0105] (3) Deep features: A global feature that is the result of statistical calculation of multiple pixels, describing the deep-level properties of an image or image region;

[0106] (4) Instance features: Feature information used to distinguish different instances in an image. Instance segmentation tasks not only need to identify the object categories in an image, but also need to generate accurate segmentation masks for each object to distinguish different instances.

[0107] (5) Upsampling: The purpose of upsampling is to restore a low-resolution feature map or image to a higher resolution. The purpose of upsampling is to increase (6) the number of pixels in the output image so that its size is the same as the input image, or to reach the target size.

[0108] (7) Multilayer Perceptron (MLP): A basic feedforward artificial neural network consisting of multiple neuron layers. It is a fundamental model for various tasks, including classification and regression.

[0109] (8) Normalization factor; In machine learning and deep learning, the normalization factor is used to adjust the scale of feature data. For example, in feature scaling, the normalization factor may be the maximum, minimum or standard deviation of the data, used to standardize the feature data to a specific range (such as 0 to 1) or a distribution with a mean of 0 and a standard deviation of 1.

[0110] This invention aims to address the problems of unreasonable information allocation and feature conflicts in the decoding process of instance segmentation and depth estimation tasks due to the different feature requirements of the two tasks, and also takes into account the complementary relationship between the two tasks.

[0111] This invention proposes a device for instance segmentation and depth estimation of video frame images, including a shared encoder and a dual-loop feature interactive decoding structure for joint instance segmentation and depth estimation. The structure includes: an instance segmentation decoder, a depth estimation decoder, and a cross-task interactive module.

[0112] The instance segmentation decoder generates an accurate segmentation mask for each instance from the network's feature maps, transforming the network's low-level features into usable instance segmentation results. The depth estimation decoder decodes the network's feature maps into depth maps, which represent the distance information of each pixel in the image, indicating the distance from that pixel to the camera. The decoder is responsible for converting features extracted from the image into these distance values, which can be used for ranging. The cross-task interaction module is used to enable mutual promotion between instance segmentation and depth estimation tasks, improving overall accuracy and computational efficiency.

[0113] In the dual-loop feature interactive decoding structure for joint instance segmentation and depth estimation of this invention, the inputs to the instance segmentation decoder and the depth estimation decoder are continuous video frames. as well as ,in The order of frames in the video stream is indicated. The structure of the depth estimation decoder and the instance segmentation decoder is roughly the same. The instance segmentation decoder and the depth estimation decoder take one video frame as input each time. The video frame first passes through the shared encoder to extract features, and then the extracted feature maps are sent to the instance segmentation decoder and the depth estimation decoder respectively.

[0114] The instance segmentation decoder and depth estimation decoder each consist of n decoding layers, where n = 0, 1, 2, 3, 4. These layers primarily consist of convolutional blocks, activation layers, pooling layers, and upsampling layers, with a largely consistent network structure. Each decoding layer includes an upsampling operation, increasing the spatial resolution of the feature maps by a factor of 1 with each layer. The convolutional blocks extract features from the input image. Each convolutional block consists of two convolutional layers, each using a 3×3 convolutional kernel to extract feature maps. This invention adds channel attention and spatial attention modules to each convolutional layer. The feature maps extracted by the convolutions are used for instance segmentation and depth estimation tasks, capturing contextual information from distant images in the channel and spatial dimensions, respectively.

[0115] The channel attention module of this invention, for a given input feature map, ,in This represents the number of channels in the input feature map. Indicates the feature map height. This represents the width of the feature map; global average pooling is used to compress the global channel information of the input feature map into a single channel descriptor. Where F is subscripted with c, representing the feature map of the channel dimension. The number of channels in the feature map is represented by . The feature map has a height and width of 1; then, a multilayer perceptron model is used to generate the channel attention weight map. Finally, the weighted map is multiplied element-wise with the input feature map to obtain the weighted channel-dimensional feature map. F, with subscript cout, represents the weighted feature map of the channel dimension;

[0116] Among them, the spatial attention module of the present invention is for the input feature map The elements at the same position in each channel are globally pooled to obtain a spatial descriptor. Where F is the subscript and s represents the feature map of the channel dimension, where The 1 in the middle represents the number of channels. Indicates the feature map height. The feature map width is used to capture edge information of the target image in the horizontal and vertical directions; then, a 7×7 convolution kernel is used to convolve this spatial descriptor to generate a two-dimensional spatial attention weight map. Finally, the weighted map is multiplied element-wise with each channel of the input feature map to obtain the weighted feature map of spatial dimension. The subscript F, denoted as sout, represents the feature map with weighted spatial dimensions.

[0117] The weighted channel dimension feature map Weighted spatial dimension feature map The features are summed to obtain a fused feature map. Activation layers add non-linearity to the fused feature map, pooling layers extract key features while reducing the feature map resolution, and upsampling layers perform deconvolution on the feature map to increase its resolution. The decoder uses bilinear interpolation for upsampling; similarly, each upsampling doubles the image scale and correspondingly reduces the number of channels by half. This allows video frame image feature information to be integrated into both the instance segmentation decoder and the depth estimation decoder for prediction output. Simultaneously, feature information is exchanged through a cross-task interaction module.

[0118] In this invention, the feature cross-task interaction module is located between each decoding layer of the depth estimation decoder and the instance segmentation decoder. It uses features from the two decoders to interact and generate instance-aware depth features, which refine depth prediction. At the same time, it uses scene geometry information in the depth features to enhance instance features.

[0119] Since the cross-task interaction module is a unidirectional data flow, in order to achieve bidirectional feature enhancement, two cross-task interaction modules are embedded simultaneously between each decoder layer, including a first feature cross-task interaction module and a second feature cross-task interaction module.

[0120] The first feature is the input feature of the cross-task interaction module. That is, the output of the instance segmentation decoder. The second feature is the input feature of the cross-task interaction module. This refers to the output of the depth estimation decoder. Here, the subscript 'i' represents the instance segmentation task, and 'd' represents the depth estimation task.

[0121] This invention uses one of two feature maps as a reference feature and the other as a target feature, refining the target feature using the reference feature. The following uses instance segmentation as an example to illustrate the module's operation:

[0122] For the first feature cross-task interaction module, the input features will be... Input features as target features For reference features. For the second feature cross-task interaction module, the input features will be... Input features as target features For reference, this invention employs a multi-head cross-attention mechanism to build a bridge for information propagation between tasks in order to model cross-task global correlations.

[0123] The first feature cross-task interaction module processes input features. Flatten and deform along the spatial dimension ,in , Indicates the feature map height. This indicates the width of the feature map. The second feature cross-task interaction module interacts with the input features. Flatten and deform along the spatial dimension ,in , Indicates the feature map height. Indicates the width of the feature map.

[0124] The first feature cross-task interaction module will be based on the input features. Acquired feature maps After processing by the LayerNorm (LN) layer and the linear layer, the first matrix required for the cross-attention operation of the instance segmentation task is generated. The second feature cross-task interaction module will be based on the input features. Acquired feature maps After processing by the LayerNorm (LN) layer and the linear layer, the second matrix required for the cross-attention operation of the instance segmentation task is generated. .in This represents the query vector at each position in the current input sequence. This represents the key vector at each position in the input sequence. This represents a vector of values ​​at each position in the input sequence.

[0125] The first feature cross-task interaction module calculates the first space global correlation weight matrix based on the matrix. subscript The correlation between instance segmentation and depth estimation is expressed as:

[0126]

[0127] The second feature cross-task interaction module calculates the second space global correlation weight matrix based on the matrix. subscript The correlation between depth estimation and instance segmentation is represented by the expression:

[0128]

[0129] In the formula: The scaling factor for solving the amplitude problem, i.e. The dimension of a vector This is the normalization function. , The value of the element can reflect the degree of dependence of the depth feature on the segmentation feature at that pixel. The larger the value, the greater the correlation, which implies that the segmentation feature transmits a greater proportion of information to the depth estimation branch at that pixel position.

[0130] The first feature of the cross-task interaction module will The first matrix of task branch values ​​split by instance Multiplication followed by linear layer projection can adaptively capture the first favorable information features passed from the segmentation features to the depth estimation branch. The second feature, the cross-task interaction module, will... The second matrix with the branch values ​​of the depth estimation task Multiplication followed by linear layer projection can adaptively capture second, advantageous information features passed from deep features to the instance segmentation branch. .

[0131] The instance segmentation decoder outputs the second advantageous information feature map from the second feature cross-task interaction module. Input features The summation, followed by a LayerNorm layer and a multilayer perceptron, outputs an instance segmentation map that fuses the second advantageous information feature map. The depth estimation decoder outputs the first advantageous information feature map from the first feature cross-task interaction module. Input features The summation, followed by LayerNorm layers and a multilayer perceptron, outputs a depth estimation map that fuses the first favorable information feature maps. .

[0132] Therefore, this invention uses the beneficial information features transmitted between tasks as residuals and superimposes and fuses them with the original specific task features, and integrates the information through a multilayer perceptron to obtain the output features after the interaction and fusion of each task branch. and The expressions are as follows:

[0133]

[0134]

[0135] In the formula: MLP() is a multilayer perceptron. The bidirectional cross-task interaction module represents the correlation attributes between specific task representations through a cross-attention mechanism, enabling the network to efficiently and reasonably extract complementary pattern information between tasks during feature interaction, reducing interference from noise and redundant information, and mitigating the performance degradation caused by conflicting task optimization objectives. Finally, in the independent decoders, the depth estimation decoder outputs a dense depth map, while the instance segmentation decoder generates a corresponding instance segmentation map.

[0136] Using the above scheme, this invention is applicable to target segmentation under severe occlusion conditions. The instance segmentation loop employs a dual attention mechanism of channel and spatial attention. Channel weights are adjusted globally, and spatial information is refined locally, thereby maximizing the expressive power of features. Channel attention enhances channel features related to the target object, reducing irrelevant information introduced by occlusion. Spatial attention guides the model to focus on unoccluded areas, enhancing the response to important regions. Thus, the dual attention mechanism can automatically adjust according to the input features, adapting to different occlusion scenarios.

[0137] Using the above-described solution, this invention is applicable to autonomous driving technology and can be applied to computer vision environmental perception, making it suitable for the field of autonomous driving. This invention can extract and classify instance target information such as pedestrians, vehicles, and buildings in the driving environment, as well as the depth information of the entire driving environment, providing comprehensive feature information for the network model and offering crucial safety assurance for normal driving.

[0138] Using the above solution, this invention is applicable to public transportation monitoring systems. The effective identification of pedestrians, vehicles, and road environment by this invention meets the needs of road traffic scenarios and provides drivers with auxiliary means for safe driving. This invention can effectively judge vehicles that violate traffic rules and pedestrians that do not abide by traffic rules, thereby improving the working efficiency of public transportation systems.

[0139] As described above, the method for instance segmentation and depth estimation of video frame images based on the above network structure includes...

[0140] The shared encoder extracts the first feature map from frame images in the input video stream;

[0141] The instance segmentation decoder receives the first feature map from the shared encoder and obtains the second feature map;

[0142] The depth estimation decoder receives the first feature map from the shared encoder and obtains the third feature map;

[0143] The first feature cross-task interaction module obtains the first advantageous information feature map passed to the depth estimation decoder from the features of the instance segmentation decoder based on the second feature map and the query vector of the position obtained by the second feature cross-task interaction module based on the third feature map.

[0144] The second feature cross-task interaction module obtains the second advantageous information feature map passed to the instance segmentation decoder from the features of the depth estimation decoder based on the third feature map and the query vector of the position obtained by the first feature cross-task interaction module based on the second feature map.

[0145] in:

[0146] The instance segmentation decoder outputs an instance segmentation map that fuses the second advantageous information feature map, based on the second feature map and the second advantageous information feature map of the second feature cross-task interaction module.

[0147] The depth estimation decoder outputs a depth estimation map that fuses the first advantageous information feature map, based on the third feature map and the first advantageous information feature map of the first feature cross-task interaction module.

[0148] in:

[0149] The first feature cross-task interaction module obtains the first matrix based on the second feature map. ;

[0150] The second feature cross-task interaction module obtains the second matrix based on the third feature map. ;

[0151] The first feature cross-task interaction module is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the first space. ;

[0152] The second feature, the cross-task interaction module, is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the second space. ;

[0153] The first feature cross-task interaction module will use the first space global correlation weight matrix. The first matrix Multiply them to obtain the first advantageous information feature map;

[0154] The second feature cross-task interaction module will use the second space global correlation weight matrix. The second matrix Multiplying them together yields the second advantageous information feature map;

[0155] Among them, the first space global correlation weight matrix As shown in the following formula:

[0156]

[0157] Among them, the second space global correlation weight matrix As shown in the following formula:

[0158]

[0159] In the formula, Indicates the scaling factor;

[0160] in, This represents the query vector at each position in the current input sequence within the first matrix. This represents the key vector at each position in the input sequence within the first matrix. This represents the value vector at each position in the input sequence within the first matrix; This represents the query vector at each position in the current input sequence within the second matrix. This represents the key vector at each position in the input sequence within the second matrix. This represents the value vector at each position in the input sequence within the second matrix;

[0161] in:

[0162] Second feature map Third feature map Spatial dimensions are expressed as ,in , Indicates the feature map height. Indicates the width of the feature map;

[0163] Second feature map After passing through LayerNorm and linear layers, the first matrix is ​​obtained; the third feature map is obtained. After passing through LayerNorm and linear layers, the first matrix is ​​obtained;

[0164] The first feature cross-task interaction module outputs the first advantageous information feature map through a linear layer, which is used to output the first advantageous information feature map.

[0165] The second feature cross-task interaction module outputs the second advantageous information feature map through a linear layer, which is used to output the second advantageous information feature map.

[0166] The instance segmentation decoder adds the second advantageous information feature map output by the second feature cross-task interaction module to the second feature map, passes through the LayerNorm layer and the multilayer perceptron, and outputs an instance segmentation map that fuses the second advantageous information feature map;

[0167] The depth estimation decoder adds the first advantageous information feature map output by the first feature cross-task interaction module to the third feature map, passes through the LayerNorm layer and the multilayer perceptron, and outputs a depth estimation map that fuses the first advantageous information feature map.

[0168] In this invention, the instance segmentation and depth estimation decoders utilize their respective advantageous information features to perform upsampling or decoding operations to recover the spatial resolution of the original image. The depth estimation decoder ultimately outputs a depth map, while the instance segmentation decoder generates a corresponding segmentation mask map.

[0169] Therefore, this invention proposes a dual-loop feature interaction decoding structure capable of simultaneous instance segmentation and depth estimation. This decoding structure includes an instance segmentation decoding loop and a depth estimation decoding loop. The instance segmentation decoding loop is used to acquire instance feature information; the depth estimation decoding loop is used to acquire depth information. This structure, through a feature interaction mechanism, improves the ability to segment objects and extract depth information in complex scenes, thereby enhancing the model's performance in visual understanding tasks. To accommodate feature details at different scales and avoid directly transmitting high-resolution features during network decoding, which would increase computational complexity, the output resolution of multi-scale features is set to 1 / 16 of the input image resolution. The input features of the cross-task interaction module... and middle, , These are the height and width of the 1 / 16 scale feature map, respectively. The number of channels is a characteristic.

[0170] The following is a summary of embodiments of the present invention:

[0171] (1) Example 1: Traffic road image segmentation and ranging

[0172] This embodiment inputs features extracted from road traffic scenes into the architecture to perform instance segmentation and depth estimation on instance targets such as vehicles and streetlights on the road. The traffic road image segmentation and ranging process is as follows: Figure 3 As shown.

[0173] (2) Example 2: Image segmentation of autonomous driving system

[0174] This example demonstrates inputting a road traffic scene into a network model to perform instance segmentation and ranging of pedestrians, vehicles, and other targets at an intersection within the traffic scene. The image segmentation process of an autonomous driving system is as follows: Figure 4 As shown.

[0175] Based on the above embodiments, this application also provides a computer program that, when run on a computer, causes the computer to execute the methods provided in the above embodiments.

[0176] Based on the above embodiments, this application also provides a computer storage medium storing a computer program, which, when executed by a computer, causes the computer to perform the methods provided in the above embodiments.

[0177] The storage medium can be any available medium that a computer can access. For example, but not limited to, a computer-readable medium can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer.

[0178] Based on the above embodiments, this application also provides a chip for reading a computer program stored in a memory to implement the method provided in the above embodiments.

[0179] Based on the above embodiments, this application provides a computer program product that implements the methods provided in the above embodiments when the computer program product is run on an electronic device.

[0180] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0181] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0182] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0183] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0184] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A device for instance segmentation and depth estimation of video frame images, characterized in that, include A shared encoder takes frame images from the input video stream and uses them to extract the first feature map. The instance segmentation decoder receives the first feature map from the shared encoder and obtains the second feature map. The depth estimation decoder receives the first feature map from the shared encoder and obtains the third feature map. The first feature cross-task interaction module obtains the first advantageous information feature map passed to the depth estimation decoder from the features of the instance segmentation decoder, based on the second feature map and the query vector of the position obtained by the second feature cross-task interaction module based on the third feature map. The second feature cross-task interaction module obtains the second advantageous information feature map passed to the instance segmentation decoder from the features of the depth estimation decoder, based on the third feature map and the query vector of the position obtained by the first feature cross-task interaction module based on the second feature map. in: The instance segmentation decoder outputs an instance segmentation map that fuses the second advantageous information feature map, based on the second feature map and the second advantageous information feature map of the second feature cross-task interaction module. The depth estimation decoder outputs a depth estimation map that fuses the first advantageous information feature map, based on the third feature map and the first advantageous information feature map of the first feature cross-task interaction module.

2. The device for instance segmentation and depth estimation of video frame images according to claim 1, characterized in that, in: The first feature cross-task interaction module obtains the first matrix based on the second feature map. ; The second feature cross-task interaction module obtains the second matrix based on the third feature map. ; The first feature cross-task interaction module is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the first space. ; The second feature, the cross-task interaction module, is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the second space. ; The first feature cross-task interaction module will use the first space global correlation weight matrix. The first matrix Multiply them to obtain the first advantageous information feature map; The second feature cross-task interaction module will use the second space global correlation weight matrix. The second matrix Multiplying them together yields the second advantageous information feature map; Among them, the first space global correlation weight matrix As shown in the following formula: Among them, the second space global correlation weight matrix As shown in the following formula: In the formula, Indicates the scaling factor; in, This represents the query vector at each position in the current input sequence within the first matrix. This represents the key vector at each position in the input sequence within the first matrix. This represents the value vector at each position in the input sequence within the first matrix; This represents the query vector at each position in the current input sequence within the second matrix. This represents the key vector at each position in the input sequence within the second matrix. This represents the value vector at each position in the input sequence within the second matrix.

3. The device for instance segmentation and depth estimation of video frame images according to claim 1, characterized in that, in: Second feature map Third feature map Spatial dimensions are expressed as ,in , Indicates the feature map height. Indicates the width of the feature map. Indicates the number of channels; Second feature map After passing through LayerNorm and linear layers, the first matrix is ​​obtained; the third feature map is obtained. After passing through LayerNorm and linear layers, the first matrix is ​​obtained; The first feature cross-task interaction module outputs the first advantageous information feature map through a linear layer, which is used to output the first advantageous information feature map. The second feature cross-task interaction module outputs the second advantageous information feature map through a linear layer, which is used to output the second advantageous information feature map. The instance segmentation decoder adds the second advantageous information feature map output by the second feature cross-task interaction module to the second feature map, passes through the LayerNorm layer and the multilayer perceptron, and outputs an instance segmentation map that fuses the second advantageous information feature map; The depth estimation decoder adds the first advantageous information feature map output by the first feature cross-task interaction module to the third feature map, passes through the LayerNorm layer and the multilayer perceptron, and outputs a depth estimation map that fuses the first advantageous information feature map.

4. The device for instance segmentation and depth estimation of video frame images according to claim 1, characterized in that, in, Instance segmentation decoder, including The convolutional layer extracts features from the first feature map to obtain the fourth feature map. The channel intention module obtains a weighted feature map of the channel dimension based on the fourth feature map. The spatial attention module uses the fourth feature map as the weighted feature map of the spatial dimension. The fusion layer adds the weighted channel dimension feature map and the weighted spatial dimension feature map to obtain the fused feature map; The activation layer adds a non-linear factor to the fused feature map; Pooling layers reduce the resolution of the fused feature maps and extract the main features; An upsampling layer is used to increase the resolution of the fused feature map.

5. The device for instance segmentation and depth estimation of video frame images according to claim 1, characterized in that, in, Depth estimation decoder, including The convolutional layer extracts features from the first feature map to obtain the fourth feature map. The channel intention module obtains a weighted feature map of the channel dimension based on the fourth feature map. The spatial attention module uses the fourth feature map as the weighted feature map of the spatial dimension. The fusion layer adds the weighted channel dimension feature map and the weighted spatial dimension feature map to obtain the fused feature map; The activation layer adds a non-linear factor to the fused feature map; Pooling layers reduce the resolution of the fused feature maps and extract the main features; An upsampling layer is used to increase the resolution of the fused feature map.

6. The device for instance segmentation and depth estimation of video frame images according to claim 4 or 5, characterized in that, in, The channel intention module uses global average pooling to process the input fourth feature map. Global channel information is compressed into channel descriptors. ; The channel attention module generates a channel attention weight map using a multilayer perceptron model. ; The channel intention module multiplies the weight map element-wise with the input feature map to obtain a weighted feature map of the channel dimensions. ; in, Indicates the number of channels. Indicates the feature map height. This indicates the width of the feature map, with the subscript 'c' representing the feature map in the channel dimension. The number of channels in the feature map is represented by . The feature map has a height and width of 1.

7. The device for instance segmentation and depth estimation of video frame images according to claim 4 or 5, characterized in that, in, The spatial attention module uses the fourth feature map Elements at the same position in each channel are globally pooled to obtain spatial descriptors. The spatial attention module uses 7×7 convolutional kernels to process spatial descriptors. Perform convolution operations to generate a two-dimensional spatial attention weight map. ; The spatial attention module will use the spatial attention weight graph. With the fourth feature map Each channel is multiplied element-wise to obtain a weighted feature map with spatial dimensions. ; in, The 1 in the middle represents the number of channels. Indicates the feature map height. Indicates the width of the feature map. Indicates the number of channels.

8. The device for instance segmentation and depth estimation of video frame images according to claim 1, characterized in that, in, The videos include external videos taken while driving or traffic surveillance videos.

9. A method for instance segmentation and depth estimation of video frame images, characterized in that, include The shared encoder extracts the first feature map from frame images in the input video stream; The instance segmentation decoder receives the first feature map from the shared encoder and obtains the second feature map; The depth estimation decoder receives the first feature map from the shared encoder and obtains the third feature map; The first feature cross-task interaction module obtains the first advantageous information feature map passed to the depth estimation decoder from the features of the instance segmentation decoder based on the second feature map and the query vector of the position obtained by the second feature cross-task interaction module based on the third feature map. The second feature cross-task interaction module obtains the second advantageous information feature map passed to the instance segmentation decoder from the features of the depth estimation decoder based on the third feature map and the query vector of the position obtained by the first feature cross-task interaction module based on the second feature map. in: The instance segmentation decoder outputs an instance segmentation map that fuses the second advantageous information feature map, based on the second feature map and the second advantageous information feature map of the second feature cross-task interaction module. The depth estimation decoder outputs a depth estimation map that fuses the first advantageous information feature map, based on the third feature map and the first advantageous information feature map of the first feature cross-task interaction module.

10. The method for instance segmentation and depth estimation of video frame images according to claim 9, characterized in that, in: The first feature cross-task interaction module obtains the first matrix based on the second feature map. ; The second feature cross-task interaction module obtains the second matrix based on the third feature map. ; The first feature cross-task interaction module is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the first space. ; The second feature, the cross-task interaction module, is based on the first matrix. The second matrix Calculate the global correlation weight matrix in the second space. ; The first feature cross-task interaction module will use the first space global correlation weight matrix. The first matrix Multiply them to obtain the first advantageous information feature map; The second feature cross-task interaction module will use the second space global correlation weight matrix. The second matrix Multiplying them together yields the second advantageous information feature map; Among them, the first space global correlation weight matrix As shown in the following formula: Among them, the second space global correlation weight matrix As shown in the following formula: In the formula, Indicates the scaling factor; in, This represents the query vector at each position in the current input sequence within the first matrix. This represents the key vector at each position in the input sequence within the first matrix. This represents the value vector at each position in the input sequence within the first matrix; This represents the query vector at each position in the current input sequence within the second matrix. This represents the key vector at each position in the input sequence within the second matrix. This represents the value vector at each position in the input sequence within the second matrix; in: Second feature map Third feature map Spatial dimensions are expressed as ,in , Indicates the feature map height. Indicates the width of the feature map; Second feature map After passing through LayerNorm and linear layers, the first matrix is ​​obtained; the third feature map is obtained. After passing through LayerNorm and linear layers, the first matrix is ​​obtained; The first feature cross-task interaction module outputs the first advantageous information feature map through a linear layer, which is used to output the first advantageous information feature map. The second feature cross-task interaction module outputs the second advantageous information feature map through a linear layer, which is used to output the second advantageous information feature map. The instance segmentation decoder adds the second advantageous information feature map output by the second feature cross-task interaction module to the second feature map, passes through the LayerNorm layer and the multilayer perceptron, and outputs an instance segmentation map that fuses the second advantageous information feature map; The depth estimation decoder adds the first advantageous information feature map output by the first feature cross-task interaction module to the third feature map, passes through the LayerNorm layer and the multilayer perceptron, and outputs a depth estimation map that fuses the first advantageous information feature map.

Citation Information

Patent Citations

  • Semantic segmentation assisted binocular vision unsupervised depth estimation method

    CN113096176A

  • Depth estimation method and device based on pyramid segmentation attention

    CN114565655A

  • Monocular three-dimensional instance segmentation method based on depth information guidance

    CN116258734A