Unmanned aerial vehicle target tracking method based on SiamMask
Through the SiamMask-based drone target tracking method, the lightweight network and adaptive area suggestion network are used to solve the accuracy and real-time problem of target tracking in counter-drone applications, and efficient and accurate target tracking effect is achieved.
Patent Information
- Application Number
- CN202510422803.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
AI Technical Summary
The existing target tracking methods are insufficient in accuracy and real-time in anti-UAV applications, making it difficult to effectively track in complex scenarios.
SiamMask-based drone target tracking method is adopted, including lightweight network MobileNetV3, deep separable convolution, adaptive area suggestion network and multi-task tracking head module. Combined with feature extraction, area suggestion and multi-task tracking strategies, the accuracy and tracking accuracy of candidate areas are improved through adaptive anchor box generation and attention mechanisms.
Efficient and accurate target tracking is achieved in complex scenarios, with a processing frame rate of 50 frames/second, adapting to target appearance changes and complex backgrounds, and meeting real-time tracking requirements.
Smart Images

Figure CN120339331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anti - UAV, and particularly to a UAV target tracking method based on SiamMask. Background Art
[0002] With the rapid development of UAV technology, UAVs are increasingly widely used in various fields, and their application scenarios are extensive and diverse. In the military field, UAVs can be used for reconnaissance and surveillance, etc. In the civilian field, UAVs can be used for aerial photography, logistics distribution, agricultural plant protection, etc. However, the popularization of UAVs has also brought a series of security risks and challenges. On the one hand, unauthorized UAVs may break into sensitive areas such as airports, military bases, and important event venues, posing a serious threat to public safety and national security. On the other hand, malicious use of UAVs for espionage, privacy infringement and other behaviors also occur from time to time. In the field of anti - UAV, accurately tracking UAV targets is one of the key links to achieve effective interception and control. By real - time tracking of UAVs, key information such as their flight trajectories, speeds, and altitudes can be obtained, providing a basis for subsequent decision - making and actions. Only by achieving precise tracking of UAVs can effective countermeasures such as interference, interception, and capture be taken in a timely manner to ensure the safety of sensitive areas.
[0003] Single - object tracking algorithms are an important research direction in the field of computer vision. At present, there are a wide variety of target tracking algorithms, which can be mainly divided into two categories: traditional algorithms and deep - learning - based algorithms. Traditional target tracking algorithms mainly include methods based on correlation filtering and methods based on deep learning. The method based on correlation filtering determines the target position by calculating the correlation between the target template and the search area, which has the advantage of fast calculation speed, but is easily interfered in complex scenarios and has limited tracking accuracy. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems of insufficient accuracy and real - time performance of existing target tracking methods in anti - UAV applications, and provide a UAV target tracking method based on SiamMask.
[0005] The purpose of the present invention is achieved through the following technical solutions:
[0006] The tracking method of the present invention based on SiamMask adopts an optimized and innovative network structure to achieve efficient and accurate target tracking, especially in complex scenarios such as anti - UAV.
[0007] A UAV target tracking method based on SiamMask is proposed, including the following steps:
[0008] Step 1: Construct the SiamMask object tracking network, which includes a feature extraction module, a region proposal module, and a tracking head module;
[0009] Among them, in the feature extraction module, the lightweight network MobileNetV3 is adopted to ensure the model's feature extraction ability while improving the inference real-time performance. In the region proposal module, depthwise separable convolutions are used to improve the quality of the target candidate regions. In the tracking head module, more accurate regression and segmentation strategies are introduced to improve the accuracy and fineness of tracking;
[0010] The UAV target tracking method based on SiamMask mainly includes a deep feature extraction module, an adaptive region proposal network module, and a multi-task tracking head module. The deep feature extraction module uses the lightweight network MobileNetV3 to provide strong feature extraction ability for the entire tracking network while maintaining low computational cost and memory occupancy. The adaptive region proposal network module plays a key role in object tracking and is responsible for generating candidate regions where the target may exist. The multi-task tracking head module is the core part of the tracking network and is responsible for classifying, regressing, and segmenting the candidate regions to determine the position and shape of the target.
[0011] The deep feature extraction module mainly includes depthwise separable convolutions, inverted residual structures, and attention mechanisms. MobileNetV3 extensively uses depthwise separable convolutions, decomposing traditional convolution operations into depthwise convolutions and pointwise convolutions. Depthwise convolutions are responsible for processing each input channel separately, and pointwise convolutions are used to combine the outputs of depthwise convolutions. This approach significantly reduces the computational amount and the number of parameters, enabling the network to operate efficiently on resource-constrained devices. The inverted residual structure is introduced, first performing dimensionality increase through pointwise convolutions, then depthwise convolutions, and finally dimensionality reduction through pointwise convolutions. This structure reduces information loss while increasing the feature dimension, helping to extract richer features. MobileNetV3 combines various attention mechanisms, such as Squeeze-and-Excitation (SE) attention and the Hard Swish activation function. The SE attention mechanism weights the channel dimension to highlight important feature channels and suppress unimportant channels. The Hard Swish activation function improves the computational efficiency of the network while maintaining non-linearity.
[0012] The adaptive region proposal network module mainly consists of two key parts: the adaptive anchor box generation mechanism and the attention mechanism weighted processing. The scale and ratio of the anchor boxes are dynamically adjusted according to the size and shape of different targets. By analyzing the feature information of the input image, the anchor box parameters most suitable for the current target are automatically determined. For example, for larger targets, larger-sized anchor boxes are generated; for targets with irregular shapes, the ratio of the anchor boxes is adjusted to better fit the shape of the target. This adaptive anchor box generation mechanism can greatly improve the accuracy and recall rate of the candidate regions and reduce the computational amount of subsequent processing. The attention mechanism is introduced to perform weighted processing on the feature map. By calculating the attention weights at each position in the feature map, the regions where the target may appear are highlighted, and the influence of background noise is suppressed. The attention mechanism can automatically adjust the weights according to the features and context information of the target, making the network pay more attention to the target-related regions and improving the quality of the candidate regions.
[0013] The multi-task tracking head module includes three branches: the classification branch, the regression branch, and the segmentation branch. The classification branch adopts a structure that combines a multi-layer perceptron and attention fusion. The multi-layer perceptron can perform in-depth analysis and classification on the features of the candidate regions, while the attention fusion can highlight the important feature information and improve the accuracy of classification. The classification branch can more accurately judge the category of the target in the candidate region and distinguish the target from the background. The regression branch: combines a feature pyramid and deformable convolutions. The feature pyramid can obtain the position information of the target from feature maps of different scales, and the deformable convolutions can deform adaptively according to the shape of the target to improve the accuracy of regression. The regression branch can more precisely adjust the position of the target and output more accurate bounding box coordinates. The segmentation branch: uses the idea of a semantic segmentation network and introduces dilated convolutions and multi-scale fusion. Dilated convolutions can expand the receptive field without increasing the computational amount and obtain more context information; multi-scale fusion can fuse the segmentation results of different scales to improve the accuracy of segmentation. The segmentation branch can generate a high-precision mask of the target to more accurately describe the shape and position of the target.
[0014] Step 2: Train the target tracking network through a specific loss function, where the loss function includes classification loss, regression loss, and segmentation loss; among them, different loss weights are used to better balance the relationship between different tasks to train the model and optimize the training effect of the model;
[0015] Step 3: Model training process: Input a large number of infrared images of different drone targets in different scenarios collected, and label a large amount of mask label data on the images. Use a dataset and an optimization algorithm to optimize the network parameters; adopt data augmentation methods such as image flipping, scaling, and affine transformation, as well as an adaptive learning rate adjustment strategy to improve the training efficiency and model performance;
[0016] During the model training phase, an image dataset containing UAV targets is collected and organized. In addition to publicly available object tracking datasets, a large amount of video data containing different types of UAVs, different flight postures, different backgrounds, and lighting conditions is specifically collected for the anti-UAV scenario and finely annotated. The annotation includes the location and mask information of the target. When training the model, the model parameters pre-trained on a large-scale image classification dataset are used to initialize the feature extraction module, and the parameters of other modules are initialized using the Xavier method to ensure the rationality and stability of the model parameters. Parameters such as the batch size, learning rate, and number of iterations for training are set. The learning rate adopts an adaptive adjustment strategy, such as automatically adjusting the size and descent speed of the learning rate according to the change of the training loss, to better optimize the model parameters. At the same time, a learning rate warm-up mechanism is introduced to slowly increase the learning rate at the beginning of training to avoid the model falling into a local optimal solution.
[0017] In the model training process, the annotated image data is input into the model, and forward propagation calculation is performed to obtain the prediction results. According to the prediction results and the true annotations, the classification loss, regression loss, and segmentation loss are calculated, and backpropagation is performed according to the total loss function to update the model parameters. An optimization algorithm combining stochastic gradient descent with momentum and weight decay is used to improve the convergence speed and stability of the model. During the training process, the model is regularly evaluated and verified, and the hyperparameters and training strategies of the model are adjusted according to the evaluation results until the preset performance indicators are reached or the model converges. The model loss mainly includes three parts: classification loss, regression loss, and segmentation loss.
[0018] Step 4: Inference phase: By inputting a video frame sequence into the model, the trained model is used to accurately track the target, and the location and mask information of the target are output. The high-performance deep learning inference optimizer TensorRT is used to accelerate the inference model, providing low-latency and high-throughput deployment inference for the deep learning model to ensure the real-time and accuracy of tracking; the model can accurately locate the target in complex scenarios, and the tracking accuracy is greatly improved compared with traditional methods. It has strong adaptability to the appearance changes of the target (such as lighting changes, posture changes, etc.), partial occlusion, and complex backgrounds, and can meet the requirements of real-time tracking. The processing frame rate reaches more than 50 frames per second, ensuring the real-time and smoothness in practical applications.
[0019] When the model is inferring, the first frame of the video is used as the template frame, and the features of the target are extracted and saved. For each subsequent frame, it is used as the search frame and input into the model. Feature extraction is performed on the search frame to obtain the corresponding feature map. During the feature extraction process, a multi-scale feature fusion method is adopted to fuse features at different levels to obtain richer feature information. At the same time, the attention mechanism is used to weight the feature map to highlight the features related to the target and suppress the influence of background noise. Candidate regions are generated on the feature map of the search frame through the adaptive region proposal network. The features of the template frame and the candidate regions are input into the multi-task tracking head module to obtain the prediction results of classification, regression, and segmentation. The position of the target in the current frame is determined according to the regression prediction result, and the mask of the target is generated according to the segmentation prediction result. The position and mask information of the target are output as the tracking result of the current frame. At the same time, post-processing such as filtering and smoothing is performed on the tracking result to improve the reliability and accuracy of the tracking result.
[0020] Furthermore, the feature extraction module adopts a lightweight deep convolutional neural network architecture to extract the deep features of the template image and the search image, and ensure the real-time performance of model inference.
[0021] Furthermore, the region proposal network module is used to generate candidate regions where the target may exist.
[0022] Furthermore, the tracking head module includes a classification branch, a regression branch, and a segmentation branch, which are used to predict the category of the target, the target bounding box, and the target Mask respectively.
[0023] Furthermore, the classification loss adopts the cross-entropy loss function, the regression loss adopts the smooth L1 loss function, and the segmentation loss adopts the binary Logistics regression loss function.
[0024] Furthermore, the dataset includes publicly available object tracking datasets and UAV target datasets collected for specific scenarios.
[0025] Furthermore, the optimization algorithm adopts the stochastic gradient descent method.
[0026] Furthermore, the loss function algorithm is as follows:
[0027] ;
[0028] where is the number of samples, is the true label of the sample (the target is 1 and the background is 0), is the probability that the model predicts the sample as the target.
[0029] Further, the algorithm of the L1 smooth loss function is as follows:
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] Among them, are the center coordinates and width and height of the anchor box respectively, are the center coordinates and width and height of the ground truth box respectively.
[0035] Further, the algorithm of the binary Logistics regression loss function is as follows:
[0036] ;
[0037] Among them, is the masked width and height of each RoW, is the label of each RoW, is the target ground truth mask of the sample, the label value at the position of the nth RoW pixel coordinate is the mask predicted by the model for the nth RoW.
[0038] The beneficial effects of the present invention are as follows:
[0039] The UAV target tracking method based on SiamMask has significantly improved in terms of the accuracy, robustness, and real-time performance of target tracking, and is particularly suitable for target tracking tasks in complex scenarios such as anti-UAV. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is the network structure diagram of the SiamMask tracking model for the implementation method of the present invention;
[0041] Figure 2 is the general structure module of MobileNetV3 for the implementation method of the present invention;
[0042] Figure 3 is the algorithm flow chart for the implementation method of the present invention;
[0043] Figure 4 is the infrared helicopter target tracking effect diagram for the implementation method of the present invention;
[0044] Figure 5 is the infrared rotor UAV target tracking effect diagram for the implementation method of the present invention;
[0045] Figure 6 This is the visible light rotor UAV target tracking effect diagram of the implementation method of the present invention. Specific implementation manners
[0046] Next, in combination with the embodiments, the technical solutions of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0047] Refer to Figure 1 , the UAV target tracking method based on SiamMask mainly consists of a deep feature extraction module, an adaptive region proposal network module, and a multi-task tracking head module.
[0048] The deep feature extraction module adopts the lightweight network MobileNetV3, which provides powerful feature extraction capabilities for the entire tracking network while maintaining low computational costs and memory usage.
[0049] The adaptive region proposal network module plays a key role in target tracking and is responsible for generating candidate regions where the target may exist.
[0050] The multi-task tracking head module is the core part of the tracking network and is responsible for classifying, regressing, and segmenting the candidate regions to determine the position and shape of the target.
[0051] Refer to Figure 2 , the deep feature extraction module MobileNetV3 mainly includes depthwise separable convolutions, inverted residual structures, and attention mechanisms, and is composed of the general structure modules shown in Figure 2 . MobileNetV3 makes extensive use of depthwise separable convolutions, decomposing traditional convolution operations into depthwise convolutions and pointwise convolutions. Depthwise convolutions are responsible for separately processing each input channel, and pointwise convolutions are used to combine the outputs of depthwise convolutions. This approach greatly reduces the amount of computation and the number of parameters, enabling the network to operate efficiently on resource-constrained devices. The inverted residual structure is introduced, first performing dimensionality increase through pointwise convolutions, then depthwise convolutions, and finally dimensionality reduction through pointwise convolutions. This structure reduces information loss while increasing the feature dimension, helping to extract richer features. MobileNetV3 combines various attention mechanisms, such as Squeeze-and-Excitation (SE) attention and the Hard Swish activation function. The SE attention mechanism weights the channel dimension to highlight important feature channels and suppress unimportant channels. The Hard Swish activation function improves the computational efficiency of the network while maintaining non-linearity.
[0052] MobileNetV3 outputs feature maps of different scales at different network layers. The feature maps of the lower layers retain the detailed information of the target, while the feature maps of the higher layers contain more global semantic information. By extracting features at multiple levels, rich multi-scale feature representations can be obtained, providing more comprehensive information for subsequent tracking tasks. MobileNetV3 adopts a feature fusion strategy to fuse feature maps of different layers. The fused features contain both the details of the target and global context information, which helps to improve the accuracy of object tracking.
[0053] MobileNetV3 is a lightweight network with a small model size and low computational complexity. This makes it very suitable for running on resource-constrained devices such as embedded systems and mobile devices. Due to the adoption of depthwise separable convolutions and an optimized network structure, MobileNetV3 can achieve fast inference speed while maintaining high accuracy. This is crucial for real-time object tracking and can meet the real-time requirements of the tracking system. MobileNetV3 is trained on large-scale image datasets and has strong generalization ability. It can adapt to different types of targets and scenarios, and can effectively extract features and track various drone models, flight postures, and background environments in the anti-drone scenario.
[0054] The adaptive region proposal network module mainly includes two key parts: an adaptive anchor box generation mechanism and attention mechanism weighted processing. Dynamically adjust the scale and ratio of the anchor boxes according to the size and shape of different targets. By analyzing the feature information of the input image, automatically determine the anchor box parameters that are most suitable for the current target. For example, for larger targets, generate larger-sized anchor boxes; for targets with irregular shapes, adjust the ratio of the anchor boxes to better fit the shape of the target. This adaptive anchor box generation mechanism can greatly improve the accuracy and recall rate of the candidate regions, reducing the computational amount of subsequent processing. An attention mechanism is introduced to perform weighted processing on the feature maps. By calculating the attention weights at each position in the feature map, highlight the regions where the target may appear and suppress the influence of background noise. The attention mechanism can automatically adjust the weights according to the features and context information of the target, making the network pay more attention to the target-related regions and improving the quality of the candidate regions.
[0055] The multi-task tracking head module consists of three branches: the classification branch, the regression branch, and the segmentation branch. The classification branch adopts a structure that combines a multi-layer perceptron and attention fusion. The multi-layer perceptron can perform in-depth analysis and classification on the features of candidate regions, while attention fusion can highlight important feature information and improve the accuracy of classification. The classification branch can more accurately judge the category of the target in the candidate region and distinguish between the target and the background. Regression branch: Combines a feature pyramid and deformable convolutions. The feature pyramid can obtain the position information of the target from feature maps of different scales, while deformable convolutions can adaptively deform according to the shape of the target to improve the accuracy of regression. The regression branch can more precisely adjust the position of the target and output more accurate bounding box coordinates. Segmentation branch: Utilizes the idea of a semantic segmentation network and introduces dilated convolutions and multi-scale fusion. Dilated convolutions can expand the receptive field without increasing the computational cost and obtain more context information; multi-scale fusion can fuse segmentation results of different scales to improve the accuracy of segmentation. The segmentation branch can generate a high-precision mask of the target to more accurately describe the shape and position of the target.
[0056] See Figure 3 , the algorithm process is divided into a training stage and an inference stage. The model training mainly includes the following steps: 1) Dataset preparation: Collect and organize an image sequence dataset containing targets, including publicly available object tracking datasets and specific datasets collected for the anti-drone scenario. Annotate the images in the dataset, and the annotation content includes the position, category, and mask information of the target. Data augmentation techniques such as random cropping, rotation, and color transformation can be used to increase the diversity of the data. 2) Initialize the model: Initialize the deep feature extraction module with the model parameters of MobileNetV3 pre-trained on a large-scale image classification dataset, and use Xavier initialization for the parameters of other modules. 3) Set training parameters: Set parameters such as the batch size, learning rate, and number of iterations for training. The learning rate can adopt a dynamic adjustment strategy, such as a relatively large initial learning rate that gradually decreases as training progresses. 4) Training process: Input the annotated image sequence data into the model and perform forward propagation calculations to obtain prediction results. According to the prediction results and the ground truth annotations, calculate the classification loss, regression loss, and segmentation loss, and perform backpropagation according to the total loss function to update the model parameters. Repeat this process until the preset number of iterations is reached or the model converges. The model loss mainly consists of three parts: classification loss, regression loss, and segmentation loss. The cross-entropy loss function is used to calculate the classification loss, and the formula is as follows:
[0057] ;
[0058] where is the number of samples, is the sample The true label (target is 1, background is 0), is the probability that the model predicts the sample as the target. The regression loss is calculated using the smooth L1 loss function, and the formula is as follows:
[0059] ;
[0060] ;
[0061] ;
[0062] ;
[0063] where, are the center coordinates and width-height of the anchor box respectively, are the center coordinates and width-height of the true box respectively. The segmentation loss is calculated using the binary logistic regression loss function, and the formula is as follows:
[0064] ;
[0065] where, is the masked width-height of each RoW, is the label of each RoW, is the true target mask of the sample, the pixel coordinates of the nth RoW the label value at the position, is the mask that the model predicts for the nth RoW.
[0066] The total loss function is the weighted sum of the classification loss, regression loss, and segmentation loss, and the formula is as follows:
[0067] ;
[0068] where, is the weight coefficient for balancing different loss terms, which is adjusted according to the actual experimental results to optimize the training effect of the model.
[0069] Model inference mainly includes the following steps: 1) Input video frames: Take the first frame of the video as the template frame, extract the features of the target and save them. For each subsequent frame, input it as the search frame into the model. 2) Feature extraction: Use the MobileNetV3 network to extract features from the search frame and output feature maps of different scales. 3) Region proposal: Generate candidate regions on the feature map of the search frame through the adaptive region proposal network. According to the anchor box generation strategy and attention mechanism, filter out the regions that may contain the target. 4) Tracking head prediction: Input the features of the template frame and the candidate regions into the multi-task tracking head module to obtain the prediction results of classification, regression, and segmentation. The classification branch determines whether the candidate region contains the target, the regression branch outputs the bounding box coordinates of the target, and the segmentation branch generates the mask of the target. 5) Object localization and mask generation: Determine the position of the target in the current frame according to the regression prediction results, and generate the mask of the target according to the segmentation prediction results. By adjusting the bounding box coordinates and generating the mask, accurately locate the target and describe its shape. 6) Output results: Output the position and mask information of the target as the tracking result of the current frame. According to the output score and the set threshold, determine whether it is necessary to re-update the target tracking template. The tracking results can be visually displayed to intuitively observe the tracking effect of the target. As Figure 4 , Figure 5 , Figure 6 shown in the video image tracking effect diagram.
[0070] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And any changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A drone target tracking method based on SiamMask, characterized in that Comprising the following steps: Step 1: Construct a SiamMask object tracking network, which includes a feature extraction module, a region proposal module, and a tracking head module; Among them, the lightweight network MobileNetV3 is adopted in the feature extraction module, and depthwise separable convolution is used in the region proposal module; Step 2: Train the object tracking network through a loss function, and the loss function includes classification loss, regression loss, and segmentation loss; Step 3: Model training process: Input infrared images of different drone targets in different scenarios, and label mask tag data on the images, and use a dataset and an optimization algorithm to optimize network parameters; Step 4: Inference stage: Input a video frame sequence into the model, use the trained model to achieve accurate tracking of the target, and output the position and mask information of the target.
2. The method for unmanned aerial vehicle target tracking based on SiamMask according to claim 1, characterized in that The feature extraction module adopts a deep convolutional neural network architecture for extracting deep features of the template image and the search image.
3. A method for UAV target tracking based on SiamMask according to claim 1, characterized in that, The region proposal network module is used to generate candidate regions where the target may exist.
4. A drone target tracking method based on SiamMask according to claim 1, characterized in that, The tracking head module includes a classification branch, a regression branch, and a segmentation branch, which are respectively used to predict the category of the target, the target bounding box, and the target Mask.
5. A method for UAV target tracking based on SiamMask according to claim 1, characterized in that The classification loss adopts a cross-entropy loss function, the regression loss adopts a smooth L1 loss function, and the segmentation loss adopts a binary Logistics regression loss function.
6. The method for tracking an unmanned aerial vehicle target based on SiamMask according to claim 1, characterized in that, The dataset includes a publicly available object tracking dataset and a drone target dataset collected for specific scenarios.
7. A method for UAV target tracking based on SiamMask according to claim 1, characterized in that, The optimization algorithm adopts the stochastic gradient descent method.
8. A method for UAV target tracking based on SiamMask according to claim 5, characterized in that, For the cross-entropy loss function, the algorithm is: ; Among them, is the number of samples, is the sample 's true label (target is 1, background is 0), is the probability that the model predicts the sample as the target.
9. A method for UAV target tracking based on SiamMask according to claim 5, characterized in that, For the L1 smooth loss function, the algorithm is: ; ; ; ; wherein, are the center point coordinates and the width and height of the anchor box respectively, are the center point coordinates and the width and height of the ground truth box respectively.
10. A method for UAV target tracking based on SiamMask according to claim 5, characterized in that, For the binary Logistics regression loss function, the algorithm is: ; Among them, is the masked width and height for each RoW, is the label for each RoW, is the ground truth mask of the sample, the n-th RoW pixel coordinates the label value at the position, is the mask predicted by the model for the n-th RoW.
Citation Information
Cited By
Data encryption transmission method and system for remote monitoring equipment
CN122053790A