Photoelectric pod target identification and tracking system based on multi-scale attention mechanism
By employing a multi-scale attention mechanism, combined with multi-scale convolution and self-attention mechanisms, the limitations of existing optoelectronic pod systems in target recognition and tracking are overcome, enabling accurate target detection and tracking and meeting real-time requirements.
Patent Information
- Application Number
- CN202510950167.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-28
AI Technical Summary
Existing optoelectronic pod systems suffer from insufficient local feature extraction capabilities (CNN) or high computational complexity and insufficient real-time performance (Transformer) in target recognition and tracking, especially when the target scale changes, leading to tracking failure or false detection.
Employing a multi-scale attention mechanism, combining a multi-scale convolution module, a multi-head self-attention module, and a scale aggregation module, it achieves local feature extraction and capture of global dependencies through dynamic kernel adaptive feature extraction and cross-scale feature interaction, and optimizes computational complexity through a lightweight model.
The system improves the target recognition accuracy and real-time performance of the optoelectronic pod in complex scenarios, can adapt to changes in target scale, enhances robustness, and meets the robustness requirements for the recognition and tracking of different targets. The system can adapt to changes in target size, improves robustness, meets the robustness requirements for the recognition and tracking of different targets, reduces the impact of changes in target scale, and improves the real-time performance and adaptability of the system.
Smart Images

Figure CN120853002A_ABST
Abstract
Description
Technical Field
[0001] This invention provides an optoelectronic pod target recognition and tracking system based on a multi-scale attention mechanism, belonging to the field of intelligent vision technology. Background Technology
[0002] Optoelectronic pods have wide applications in military, security, and aerospace fields, with their main function being target detection, identification, and tracking. However, traditional optoelectronic pod technology has some limitations in target detection and identification. Traditional optoelectronic pods primarily rely on conventional optical and electronic technologies. In optical system design, the focus is on optimizing parameters such as the focal length and aperture of optical lenses to obtain clear images. However, in complex environments, these parameters alone are insufficient to meet the accuracy requirements for target identification. In electronic processing, although basic algorithms such as histogram equalization and filtering are used to enhance image contrast and remove noise, their adaptability to different types of targets is insufficient, and their processing efficiency is low.
[0003] With the development of intelligent vision technology in the field of computer vision, such as the emergence of technologies like convolutional neural networks (CNN) and visual transformers (ViT), new ideas and methods have been provided for the improvement of optoelectronic pod technology.
[0004] Convolutional Neural Networks (CNNs) are a commonly used architecture in deep learning, particularly suitable for image processing tasks. Electro-optical pod systems widely employ CNNs to handle target recognition tasks in video images. Their basic workflow involves extracting local features from images, such as texture and edges, through layer-by-layer convolutional operations. These local features aid in target recognition and tracking within images. By sharing convolutional kernel weights, local features can be extracted quickly and efficiently, making them particularly suitable for processing low-level features (such as edges and textures) in high-resolution images. This provides excellent results for target detection tasks in electro-optical pods, such as detecting vehicles, buildings, and ground targets in drones.
[0005] Because the receptive field of convolution operations is fixed, CNNs can only capture features within a local area of an image. For tracking long-distance or large-scale targets in electro-optical pods, CNNs cannot effectively capture the global dependencies between targets, especially when dealing with significant changes in target scale, which can easily lead to tracking failures or target loss. Furthermore, when the target size changes, single-scale convolutional kernels cannot flexibly adapt, resulting in insufficient robustness of the model to targets of different sizes. This is particularly evident when electro-optical pods need to track distant targets, especially when switching between aerial and ground targets, which can easily lead to inaccurate detection.
[0006] The Transformer model initially achieved great success in natural language processing and has been widely applied in image processing in recent years. Its core advantage lies in its self-attention mechanism, which models the relationships between each pixel and other pixels in an image, giving it global feature capture capabilities. Through self-attention, the model can capture long-range dependencies and global features. In electro-optical pod applications, especially for large-area aerial surveillance, the Transformer effectively handles multi-target monitoring and tracking in large scenes, overcoming the shortcomings of CNNs in global information capture. Because of its advantage in modeling the relationships between targets and background, and between targets themselves, the self-attention mechanism maintains high detection accuracy in target detection tasks, making it particularly suitable for multi-target tracking tasks in electro-optical pods. In complex environments (such as changing lighting and weather conditions), the Transformer exhibits better robustness than traditional convolutional networks.
[0007] The computational complexity of the self-attention mechanism increases quadratically with image resolution. This means that in electro-optical pods, when faced with high-resolution, large-scene images, the computational demands of the Transformer increase rapidly, leading to a significant drop in model inference speed and failing to meet real-time requirements. This is unacceptable for mission-critical applications such as UAVs and battlefield reconnaissance. Furthermore, processing large-scale image data requires substantial computational resources, easily causing inference delays. This poses a significant challenge in real-time target tracking applications for electro-optical pods, especially in highly dynamic scenarios where the system cannot respond promptly to target movement and changes. Summary of the Invention
[0008] This invention provides an optoelectronic pod target recognition and tracking system based on a multi-scale attention mechanism, aiming to solve the following technical problems:
[0009] 1) Existing optoelectronic pod systems based on convolutional neural networks (CNN) or Transformer self-attention mechanisms have certain advantages in local or global feature extraction, but they each have their own shortcomings: CNN is good at extracting local features, but it is difficult to capture global features; Transformer can effectively capture global features, but it has high computational complexity and insufficient real-time performance.
[0010] 2) In practical applications of optoelectronic pods, the size and distance of targets often change dynamically. Traditional single-scale detection methods struggle to handle these changes, leading to tracking failures or false detections. For example, as a target gradually approaches from a distance, its proportion in the image changes rapidly, and traditional convolutional networks or detectors with fixed receptive fields struggle to adapt to this change.
[0011] The main technical problem of this invention is how to apply a multi-scale attention mechanism in an optoelectronic pod, so that the system can effectively extract local features and capture global dependencies, achieve accurate detection and tracking of targets, and enable the system to dynamically adapt to changes in target scale, thereby improving the robustness of target detection and tracking and meeting the real-time requirements of the optoelectronic pod.
[0012] This invention proposes a target recognition and tracking system for an optoelectronic pod based on a multi-scale attention mechanism. It employs a hybrid multi-scale convolution and self-attention mechanism, and improves the system's target recognition accuracy and real-time performance through specific structural design. This technical solution combines the local feature extraction capability of convolutional neural networks with the global feature capture capability of self-attention mechanisms, effectively handling changes in target size, and ensuring real-time performance by optimizing computational complexity. The specific technical solution is as follows:
[0013] The optoelectronic pod target recognition and tracking system based on multi-scale attention mechanism consists of the following parts:
[0014] Multi-scale convolution module: used to extract local features in an image, capable of capturing detailed information at different scales.
[0015] Multi-head self-attention module: used to extract global features and model the dependencies between distant targets through a self-attention mechanism, enabling it to handle target detection in complex scenes.
[0016] Scale aggregation module: It fuses the features extracted by the multi-scale convolution module and the self-attention module to ensure the robustness of the system to the target at different scales, and performs information fusion through an inverse bottleneck structure.
[0017] Output module: Responsible for generating target location, category, and other information for subsequent target tracking.
[0018] Furthermore, the multi-scale convolution module consists of two parts: a Dynamic Kernel Adaptive Feature Extraction Structure (DKAFE) responsible for geometric adaptive enhancement of single-scale features, and a Cross-Scale Feature Interaction Network (CSFIN) to achieve global information complementarity of multi-scale features. The Dynamic Kernel Adaptive Feature Extraction Structure (DKAFE), based on the parallel processing of traditional multi-scale convolutional kernels (3×3 / 5×5 / 7×7), introduces a deformable convolutional kernel dynamic generation unit to achieve adaptive matching of the convolutional kernel shape with the target scale.
[0019] Structural design:
[0020] (1) Scale-aware branch: The scale prior of the input features is obtained through global average pooling and input into a two-layer fully connected network to generate convolution kernel size adjustment parameters;
[0021] (2) Deformable convolution generation: Based on the adjustment parameters, the offset is generated to deform the standard convolution kernel so that it can adapt to the geometry of the target (such as the narrow shape of a distant target or the irregular outline of a close target).
[0022] (3) Hybrid feature fusion: Deformation convolution features and standard convolution features are weighted and fused through a gating mechanism.
[0023] Mathematical expression:
[0024] Scale-aware parameter calculation:
[0025] s=FC2(ReLU(FC1(GlobalAvgPool(X)))) (1)
[0026] Where s∈R 3 The scaling factor corresponding to the 3×3 / 5×5 / 7×7 convolution kernel; X∈R H×W×C The input feature map is GlobalAvgPool(X)∈R. C Global average pooling is used to compress the spatial dimension; ReLU is the activation function; FC1 and FC2 are fully connected layers.
[0027] Deformable convolution offset generation:
[0028] Δp i =s i ·Tanh(Conv 3×3 (X)), (i = 1, 2, 3 correspond to convolution kernels of different scales) (2)
[0029] Where Δp i The offset of the sampling points of the convolution kernel is limited to the range [-1, 1] by the Tanh function. i Let be the scaling factor for the i-th convolutional kernel. 3×3 This is a convolution operation.
[0030] Deformation convolution calculation:
[0031]
[0032] in Let Ω represent the feature map generated after deformation convolution operation at the i-th scale; Ω is the sampling region of the standard convolution kernel, p0 is the center pixel coordinate, and w i (p) represents the weights of the deformed convolution kernel.
[0033] Hybrid Feature Fusion:
[0034]
[0035] G iσ represents the gating weight, which controls the fusion ratio of the standard and deformation features; σ is the Sigmoid activation function. The output features are those from standard convolution; Concat is the concatenation operation.
[0036] The deformable convolution and gating fusion mechanism enables the module to adapt to the geometric changes of the target, improving feature extraction accuracy by 21.3% in scenarios where the target is rotated or partially occluded.
[0037] The multi-scale convolutional module includes a cross-scale feature interaction network (CSFIN). To address the information silo problem between convolutional features of different scales, a cross-scale feature interaction module is designed to achieve multi-scale information complementarity through hierarchical feature routing.
[0038] Structural design:
[0039] (1) Establish a 3-layer scale feature pyramid (corresponding to 3×3 / 5×5 / 7×7 convolution output);
[0040] (2) Each layer of features is aligned with other layer features through upsampling / downsampling, and a gating mechanism is used to control the flow of information across scales;
[0041] (3) Introduce residual connections to avoid information loss during feature fusion.
[0042] Mathematical expression:
[0043] Scale feature alignment:
[0044]
[0045] This represents the feature map at the i-th scale. The feature map is magnified to its maximum scale through upsampling. Same space dimensions; This represents the feature map at the i-th scale. Max pooling is used to progressively downsample to the minimum scale to achieve spatial alignment of features at different scales; whereby... These correspond to the outputs of 3×3, 5×5, and 7×7 convolutions, respectively; Upsample is the upsampling operation; MaxPool is the max pooling operation, which gradually downsamples to the minimum scale.
[0046] Cross-scale gating weights:
[0047]
[0048] G i,j The interaction weights from layer i to layer j (i,j = 1, 2, 3 represent the interaction weights between different scales)
[0049] Feature interaction fusion:
[0050]
[0051] in This indicates element-wise multiplication. Through cross-scale interaction, the module can capture the associated features of the target at different scales (such as the contour of a distant target being complementary to the texture of a nearby target).
[0052] Furthermore, the multi-head self-attention module employs a multi-head self-attention mechanism, the core of which lies in calculating the relationship between queries, keys, and values. Given input features X, the process of calculating self-attention is as follows:
[0053] Q = W Q X,K=W K X,V=W V X(10)
[0054] Q, K, V are the generated query, key, and value vectors; where W Q 、W K and W V These are the weight matrices for the query, key, and value, respectively. Next, we calculate the attention weights:
[0055]
[0056] Where, d k The dimension of the key vector is used for scaling to ensure the dot product is not too large. The self-attention mechanism can capture dependencies between distant features, improving the detection and tracking capabilities of global targets.
[0057] Furthermore, in order to fuse features across multiple scales, a lightweight scale aggregation module was designed, which performs information fusion through an inverse bottleneck structure.
[0058] First, features at different scales are reduced in dimensionality using 1x1 convolutions, followed by aggregation of scale information:
[0059] Feature agg =Concat(Feature) conv Feature attn (12)
[0060] Among them, Feature conv It is the output of multi-scale convolution, Feature attn This is the output of the self-attention module, which is fused together using a cascading operation. Concat is the concatenation operation.
[0061] Subsequently, the number of channels is increased by utilizing an inverse bottleneck structure, and finally, the fused multi-scale features are obtained.
[0062] Furthermore, in the output module, the fused features are used to predict the target's category and location. For target location prediction, bounding box regression is employed, with the specific formula as follows:
[0063]
[0064] Among them, W b It is the regression weight matrix. These are the predicted bounding box coordinates. Classification is performed using a fully connected layer, and the cross-entropy loss function is used to calculate the loss.
[0065] This invention can also employ lightweight deep learning models, such as MobileNetV3 and ShuffleNet, which are designed specifically for mobile devices and low-resource scenarios, offering good computational efficiency. The system uses these lightweight models for object detection and tracking, supplemented by some traditional image processing techniques to reduce the computational burden.
[0066] The present invention can also employ a multi-scale feature pyramid network (FPN). By constructing feature pyramids at multiple levels of the CNN, FPN can effectively capture feature information at different scales and take into account the global information of the target while maintaining the ability to extract local features.
[0067] The workflow of the optoelectronic pod target recognition and tracking system based on a multi-scale attention mechanism provided by this invention is as follows:
[0068] S1. Image Input: The system for inputting images or video streams captured by the photoelectric pod.
[0069] S2. Multi-scale convolution feature extraction: The input image is processed by a multi-scale convolution module to extract local features.
[0070] S3. Self-attention feature extraction: Extract global features through a multi-head self-attention module to model the dependencies between distant targets.
[0071] S4. Scale Aggregation: Aggregates multi-scale convolutional features with self-attention features, fusing features at different scales.
[0072] S5. Target Recognition and Tracking: Generate target location and category information based on the fused features.
[0073] S6. Output Results: The system outputs the target recognition results for subsequent target tracking.
[0074] The technical effects of the technical solution provided by this invention are as follows:
[0075] (1) This invention effectively addresses the limitations of traditional convolutional neural networks (CNNs) in handling changes in target scale by introducing a multi-scale convolutional module. This module can simultaneously process target features at different scales, accurately detecting and recognizing targets regardless of whether they are small objects at a distance or large objects at close range in the image. The system can adapt to changes in the size of different targets, improving the robustness of target recognition and tracking in complex scenes and reducing recognition errors caused by changes in target size.
[0076] (2) This invention overcomes the limitation of traditional convolutional networks in capturing global features by employing a multi-head self-attention mechanism. The self-attention mechanism enables the system to establish long-distance dependencies between targets, making it particularly suitable for multiple target tracking tasks in large-scale optoelectronic pod monitoring scenarios. Even in complex scenes with widely distributed targets and significant background noise, the system can still accurately capture and track targets, significantly improving the accuracy of target recognition. Attached Figure Description
[0077] Figure 1 This is a flowchart of the optoelectronic pod target recognition and tracking system based on a multi-scale attention mechanism of the present invention. Detailed Implementation
[0078] The specific technical solutions of the present invention will be described with reference to the embodiments.
[0079] The optoelectronic pod target recognition and tracking system based on multi-scale attention mechanism consists of the following parts:
[0080] Multi-scale convolution module: used to extract local features in an image, capable of capturing detailed information at different scales.
[0081] Multi-head self-attention module: used to extract global features and model the dependencies between distant targets through a self-attention mechanism, enabling it to handle target detection in complex scenes.
[0082] Scale aggregation module: It fuses the features extracted by the multi-scale convolution module and the self-attention module to ensure the robustness of the system to the target at different scales, and performs information fusion through an inverse bottleneck structure.
[0083] Output module: Responsible for generating target location, category, and other information for subsequent target tracking.
[0084] like Figure 1 The process shown in this embodiment is as follows:
[0085] Input Example: A drone equipped with an electro-optical pod monitors ground traffic scenes in real time. The input is an RGB video stream with a resolution of 1024×768, containing vehicles (cars, trucks), pedestrians, and other targets at different distances. Targets gradually approach from a distance (occupying approximately 1% of the image pixel area, such as a car 200 meters away) to a closer distance (occupying approximately 15% of the image pixel area, such as a truck within 50 meters), accompanied by the staggered movement of multiple targets (3-5 in total), with complex background interference including trees, buildings, etc.
[0086] Step 1: Multi-scale convolution feature extraction;
[0087] Video frames are normalized to [-1, 1] and input into a multi-scale convolutional module. Three types of convolutional kernels (3×3, 5×5, and 7×7) are used in parallel to extract local features at different scales. The 3×3 kernel captures detailed features (such as pedestrian outlines and vehicle edges); the 5×5 kernel extracts medium-scale features (such as the overall shape of vehicles and pedestrian poses); and the 7×7 kernel obtains features over a larger range (such as the relative positions of multiple objects and the layout of the background environment). The feature maps output by different convolutional kernels are summed along their channel dimensions to obtain a feature map `Feature_conv` (128×128×256 pixels) containing multi-scale local information.
[0088] Step 2: Multi-head self-attention feature extraction;
[0089] The Feature_conv function is input into a multi-head self-attention module, which generates a query (Q), key (K), and value (V) matrix through linear transformation (Equation 10). The number of heads is set to 8, and each head has a dimension d. k =32 attention weights are calculated independently for each head: It captures global dependencies between targets and between targets and background (e.g., the spatial relationship between a car at a distance and a truck at a close distance, and the contextual relationship between pedestrians and the road). The outputs of the 8 heads are concatenated and passed through a linear layer to obtain the global feature Feature_attn (size 128×128×256).
[0090] Step 3: Scale aggregation and feature fusion;
[0091] We reduce the dimensionality of Feature_conv and Feature_attn to 128 channels using 1×1 convolutions to reduce computation. We then concatenate the dimensionality-reduced features along the channel dimension to obtain Concat(Feature_conv, Feature_attn) (size 128×128×256). Finally, we increase the number of channels to 512 using dilated convolutions to enhance feature representation and output the fused multi-scale feature Feature_agg.
[0092] Step 4: Target identification and location regression;
[0093] The Feature_agg is input into a fully connected layer, and a Softmax classifier outputs the target class probability (e.g., vehicle, pedestrian, background). Cross-entropy loss is used to optimize classification accuracy. A regression layer processes the Feature_agg (Equation 5) to predict the target bounding box coordinates (x1, y1, x2, y2), and smoothed L1 loss is used to optimize localization error. Duplicate detection boxes for the same target are filtered, and the result with the highest confidence is retained.
[0094] Step 5: Target tracking and result output;
[0095] The detection results of adjacent frames are correlated using the Hungarian algorithm to match target IDs, handling occlusion and scale-changing scenarios (such as when a truck approaches from a distance, the tracker maintains ID consistency through multi-scale features). The identified target category, location coordinates, and tracking ID are output to the UAV control system for path planning or target locking. The system frame rate is maintained above 30 FPS to meet real-time requirements.
[0096] The multi-scale convolution module proposed in this invention uses multiple convolution kernels of different sizes to process image features simultaneously, adapting to changes in target scale. This design effectively addresses the limitations of traditional convolutional networks in processing targets at a single scale.
[0097] This invention innovatively combines self-attention mechanisms with multi-scale convolution for target detection and tracking tasks in optoelectronic pods. Through this combination, the system can extract local features while simultaneously capturing global information, thus improving its recognition capabilities in complex scenes.
Claims
1. A photoelectric pod target recognition and tracking system based on a multi-scale attention mechanism, characterized in that, include: Multi-scale convolution module: used to extract local features in an image, capable of capturing detailed information at different scales; Multi-head self-attention module: used to extract global features and model the dependencies between distant targets through a self-attention mechanism, enabling it to handle target detection in complex scenes; Scale aggregation module: It fuses the features extracted by the multi-scale convolution module and the self-attention module to ensure the robustness of the system to the target at different scales, and performs information fusion through an inverse bottleneck structure; Output module: Responsible for generating target location, category, and other information for subsequent target tracking.
2. The photoelectric pod target recognition and tracking system based on a multi-scale attention mechanism according to claim 1, characterized in that, The multi-scale convolution module is a dynamic kernel adaptive feature extraction structure, introducing a deformable convolution kernel dynamic generation unit to achieve adaptive matching between the convolution kernel shape and the target scale, including: (1) Scale-aware branch: The scale prior of the input features is obtained through global average pooling and input into a two-layer fully connected network to generate convolution kernel size adjustment parameters; Scale-aware parameter calculation: s=FC2(ReLU(FC1(GlobalAvgPool(X)))) (1) Where s∈R 3 The scaling factor corresponding to the 3×3 / 5×5 / 7×7 convolution kernel; X∈R H×W×C The input feature map is GlobalAvgPool(X)∈R. C Global average pooling is used to compress the spatial dimension; FC1 and FC2 are fully connected layers; ReLU is the activation function. (2) Deformable convolution generation: Based on the adjustment parameters, the offset is generated to deform the standard convolution kernel so that it can adapt to the geometry of the target (such as the narrow shape of a distant target or the irregular contour of a close target). Deformable convolution offset generation: Δp i =s i ·Tanh(Conv 3×3 (X)), (i = 1, 2, 3 correspond to convolutional kernels of different scales) (2) Where Δp i The offset of the convolution kernel sampling points is limited to the range [-1, 1] by the Tanh function; s i The scaling factor for the i-th convolutional kernel; Conv 3×3 This is a convolution operation; Deformation convolution calculation: in Let Ω represent the feature map generated after deformation convolution operation at the i-th scale; Ω is the sampling region of the standard convolution kernel, p0 is the center pixel coordinate, and w i (p) represents the weights of the deformed convolution kernel; (3) Hybrid feature fusion: Deformation convolution features and standard convolution features are weighted and fused through a gating mechanism; G i σ represents the gating weight, which controls the fusion ratio of the standard and deformation features; σ is the Sigmoid activation function. represents the standard convolution output features; Concat is the concatenation operation.
3. The photoelectric pod target recognition and tracking system based on a multi-scale attention mechanism according to claim 2, characterized in that, The multi-scale convolution module includes a cross-scale feature interaction module, which achieves multi-scale information complementarity through hierarchical feature routing; The structure includes: (1) Establish a 3-layer scale feature pyramid, corresponding to 3×3 / 5×5 / 7×7 convolution output; (2) Each layer of features is aligned with other layer features through upsampling / downsampling, and a gating mechanism is used to control the flow of information across scales; (3) Introduce residual connections to avoid information loss during feature fusion; Mathematical expression: Scale feature alignment: This represents the feature map at the i-th scale. The feature map is magnified to its maximum scale through upsampling. Same space dimensions; This represents the feature map at the i-th scale. Max pooling is used to progressively downsample to the minimum scale to achieve spatial alignment of features at different scales; whereby... These correspond to the outputs of 3×3, 5×5, and 7×7 convolutions, respectively; Upsample is the upsampling operation; MaxPool is the max pooling operation, which gradually downsamples to the minimum scale. Cross-scale gating weights: G i,j Let represent the interaction weights from layer i to layer j, where i,j = 1, 2, 3 represent the interaction weights between different scales; Feature interaction fusion: in This indicates element-wise multiplication. Through cross-scale interaction, the module can capture the correlation features of the target at different scales.
4. The photoelectric pod target recognition and tracking system based on a multi-scale attention mechanism according to claim 1, characterized in that, The multi-head self-attention module employs a multi-head self-attention mechanism, the core of which lies in calculating the relationship between queries, keys, and values; given input features X, the process of calculating self-attention is as follows: Q=W Q X,K=W K X,V=W V X(10) Q, K, V are the generated query, key, and value vectors; where W Q 、W K and W V These are the weight matrices for the query, key, and value, respectively; next, we calculate the attention weights: Where, d k The dimension of the key vector is used for scaling to ensure that the dot product is not too large; the self-attention mechanism can capture the dependencies between distant features, improving the ability to detect and track global targets.
5. The optoelectronic pod target recognition and tracking system based on a multi-scale attention mechanism according to claim 1, characterized in that, The scale aggregation module described above performs information fusion through an inverse bottleneck structure; First, features at different scales are reduced in dimensionality using 1x1 convolutions, followed by aggregation of scale information: Feature agg =Concat(Feature conv ,Feature attn ) (12) Among them, Feature conv It is the output of multi-scale convolution, Feature attn It is the output of the self-attention module, which is fused together through cascading operations; Subsequently, the number of channels is increased by utilizing an inverse bottleneck structure, and finally, the fused multi-scale features are obtained.
6. The photoelectric pod target recognition and tracking system based on a multi-scale attention mechanism according to claim 1, characterized in that, The output module, after fusing the features, predicts the target's category and location. The target location prediction uses bounding box regression, with the specific formula as follows: Among them, W b It is the regression weight matrix. The coordinates of the bounding box are the predicted values; the category prediction uses a fully connected layer for classification, and the cross-entropy loss function is used to calculate the loss.
Citation Information
Cited By
Low-altitude defense scene low-slow small target identification method, terminal, medium and product
CN121170278A
Light and small multispectral photoelectric recognition system and method based on intelligent algorithm
CN121500326A
Size perception visual detection tracking device and method for mobile terminal
CN122115899A
Vehicle visual recognition method and system based on multi-scale feature fusion
CN122368964A