A small target detection method and system based on cross-scale semantic alignment
Patent Information
- Application Number
- CN202610828691.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-01
AI Technical Summary
[0003]为了解决现有技术中小目标检测精度低、标注成本高且难以实现跨尺度语义对齐的不足,本发明提供了一种基于跨尺度语义对齐的小目标检测方法及系统
本发明通过构建跨尺度统一语义空间,利用单点双拍策略获取的大尺度特写图像辅助小尺度图像检测,从根本上解决了小目标特征稀缺问题;同时借助实例级对比学习在对象查询层面拉近同一实例的跨尺度表示、推远不同实例的表示,实现了真正意义上的跨尺度知识迁移,并在训练时仅需标注大尺度图像即可自动生成小尺度图像的精确标注,推理时仅需输入小尺度图像即可输出高精度检测结果。
Smart Images

Figure CN122676151A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of agricultural information technology and computer vision, and in particular to a method and system for small target detection based on cross-scale semantic alignment. Background Technology
[0002] In the monitoring of pests and diseases in facility agriculture, drones or fixed cameras typically capture small-scale images with a wide field of view. In these images, pest and disease targets constitute a very small percentage of pixels, have blurred textures, and lack features, making them difficult for mainstream target detectors to effectively identify. To address this issue, existing technologies often focus on feature enhancement, such as multi-scale feature fusion, super-resolution reconstruction, and slice-assisted inference. However, these methods are limited to enhancing pixels or features at a single scale, failing to introduce clear texture and structural information of the corresponding targets at a larger scale, thus offering limited improvement to small target features. Furthermore, fine annotation of small-scale images is extremely costly. Manually locating tiny lesions in low-resolution images is not only time-consuming and laborious but also difficult to guarantee in terms of annotation accuracy, severely restricting the training efficiency and detection performance of models. Some studies have attempted to bridge the feature distribution differences between large and small-scale images using domain adaptation methods, but these methods only align macroscopic feature distributions and fail to establish precise semantic correspondences between the same pest or disease instance at different scales, thus failing to achieve true cross-scale knowledge transfer. Summary of the Invention
[0003] To address the shortcomings of existing technologies, such as low accuracy, high annotation costs, and difficulty in achieving cross-scale semantic alignment for small target detection, this invention provides a small target detection method and system based on cross-scale semantic alignment.
[0004] On the one hand, a small target detection method based on cross-scale semantic alignment is provided; the method includes: Obtain large-scale close-up images and small-scale scene images of the same target instance, generate cross-scale paired annotation data with pixel-level accuracy, and form a training set; Two encoder networks with identical structures and shared weights are constructed. Large-scale images and small-scale images from the training set are input into the encoder networks respectively, and large-scale feature maps and small-scale feature maps are output. A cross-view attention module is constructed, using the small-scale feature map as the query matrix and the large-scale feature map as the key matrix and value matrix, and a semantically enhanced small-scale aligned feature map is generated through a multi-head cross-attention mechanism. A decoder network based on the Transformer architecture is constructed. The semantically enhanced small-scale aligned feature maps and large-scale feature maps are input into the decoder to generate an object query set. The object query is matched with the real target annotation through the Hungarian algorithm to construct cross-scale positive and negative sample pairs. The cross-scale contrast loss is calculated based on the positive and negative sample pairs. The object detection loss function and the cross-scale contrast loss function are jointly optimized, and the parameters of the encoder network, cross-view attention module and decoder network are updated through backpropagation algorithm until the model converges.
[0005] On the other hand, a small target detection system based on cross-scale semantic alignment is provided, including a cross-scale data extraction module, a dual-stream shared encoding module, a cross-view attention alignment module, a decoder module, an instance matching module, a contrastive learning optimization module, and an inference detection module; The cross-scale data acquisition module is used to acquire large-scale close-up images and small-scale scene images of the same target instance through a single-point dual-shot strategy, and generate cross-scale paired annotation data with pixel-level accuracy to form a training set. The dual-stream shared coding module contains two encoder networks with identical structures and shared weights, which extract features from the large-scale image and the small-scale image respectively, and output large-scale feature maps and small-scale feature maps. The cross-view attention alignment module uses the small-scale feature map as the query matrix and the large-scale feature map as the key matrix and value matrix, and calculates through a multi-head cross-attention mechanism to generate a semantically enhanced small-scale aligned feature map. The decoder module, based on the Transformer architecture, decodes the small-scale aligned feature maps and large-scale feature maps of semantic enhancement respectively, generating the corresponding small-scale object query sets and large-scale object query sets. The instance matching module uses the Hungarian algorithm to match object queries with real target annotations, and filters out cross-scale positive sample object query pairs and negative sample object queries that correspond to the same physical instance. The contrastive learning optimization module calculates instance-level contrastive loss in the semantic feature space and performs end-to-end training in conjunction with object detection loss to update network parameters; The inference detection module takes only small-scale images as input during the inference stage. After processing by the encoder, cross-view attention module, and decoder, it outputs the target detection results.
[0006] The above technical solution has the following advantages or beneficial effects: This invention fundamentally solves the problem of scarce features for small targets by constructing a unified semantic space across scales and using large-scale close-up images obtained through a single-point dual-shot strategy to assist in the detection of small-scale images. At the same time, it leverages instance-level contrastive learning to bring the cross-scale representations of the same instance closer together and push the representations of different instances further apart at the object query level, achieving true cross-scale knowledge transfer. During training, only large-scale images need to be labeled to automatically generate accurate annotations for small-scale images, and during inference, only small-scale images need to be input to output high-precision detection results. Attached Figure Description
[0007] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0008] Figure 1 This is a flowchart of the method in Example 1; Figure 2 This is an example image showing the results of pest and disease detection in a small-scale image. Detailed Implementation
[0009] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0010] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the invention. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0011] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0012] All data acquisition in this embodiment is carried out in accordance with laws and regulations and with user consent, and the data is used legally.
[0013] Example 1 Figure 1 This is a flowchart illustrating a small-scale crop disease and pest detection method based on cross-scale semantic alignment, provided as an embodiment of the present invention. Figure 1 The method includes the following steps: S1. A large-scale close-up image and a small-scale scene image of the same target instance are obtained through a single-point dual-shot strategy to generate cross-scale paired annotation data with pixel-level accuracy, which constitutes the training set.
[0014] Specifically, in an agricultural environment, the single-point dual-shot strategy involves rapidly and continuously acquiring two images from different perspectives for the same pest or disease instance: one is a large-scale close-up image taken at a distance of 0.3 to 0.5 meters. One image occupies more than 10% of the image and is used to capture clear texture details; the other image is a small-scale scene image at a distance of 1.5 to 2 meters. The target occupies less than 1% of the image area. Since the interval between the two shots is extremely short and the environment is stable, the scene can be considered static.
[0015] The process of generating cross-scale paired annotation data with pixel-level accuracy involves manually annotating the large-scale close-up image to obtain target bounding boxes, and then automatically mapping the annotation information of the large-scale image to the small-scale image through feature point matching and homography matrix transformation. By only annotating the large-scale close-up image, accurate annotations for the small-scale image can be automatically obtained, minimizing the cost of manual annotation and significantly improving the efficiency of model iteration and application.
[0016] The process of automatically mapping annotation information from large-scale images to small-scale images through feature point matching and homography matrix transformation specifically involves: Large-scale close-up images and small-scale scene images Local feature points (such as SIFT, ORB, etc.) are extracted from each of the above methods to obtain the feature point set: ; in, Indicates a large-scale close-up image The set of local feature points extracted from the image, where m is the total number of feature points extracted from the image; Represents small-scale scene images The set of local feature points extracted from the image, where n is the total number of feature points extracted from the image.
[0017] Reliable matching pairs are obtained by combining nearest neighbor matching and ratio testing with RANSAC to eliminate false matches: ; Assuming that the two images satisfy planar perspective transformation (homography), there exists a homography matrix This makes it possible for any pair of matching points have:
[0018] in, This indicates homogeneity (with a non-zero scale factor). Indicates a large-scale close-up image The i-th local feature point extracted above has homogeneous coordinates in the form of: , These are the pixel coordinates of the feature point in a large-scale image; Represents small-scale scene images The extracted first A local feature point, whose homogeneous coordinate form is: , These are the pixel coordinates of the feature point in a small-scale image.
[0019] It is The homography matrix describes the planar perspective transformation relationship between large-scale and small-scale images: the transformation can be specifically written as: , usually fixed Each set of matching points provides two linear equations; these are solved iteratively using RANSAC. The model with the most interior points is selected as the final homography matrix; thus, the following is obtained. Then, the four corner points of the manually annotated target bounding box on the large-scale image are... Transform to a smaller image: in Using homogeneous coordinates, the transformed and normalized bounding boxes are obtained on the small-scale image. Because the "single-point double-shot" strategy ensures the scene is approximately static, this mapping has pixel-level accuracy.
[0020] S2. Construct two encoder networks with identical structures and shared weights. Input large-scale and small-scale images from the training set into the encoder networks respectively. Extract scale-invariant low-level visual features in parallel using the shared network parameters, and output a large-scale feature map F. l and small-scale feature map F s .
[0021] Specifically, the encoder network can be constructed using either a single-instance dual-forward approach or explicit parameter binding; The single-instance dual-forward method defines only one encoder network instance (Encoder) at the code level, which contains all trainable parameters. Large-scale and small-scale images are input into this same instance sequentially or in parallel. The display parameter binding constructs two independent encoder objects: a large-scale encoder. and small-scale encoders After initialization, parameter binding is performed. During training, the gradients of the two encoders are merged and updated, or directly set. gradient backpropagation On the parameters, ensure that the two sets of parameters remain consistent at all times; during backpropagation, the gradients from the large-scale branch and the small-scale branch will accumulate on the same set of parameters and then be updated uniformly once. In this way, the two branches not only share weights in forward computation, but also fully share parameter updates during the training process. By sharing weights, the encoder is forced to learn scale-invariant low-level visual features (such as edges, textures, and simple shapes) on both large-scale and small-scale inputs, avoiding the scale bias introduced by separate training. This is an important prerequisite for the subsequent cross-view attention module to effectively align cross-scale semantics.
[0022] Preferably, the encoder adopts a Vision Transformer (ViT) architecture, specifically including: Image segmentation: dividing the input image Cut into A fixed-size image patch, for example Each block is flattened into a one-dimensional vector; Linear projection: Each block vector is mapped to a learnable linear transformation. 3D embedding space, to obtain Embed a patch; Position encoding: Add learnable or fixed sine / cosine position encoding to each patch embedding to preserve spatial position information; Transformer encoder: Stacks multiple identical layers, each layer containing multi-head self-attention (MSA), a multilayer perceptron (MLP), layer normalization (LayerNorm), and residual connections, outputting... The feature sequence can be reshaped into The feature map is used by the subsequent detection head.
[0023] Through a weight-sharing mechanism, the encoder is forced to learn consistent semantic representations on inputs at different scales, thereby effectively extracting low-level visual features that are robust to scale changes.
[0024] S3. Construct a cross-view attention module to process the small-scale feature map F. s As a query matrix, the large-scale feature map F l As the key matrix and value matrix, a semantically enhanced small-scale aligned feature map is generated through a multi-head cross-attention mechanism. ,
[0025] in, is a small-scale feature map; Q is the query matrix, representing the target features that need to be enhanced; This represents a large-scale feature map; K is the key matrix; V is the value matrix; It serves as both a key matrix (Key, K) and a value matrix (Value, V), providing clear texture and structural information. This employs a multi-head cross-attention mechanism. Input features are mapped to multiple subspaces, attention weights are computed in parallel, and finally, these weights are concatenated and fused. Each position in the text is determined based on semantic similarity. This allows for the aggregation of information in the corresponding regions, thereby enhancing the representational ability of small-scale features.
[0026] This operation makes The features at each position can be derived from semantic relevance. The corresponding regions aggregate clearer texture and structural information, thereby enhancing the representation ability of small-scale features. Even with slight displacement due to device jitter, the attention mechanism can find the correct correspondence based on the content, demonstrating robustness.
[0027] S4. Construct a decoder network based on the Transformer architecture, input the semantically enhanced small-scale aligned feature map and large-scale feature map into the decoder to generate an object query set, match the object query with the real target annotation using the Hungarian algorithm, construct cross-scale positive and negative sample pairs, and calculate the cross-scale contrast loss based on the positive and negative sample pairs.
[0028] Specifically, S4 includes: S401. Construct a decoder network based on the Transformer architecture, inputting the semantically enhanced small-scale aligned feature maps and large-scale feature maps into the value decoder to generate the corresponding small-scale object query sets. and large-scale object query sets .
[0029] Specifically, the decoder network consists of six Transformer decoder layers stacked sequentially. Each decoder layer includes a multi-head self-attention sub-layer, a multi-head cross-attention sub-layer, and a feedforward network sub-layer. The multi-head self-attention sub-layer is used to perform global interaction modeling on the input object query set. The multi-head cross-attention sub-layer uses the object query as the query matrix and the feature map output by the encoder as the key matrix and value matrix to aggregate target context information from image features. The feedforward network sub-layer consists of two fully connected layers and a ReLU activation function, where the first fully connected layer maps the input dimension from 256 to 2048, and the second fully connected layer maps it back from 2048 to 256. Each sub-layer is followed by a residual connection and a layer normalization operation.
[0030] The input to the decoder network is a fixed number of learnable object query embedding vectors, preferably 100. The dimension of each query vector is the same as the dimension of the encoder output features, preferably 256. After 6 layers of decoder iteration and update, the same number of object query sets are output. Each query corresponds to a potential target instance, which is used by the subsequent classification head and regression head to predict the target category and bounding box coordinates.
[0031] S402, Using the Hungarian algorithm respectively and True bounding box matching, and Based on true bounding box matching, filter out valid queries that correspond to real instances. .
[0032] Specifically, the Hungarian algorithm is used to respectively... and True bounding box matching, including: Calculate the cost matrix: For the current branch, the decoder outputs a fixed number (e.g., 100) object queries, each query corresponding to a prediction (class probability distribution and bounding box). The ground truth annotations include the bounding boxes and classes of all object instances in the image. Define the prediction. With the real target The matching cost between them is:
[0033] in, For the i-th real target, including categories and bounding box ; For the first Each prediction target (i.e., a prediction output by the model) includes a class probability distribution and a prediction bounding box; The optimal matching function obtained by the Hungarian algorithm; Let be the category label for the i-th real target. If there is no real target at that location, ; Let be the bounding box coordinates of the i-th real target; For the first The bounding box coordinates of the predicted target output; To predict class probabilities, The bounding box loss is a linear combination of L1 loss and GIoU loss, used for predictions without a match (corresponding to...). The cost is fixed as a constant. It is an indicator function, also often written as or ; Finding the optimal allocation: Construct a bipartite graph by combining the predicted set and the true target set, and use the cost matrix as weights to find the unique match that minimizes the total cost using the Hungarian algorithm;
[0034] Filtering valid queries: Based on the optimal matching results, predicted queries that successfully match the true target are marked as valid queries (positive samples), and unmatched predictions are marked as negative samples, thus obtaining a large-scale set of valid queries. and small-scale efficient query set ; Constructing cross-scale positive and negative sample pairs: Based on the matching results, in and Two valid queries corresponding to the same physical instance A positive sample pair is formed across scales; queries of different instances within the same image or queries of instances in different images are formed into negative sample pairs, which are used for subsequent calculation of cross-scale contrast loss.
[0035] S403. Based on the matching results, construct cross-scale positive and negative sample pairs. Specifically, for those in... and The same physical instance that exists simultaneously in the query ( , A query for a single instance in the same image or a query for instances in different images constitutes a positive sample pair across scales. A query for a single instance in the same image or a query for an instance in different images constitutes a negative sample pair.
[0036] S404. Apply instance-level cross-scale contrastive loss to the semantic feature space of the object query embedding. This embodiment uses the InfoNCE loss function, which is defined as follows:
[0037] in For cosine similarity, ; Temperature coefficient; The number of positive sample pairs, i.e., the summation sign. The total number of positive sample pairs, used to average the loss; It is a set of positive sample pairs; each positive sample pair Embedded by a large-scale object query (from) ) and a small-scale object query embedding (from) The two are composed of the same physical instance (i.e., they are matched to the same pest target). Given a negative sample set, for the current positive sample pair... The negative sample set contains all object query embeddings that do not match: queries corresponding to other instances in the same image, instance queries in different images (other batches or momentum queues), and so on. Any other query that does not belong to the same physical instance; negative sample set For any negative sample in the query embedding vector; for all Calculate similarity And sum them up for normalization.
[0038] This loss directly applies to the semantic embedding at the instance level, forcing cross-scale representations of the same instance to be close together in the semantic space while pushing different instances apart, thereby constructing a scale-independent, highly discriminative unified semantic space. To increase the diversity of negative samples, a momentum queue mechanism can also be introduced to sample negative samples from historical batches.
[0039] S5. Jointly optimize the target detection loss function and the cross-scale contrast loss function, and update the parameters of the encoder network, cross-view attention module and decoder network through the backpropagation algorithm until the model converges.
[0040] Specifically, S5 includes: S501. Combine object detection loss and cross-scale contrast loss to construct the total loss function:
[0041] in It is the standard detection loss of DETR (including L1 loss, GIoU loss and classification cross-entropy loss). It is a balancing factor.
[0042] S502, Transfer the small-scale images of the current batch and large-scale images The data is passed through a two-stream encoder, a cross-view attention module, and a decoder in sequence for forward propagation, resulting in a small-scale object query set. and large-scale object query sets The total loss for the current batch is calculated based on the predicted bounding box and category.
[0043] S503. With the goal of minimizing the total loss, calculate the gradient of the loss function with respect to all trainable parameters in the network.
[0044] S504. Parameter update via gradient backpropagation S505. Repeat steps S501-S504 to perform multiple iterations on all batches of data in the training set. After each iteration, evaluate the model performance on the validation set. When the validation set metric no longer improves for several consecutive iterations or reaches the preset number of training iterations, the model is considered to have converged, and the current model weights are saved as the final training result.
[0045] S6. During the inference phase, only the small-scale image to be detected is input. After processing by the encoder network, cross-view attention module, and decoder network, the target detection result, including the target category and bounding box coordinates, is output. The entire inference process does not require the participation of large-scale images, maintaining the same inference efficiency as conventional detectors.
[0046] The specific process is as follows: S601, The small-scale scene image to be detected (size is) Input it into the trained model.
[0047] S602. Extract small-scale feature maps using a weight-sharing encoder.
[0048] Although the encoder processes both large and small-scale images simultaneously during training, it only uses one branch during inference; The data is fed into a weight-sharing encoder network (ViT architecture) to extract small-scale feature maps. ; The encoder output shape is The feature map, where For patch size, For feature dimensions.
[0049] S603, Cross-view attention module processing During training, this module uses small-scale features as queries. Large-scale features are the key Sum of values $V$, enabling cross-scale enhancement; During inference, since there are no large-scale images, small-scale feature maps are used simultaneously. , , Input to the same multi-head cross-attention module (i.e., degenerate into self-attention): ; or equivalently, directly order (Skip module); This design ensures that large-scale data is not required during the inference phase and does not incur additional computational overhead.
[0050] S604, Decoder generates object query Align the semantically enhanced small-scale feature maps The input is fed into a Transformer decoder network. The decoder consists of 6 stacked (tunable) decoder layers, each containing self-attention and cross-attention (using...). for ) and feedforward networks.
[0051] The decoder takes a fixed number of learnable object query embeddings (e.g., 100) as input and outputs the same number of object query sets. .
[0052] S605, Prediction head outputs detection results Query each object The data are fed into the classification head (linear layer + Softmax) and the regression head (linear layer) respectively to obtain the class probability distribution. and bounding box coordinates ; in, , Represents the center coordinates; , Represents the width and height of the sides; By setting a confidence threshold (e.g., 0.5) and non-maximum suppression (NMS), the final detection results, including the target category and bounding box coordinates, are filtered out.
[0053] S606. Output all detected pest and disease targets in the small-scale image, in the following format:
[0054] in, For the detected first The index of each target instance, with values ranging from 1 to K; This represents the total number of pests and diseases finally detected in the small-scale image (the result after confidence thresholding and non-maximum suppression). For the first The category label of an individual target (such as "rust", "powdery mildew", etc.) is usually represented by an integer ID or a string; For the first Confidence score for each target, range of values This indicates the reliability of the model's prediction; only targets with scores higher than a preset threshold (such as 0.5) will be output. For the first The coordinates of the bounding box of the target. The pixel coordinates of the top-left corner of the bounding box in the image; These are the pixel coordinates of the bottom right corner of the bounding box in the image.
[0055] like Figure 2 The image shown is an example of the detection results of the method of the present invention on a small-scale image. Figure 2 a and c are the original input images. Figure 2 b and d represent the identified pests and diseases. The model successfully detected Verticillium wilt lesions in small-scale images, and the method has a high accuracy rate in identifying tiny pests and diseases.
[0056] Example 2 This embodiment provides a small-scale crop pest and disease detection system based on cross-scale semantic alignment; it includes a cross-scale data extraction module, a dual-stream shared encoding module, a cross-view attention alignment module, a decoder module, an instance matching module, a contrastive learning optimization module, and an inference detection module.
[0057] The cross-scale data acquisition module is used to acquire large-scale close-up images and small-scale scene images of the same target instance through a single-point dual-shot strategy, generating cross-scale paired annotation data with pixel-level accuracy to form a training set.
[0058] The dual-stream shared coding module comprises two encoder networks with identical structures and shared weights, which extract features from the large-scale image and the small-scale image respectively, and output large-scale feature maps and small-scale feature maps.
[0059] The cross-view attention alignment module uses the small-scale feature map as the query matrix and the large-scale feature map as the key matrix and value matrix, and calculates the semantically enhanced small-scale aligned feature map through a multi-head cross-attention mechanism.
[0060] The decoder module, based on the Transformer architecture, decodes the semantically enhanced small-scale aligned feature maps and large-scale feature maps respectively, generating the corresponding small-scale object query sets and large-scale object query sets.
[0061] The instance matching module uses the Hungarian algorithm to match object queries with real target labels, filtering out cross-scale positive sample object query pairs and negative sample object queries that correspond to the same physical instance.
[0062] The contrastive learning optimization module calculates instance-level contrastive loss in the semantic feature space and performs end-to-end training in conjunction with object detection loss to update network parameters.
[0063] The inference detection module takes only small-scale images as input during the inference stage. After processing by the encoder, cross-view attention module, and decoder, it outputs the target detection results.
Claims
1. A small target detection method based on cross-scale semantic alignment, characterized in that, include: Obtain large-scale close-up images and small-scale scene images of the same target instance, generate cross-scale paired annotation data with pixel-level accuracy, and form a training set; Two encoder networks with identical structures and shared weights are constructed. Large-scale images and small-scale images from the training set are input into the encoder networks respectively, and large-scale feature maps and small-scale feature maps are output. A cross-view attention module is constructed, using the small-scale feature map as the query matrix and the large-scale feature map as the key matrix and value matrix, and a semantically enhanced small-scale aligned feature map is generated through a multi-head cross-attention mechanism. A decoder network based on the Transformer architecture is constructed. The semantically enhanced small-scale aligned feature maps and large-scale feature maps are input into the decoder to generate an object query set. The object query is matched with the real target annotation through the Hungarian algorithm to construct cross-scale positive and negative sample pairs. The cross-scale contrast loss is calculated based on the positive and negative sample pairs. The object detection loss function and the cross-scale contrast loss function are jointly optimized, and the parameters of the encoder network, cross-view attention module and decoder network are updated through backpropagation algorithm until the model converges.
2. The small target detection method based on cross-scale semantic alignment according to claim 1, characterized in that, The process of generating cross-scale paired annotation data with pixel-level accuracy involves manually annotating the large-scale close-up image to obtain the target bounding box, and then automatically mapping the annotation information of the large-scale image to the small-scale image through feature point matching and homography matrix transformation.
3. The small target detection method based on cross-scale semantic alignment according to claim 1, characterized in that, The encoder network adopts a visual transformer architecture, which extracts scale-invariant visual features through shared multi-head self-attention weight parameters.
4. The small target detection method based on cross-scale semantic alignment according to claim 1, characterized in that, The process of matching object queries with real target labels using the Hungarian algorithm includes: constructing a cost matrix, where the matching cost between each prediction and the real target consists of a classification cost and a bounding box cost; using the Hungarian algorithm to solve for the optimal bipartite graph matching between the prediction set and the real target set; and based on the matching results, marking prediction queries that successfully match the real target as valid queries and marking unmatched prediction queries as negative samples.
5. A small target detection method based on cross-scale semantic alignment according to claim 1, characterized in that, The cross-scale positive and negative sample pairs include: for the same physical instance that exists simultaneously in both large-scale and small-scale images, the matched query ( , A query for a single instance in the same image or a query for instances in different images constitutes a positive sample pair across scales; while a query for instances in different images constitutes a negative sample pair.
6. The small target detection method based on cross-scale semantic alignment according to claim 1, characterized in that, The cross-scale contrastive loss adopts the InfoNCE loss function, which applies instance-level cross-scale contrastive loss to the semantic feature space of the object query embedding. It also introduces a momentum queue mechanism to sample negative samples from historical batches, thereby improving the diversity of negative samples.
7. A small target detection method based on cross-scale semantic alignment according to claim 1, characterized in that, A total loss function is constructed by combining the target detection loss and the cross-scale contrast loss.
8. A small target detection method based on cross-scale semantic alignment according to claim 1, characterized in that, The model convergence includes: the validation set performance index no longer improving after multiple consecutive rounds, or the model training rounds reaching a preset value, and the current model weights being saved as the final training result.
9. A small target detection method based on cross-scale semantic alignment according to claim 1, characterized in that, The method further includes, during the inference phase, only a small-scale image to be detected is input, and after processing by an encoder network, a cross-view attention module and a decoder network, the target detection result is output, including the target category and bounding box coordinates.
10. A small target detection system based on cross-scale semantic alignment, characterized in that, It includes a cross-scale data extraction module, a two-stream shared encoding module, a cross-view attention alignment module, a decoder module, an instance matching module, a contrastive learning optimization module, and an inference detection module; The cross-scale data acquisition module is used to acquire large-scale close-up images and small-scale scene images of the same target instance through a single-point dual-shot strategy, and generate cross-scale paired annotation data with pixel-level accuracy to form a training set. The dual-stream shared coding module contains two encoder networks with identical structures and shared weights, which extract features from the large-scale image and the small-scale image respectively, and output large-scale feature maps and small-scale feature maps. The cross-view attention alignment module uses the small-scale feature map as the query matrix and the large-scale feature map as the key matrix and value matrix, and calculates through a multi-head cross-attention mechanism to generate a semantically enhanced small-scale aligned feature map. The decoder module, based on the Transformer architecture, decodes the small-scale aligned feature maps and large-scale feature maps of semantic enhancement respectively, generating the corresponding small-scale object query sets and large-scale object query sets. The instance matching module uses the Hungarian algorithm to match object queries with real target annotations, and filters out cross-scale positive sample object query pairs and negative sample object queries that correspond to the same physical instance. The contrastive learning optimization module calculates instance-level contrastive loss in the semantic feature space and performs end-to-end training in conjunction with object detection loss to update network parameters; The inference detection module takes only small-scale images as input during the inference stage. After processing by the encoder, cross-view attention module, and decoder, it outputs the target detection results.