Cross-modal learning methods for motion features in small object detection

By employing a cross-modal learning method combining a dynamic-static aggregation network and a motion pattern mining module, the problem of insufficient utilization of inter-frame features in small target detection is solved, achieving efficient and accurate small target detection, which is applicable to target monitoring in military satellite remote sensing and national defense fields.

CN117095452BActive Publication Date: 2025-11-14UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310810022.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-03
Publication Date
2025-11-14
Estimated Expiration
2043-07-03

AI Technical Summary

Technical Problem

Existing small target detection methods do not make sufficient use of inter-frame features, which limits their detection performance. Furthermore, multimodal schemes have high computational costs, long inference times, and require specific modal data, making them ineffective for detecting small targets.

Method used

We employ a motion feature cross-modal learning method, which combines a dynamic-static aggregation network and a motion-inspired cross-modal learning multi-layer recurrent architecture with a motion pattern mining module and a motion-visual adapter to fuse features from multiple frames of images for small object detection.

Benefits of technology

It improves the accuracy and adaptability of small target detection, reduces inference costs, and enables advanced detection performance in simple detectors, making it suitable for target monitoring in military satellite remote sensing and defense fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117095452B_ABST
    Figure CN117095452B_ABST
Patent Text Reader

Abstract

This invention discloses a cross-modal learning method for motion features in small object detection, applied to the field of small object detection technology, addressing the problem of poor performance in existing small object detection technologies. The method of this invention includes two key parts: motion pattern mining and a motion visual adapter. The former mines motion patterns from a time-varying visual representation space, while the latter is designed to associate motion patterns with visual semantics. Then, their cross-modal interactions are explored to guide the motion extractor to effectively capture the implicit motion modalities. With the cooperation of motion modalities, even using a simple detector, state-of-the-art detection performance can be achieved with almost no increase in the inference cost of small moving object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing, and specifically relates to a small target detection technique. Background Technology

[0002] Due to the shooting distance and camera mode, many real-world targets, even those that are actually large, often appear visually small in images or videos, such as vehicles in remote sensing videos and aircraft in infrared detection images. Because of their small size in the background image, these small targets may lack distinct color, shape, and texture, and may even blend into the background with low contrast.

[0003] Given the aforementioned characteristics, accurately detecting the location and bounding boxes of small objects from images or videos is typically a highly challenging problem. In small object detection tasks, detectors are usually trained on visual modality datasets to learn visual feature representations of small objects. Generally, these small object detectors face two inherent challenges. First, the training samples may contain only a small number of labeled objects, limiting the feature extractor's ability to learn sufficient and effective object features. Second, blurry small objects, represented by only a few visual features, may fail to be recognized by the object detector.

[0004] To address the problem of small object detection, numerous detection methods have been developed over the decades. Most previous methods followed popular unimodal frameworks (e.g., vision) or multimodal (e.g., vision-infrared) feature extraction models. From a training and inference perspective, these works can be divided into three groups: (a) unimodal I, (b) unimodal II, and (c) multimodal.

[0005] Single-modal I and multimodal approaches have obvious advantages such as simplicity. However, besides visual features, they often neglect inter-frame features of small targets. Therefore, their performance is usually limited due to the lack of inter-frame features. An example of single-modal I, YOLOX, only captures and focuses on the visual features of a single frame, resulting in poor performance in small target detection. Therefore, most current state-of-the-art models adopt the single-modal II approach.

[0006] To overcome the weaknesses of unimodal I approaches, unimodal II approaches are adopted to extract inter-frame visual features of small targets from image sequences, such as DSFNet, B-MCMD, and MotionRNN. However, these widely adopted approaches also have two significant limitations. First, their complex models can impose substantial computational and memory costs on the inference stage. Second, these approaches only play an auxiliary role, providing inter-frame visual features to the detector rather than "teaching" the feature extractor to capture motion features to detect small targets.

[0007] Furthermore, multimodal approaches typically use RGB and infrared images for multimodal learning, enabling the detector to learn richer small target features than single-modal I and single-modal II methods, such as UA-CMDet. While achieving good results, these are essentially still single-frame-based detection methods, neglecting inter-frame features. In practical applications, they often require long inference times and a large number of model parameters. Moreover, they require data from both modalities to function effectively. Summary of the Invention

[0008] To address the aforementioned technical problems, this invention proposes a cross-modal learning method for motion features of small targets, which enables more accurate detection of moving small targets.

[0009] The technical solution adopted in this invention is: a cross-modal learning method for motion features of small targets, and the detection system based on it includes: a motion-inspired cross-modal learning multi-layer recurrent architecture, a dynamic-static aggregation network, and a detector;

[0010] An image sequence is obtained by using multiple consecutive frames of images acquired by a camera;

[0011] The dynamic-static aggregation network performs primary visual feature extraction on the input image sequence through the third branch of the dynamic flow structure, resulting in feature map X. s ;

[0012] The motion-inspired cross-modal learning multi-layer recurrent architecture consists of multiple NODEs (l, t), where l represents the number of NODE layers and t represents the time step; the input of each NODE (l, t) layer is the output of the dynamic-static aggregation network corresponding to time step t.

[0013] Each layer of NODE(l,t) specifically includes a motion pattern mining module and a motion-visual adaptation module; the motion pattern mining module mines local motion patterns P based on the output of the dynamic-static aggregation network corresponding to time step t. t The motion-visual adaptation module adapts to local motion patterns P. t Semantic bias and modality consistency adaptation are performed to obtain the hidden state of NODE(l,t).

[0014] The final motion modal feature H is obtained by concatenating the hidden states of all layers NODE(l,t) into a matrix. s ;

[0015] Feature map X s With motion modal features H s As supervisory information for training detectors;

[0016] The feature map X extracted by the dynamic and static aggregation network sAs input to the trained detector, the small object detection results are obtained.

[0017] The beneficial effects of this invention are as follows: This invention mainly comprises two key components: motion pattern mining and a motion visual adapter. The former aims to mine motion patterns from a time-varying visual representation space. The latter is designed to associate motion patterns with visual semantics. Then, their cross-modal interactions are explored to guide the motion pattern mining module in effectively capturing implicit motion modalities. With the cooperation of motion modalities, even using a simple detector, state-of-the-art detection performance can be achieved with almost no increase in inference cost for moving small target detection. Secondly, this scheme has good adaptability and performance advantages, and can be applied to two small target-related tasks. The method of this invention makes moving small target detection more accurate, has the ability to discover unlearned small targets, and can be applied to military satellites and defense fields, such as remote sensing detection via satellite, monitoring of military targets, and all-time battlefield awareness. Attached Figure Description

[0018] Figure 1 This is a flowchart of a cross-modal learning method for motion features in small target detection according to the present invention.

[0019] Figure 2 This is a diagram of a small target detection framework based on motion-inspired cross-modal learning, as proposed in this invention.

[0020] Figure 3 This invention provides a comparison between MICML proposed in this embodiment and LSTM-based methods.

[0021] Among them, (a) is a comparison of the performance curves of the present invention and the F1 and IoU threshold based on the LSTM method, and (b) is a comparison of the F1, time cost and parameter size achieved by the present invention and the LSTM method in the detection task. Figure 4 This is a visualization of the detection box in a real-world scenario in an embodiment of the present invention.

[0022] Figure 5 These are visualizations of MPM and MVA and motion modal visualizations in embodiments of the present invention;

[0023] Among them, (a) is the ground view of four satellite video scenes, (b) is the detection result using MPM, (c) is the detection result using MPM & MVA, (d) is the 3D visualization corresponding to (b), and (e) is the 3D visualization corresponding to (c).

[0024] Figure 6 This is a feature heatmap of the model in this embodiment of the invention;

[0025] Among them, (a), (b), (c), (d), and (e) are feature heatmaps of five different scenarios. Detailed Implementation

[0026] To facilitate understanding of the technical content of this invention by those skilled in the art, the following description, in conjunction with the accompanying drawings, further illustrates the invention.

[0027] like Figure 1 The flowchart shown is for a cross-modal learning method of motion features for small target detection according to the present invention. The specific steps are as follows:

[0028] S1. A motion-inspired cross-modal learning scheme for small object detection is proposed.

[0029] S2. To further propose a motion pattern mining module based on a long short-term memory network for the scheme proposed in step S1;

[0030] S3. A motion-vision adapter is further proposed for the module proposed in step S2 to adaptively interact with cross-modal features;

[0031] S4. Use the feature extraction network from step S1 to perform primary feature extraction on the sequence of moving small targets (multiple consecutive images obtained by a camera);

[0032] S5. Motion pattern mining is performed on the sequence features extracted in step S1 using the module proposed in step S2.

[0033] S6. Using the module proposed in step S3, the motion mode features output in step S2 are subjected to motion-visual cross-modal feature adaptation, and the motion modal features are obtained after all loops.

[0034] S7. The motion modal features output in step S6 are fused with the visual modal features output in step S4 to jointly supervise the detector during the training phase.

[0035] This invention designs a detection network framework based on global contrast learning, the complete structure of which is as follows: Figure 2 As shown, the main structural components include:

[0036] (1) Small target detection framework;

[0037] This framework comprises a proposed Motion-Inspired Cross-Modal Learning (MICML) multi-layer recurrent architecture and a detector based on a Dynamic and Static Fusion Network (DSFNet). Motion modal features acquired through MICML and visual modal cross-modal features acquired through DSFNet are used together as supervision information for training the detector. Finally, the detector trained with the cross-modal features performs the detection of small moving targets. The multi-layer recurrent architecture consists of NODE(l,t), which sets the output of the last NODE layer... As a motion mode.

[0038] (2) Motion Pattern Mining (MPM) module;

[0039] This invention proposes a motion pattern mining module based on a Long Short-Term Memory (LSTM) network, comprising one cell, T nodes (l, t), and a channel attention module. The computation of a cell is represented as follows:

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046] Where l represents the cell or node level, and H represents the hidden state. The hidden state of the Cell output is represented by C, where C represents the cell state, σ is the sigmoid activation function, and * and Let represent the convolution operation and the Hadamard product, respectively, and tanh be the hyperbolic tangent function. W and b represent different weight matrices of the convolution (corresponding to subscripts xi, hi, xf, hf, xo, ho, xg, hg) and different biases, respectively. When l = 1... For X t That is, the visual feature matrix X at time step t. s ={X1, X2, ..., X t , ..., X TFurthermore, the calculation method for node(l,t) is the same as that for Cell above; the input of node(l,t) is the output of Cell, i.e. By using formulas (1)-(4) Replace with This will give you the hidden states for all time steps.

[0047] Furthermore, the model uncovers past and future motion patterns, specifically when t→(t-1), This represents the backward motion flow, and when t→(t+1), This represents the forward motion flow. Ultimately, a set of local motion patterns based on the forward or reverse motion flow can be calculated as follows:

[0048]

[0049]

[0050]

[0051]

[0052] Among them, P t This represents the local motion pattern at reference time t. To be The results are divided into c parts. concat (·) denotes the matrix concatenation operation, f maxpool (·) indicates max pooling. This represents channel-level multiplication. Channel attention is used to compute α to optimize the dependencies between local relational feature channels. α is the ReLU function, and σ is the Sigmoid function. z represents the features after global average pooling, z = [z1, z2, ..., z...]. c ] T , z c This represents the spatial features encoded on a single channel.

[0053] (3) Motion-Vision Adapter (MVA);

[0054] Motion pattern mining yields potential motion patterns containing local motion relationships across two motion flows. However, this simplistic approach to motion patterns may have two drawbacks. To overcome the semantic bias and modal inconsistency issues that may exist in motion visual features, this invention designs a motion-vision adapter.

[0055] The motion-vision adapter involves three implementation steps:

[0056] Step 1, Projection. Use max pooling to... The size is downsampled from w×h to c×c, and obtained through channel attention. For mode P t and Their features reside in different semantic spaces. To enable interaction between the two modalities, the modal representations are projected onto a new common semantic space. A linear projection matrix for modality association is defined for each modality. Specifically:

[0057] R t =Q t U t (11)

[0058]

[0059]

[0060] Among them, Q t and U t It is a projected representation located in a new space. and Let b represent two projection matrices respectively. Q and b H These represent two bias terms. The motion-visual relationship R is then learned through parameter learning. t .

[0061] Step 2, Aggregation. A motion-visual relationship-guided aggregation method is proposed to help the Motion Pattern Mining (MPM) module focus on motion-aware regions on the visual feature map and learn motion features corresponding to visual attributes, as described below:

[0062] Z t =f mean (R t ||N t (14)

[0063]

[0064] Among them, Z t It is a visually dependent motion feature that encodes the average motion-visual distribution of channels in the latent space, f mean (·) denotes the element-wise average function of the matrix. N t It is a feature of visual association, through Projecting onto a new semantic space yields N. t with U t All passed Projections obtained in different subspaces are features that exist in different semantic spaces. and b N These represent the projection matrix and the bias term, respectively.

[0065] Step 3, Adaptation. To alleviate the semantic adaptation gap between motion and vision, a two-level weighted strategy is proposed for features from both modalities. Therefore, the hidden state of NODE(l,t) (l-th layer, t-th time step) can be computed as follows:

[0066]

[0067] in, It is Z t Visually dependent motion features obtained through upsampling. and These are two linear weight matrices. α and σ are the ReLU and Sigmoid functions, respectively. It includes local motion patterns, so global motion patterns can be obtained at all time steps.

[0068] The input characteristics of the motion-vision adapter (MVA) are After passing through the adapter, the hidden state output by NODE(l,t) is: The hidden state set is obtained through the NODE at all time steps, and then the matrix is ​​concatenated to obtain the final motion modality features used for learning by the supervised detector. H through model s Learning can help the detector better capture the motion features of the target, making the detection of small targets more accurate.

[0069] The process by which DSFNet acquires visual modal features is as follows:

[0070] In DSFNet, the third branch of the dynamic flow structure performs primary visual feature extraction on the sequence of moving small targets (the sequence of moving small targets is the image sequence of the entire frame input, obtained by a camera or other monitoring equipment), resulting in a feature map of size T×c×w×h. in, Let represent the set of real numbers, c be the number of channels, T be the sequence length, w be the feature map width, and h be the feature map height. Then, this visual feature map is input into the motion pattern mining module.

[0071] The process of MICML acquiring motion modal features is as follows:

[0072] The input features of the motion pattern mining module are The output motion mode features are

[0073] In this invention, motion modal features acquired by MICML and visual modal features acquired by DSFNet are fused using direct addition. During the training phase, the fused features are used to supervise the detector (the supervision information during training is: and During the inference phase, only visual modal features are used (the information used during the inference phase is: The training employed three loss functions, one of which was the keypoint loss. (Heatmap), used to optimize the loss of bounding boxes. (Box size) and for offset The center offset is designed as a standard L1 loss. The overall loss function is defined as follows:

[0074]

[0075] Where λ represents the hyperparameter, λ size =0.1, λ off =1.

[0076] The following uses the Satellite videos and DroneCrowd datasets for small moving targets as examples to further illustrate the present invention. The Satellite videos training set contains 27,434 frames, and the test set contains 2,255 frames. Moving vehicles in the dataset are selected as small targets. DroneCrowd has a total of 33,600 frames, with 82 sequences in the training set and 30 sequences in the test set.

[0077] (1) Experiment initialization:

[0078] In the comparative experiments, the DSFNet detector was used as the baseline model. In this embodiment, the Adam optimizer was used to train both the baseline and MICML for 55 epochs, with a batch size of 2. The initial learning rate was 1.25 × 10⁻⁶. -4 The degradation rate decreases by 10% after generations 30 and 45. MICML only collaborates with the baseline detector during the training phase and is not used during the inference phase. To ensure fair comparison with other methods, the length of the input sequence is set to 5. The input frame resolution for model training is 512×512, and the input frame resolution for inference is 1024×1024.

[0079] In the tracking experiments, the STNNet tracker was used as the baseline. Following the STNNet protocol, the tracking model was trained for 100 epochs with a batch size of 4. The input frame resolution was 960×540 for both training and testing. The learning rate was set to a constant 1×10⁻⁶. -6 The total number of candidate small targets is set to 128. To ensure a fair comparison with other algorithms, the length of the input sequence is set to 2.

[0080] (2) Cross-modal learning stage:

[0081] In this embodiment, MICML uses a set of visual features from multiple frames of images. As input, features are obtained from a visual feature extractor, such as Figure 2 As shown, NODE(l,t) is magnified into motion pattern mining and motion-visual adapter, representing the computational component of the l-th layer at time step t-th in the motion feature extractor of the L-layer recurrent network. The motion pattern mining component aims to mine motion patterns, and the motion-visual adapter aims to bridge the semantic gap between the two modalities. The motion feature set H obtained by the motion extractor is... s With X s The data is then fused and transmitted to the small target detector to calculate the loss.

[0082] (3) Loss function:

[0083] In this embodiment, three loss functions are used: keypoint loss, bounding box regression loss, and offset loss. The loss is calculated using Equation (17), and the model is effectively learned through backpropagation gradient descent, ultimately achieving the result of effectively detecting small moving targets.

[0084] To verify the effectiveness of this invention in detecting small moving targets, it was trained on the publicly available dataset Satellite videos. Furthermore, this invention was compared with some state-of-the-art models (E-LSD, D&T, B-MCMD, DSFNet, etc.), as shown in Table 1.

[0085] Table 1 Comparison results of the present invention with some advanced models

[0086]

[0087] The main comparisons of the experiments are shown in Table 1. Two significant findings can be observed from Table 1. First, MICML can refresh the state-of-the-art peak performance across all three average metrics. Second, in terms of time cost, the proposed method has a significant advantage over most traditional nonparametric models. For example, the minimum time cost for detection inference achieved through D&T is 0.18 seconds (the fastest). Meanwhile, the cost of the present invention is 0.20 seconds, slightly higher than the former. However, the time costs of most comparisons even exceed 1.0 second, such as B-MCMD's 45.7 seconds and Vibe's 1.3 seconds; these methods are significantly more time-consuming than the present invention.

[0088] A comparison of a set of average accuracy metrics is shown in Table 2. Two points can be observed through this comparison. First, the moving target method consistently outperforms the stationary target detection method by a significant margin. The AP of the stationary target method... 50 The highest AP was only 48.9% (obtained by CenterNet), while the AP of the moving target method was... 50 The lowest was 69.2% (obtained from DSFNet). The latter was approximately 20.3% higher than the former.

[0089] Table 2 Comparison of a set of average accuracy indicators

[0090]

[0091] In addition, two sets of comparative experiments were designed for popular LSTM-based methods, and the comparison results are as follows: Figure 3 As shown. Through observation, in Figure 3 In (a), it can be observed that the MICML proposed in this invention consistently achieves the highest average F1 value for any IoU threshold from 0 to 0.5. The performance curves of F1 versus IoU thresholds in this invention are significantly higher than other curves. Furthermore, in Figure 3 In (b), it can be observed that the sphere of the present invention is located in the upper left corner, and its size is almost the smallest. Conversely, other spheres are always below and to the right of the sphere of the present invention. This comparison shows that the present invention achieves the best overall performance in the detection task in terms of F1, time cost, and parameter size, far superior to other models.

[0092] like Figure 4 As shown, in order to intuitively evaluate the detection results of different methods, four sets of visualization comparisons were made for some representative methods, such as DSFNet (an advanced method for detecting small moving objects), SA-ConvLSTM and MotionRNN (an advanced method based on LSTM) and YOLOX (a well-known general object detection method).

[0093] Through visual comparison, two obvious findings can be observed. First, this invention can typically and accurately detect small moving targets, while other methods often produce missed detections and false detections. For example, in the first row, the labeled image contains 6 small targets. This invention can detect all of them completely. However, SA-ConvLSTM only detects 5, missing 1. Furthermore, DSFNet detects 7, falsely detecting 1. Second, this invention can typically detect small targets that have not been learned, while most other methods typically cannot. For example, in the second row, although there is an aircraft not labeled in the test image, this invention can detect it. In contrast, the other four methods almost all fail to detect it. From the perspective of detection visualization, the above intuitively verifies the numerical experimental results in Table 1. This shows that the detection performance of this invention for small moving targets is superior to other algorithms.

[0094] To verify the adaptability of this invention to small target tracking tasks, it was compared with several popular methods, including MCNN, CAN, CSRNet, DM-Count, and STNNet. Table 3 shows a comparison of the quantitative results of the six methods on DroneCrowd.

[0095] These results demonstrate that the present invention significantly outperforms the other five methods in tracking performance across all four metrics. For example, it is readily apparent that the present invention achieves superior performance in AP. 10 AP 15 and AP 20 The AP ratios achieved were 36.17%, 34.03%, and 27.86%, respectively. In comparison, the suboptimal STNNet performed poorly in AP. 10 AP 15 and AP 20 The accuracy rates on the above methods were only 34.82%, 33%, and 26.92%, respectively. Furthermore, this invention achieved an accuracy of 32.69% on mAP, which is 1.11% higher than STNNet's 31.58%.

[0096] Table 3 Comparison of quantitative results of the method of the present invention with six existing methods on DroneCrowd.

[0097]

[0098] As shown in Table 4, two sets of ablation experiments were conducted to evaluate the effectiveness of the Motion Pattern Mining (MPM) component and the Motion-Vision Adapter (MVA) in this invention.

[0099] The comparison shows that both components are consistently effective in helping the baseline model improve detection and tracking performance. For example, on Satellite videos (detection task), the baseline's F1 score and AP... 50 The values ​​were 0.83% and 69.2%, respectively. With the help of MPM, F1 and AP... 50 These figures were improved to 0.85 and 71.5%, respectively. Once the motion MPM and MVA were incorporated into the baseline, the F1 and AP50 could be continuously refreshed, reaching 0.88 and 73.1%, respectively.

[0100] Furthermore, we found that MPM and MVA typically have different impacts on model performance. Experimental results show that MPM can help improve the baseline model's mAP on DroneCrowd (tracking task) from 31.58% to 31.94%, an improvement of 0.36%. Similarly, MVA can continuously improve the mAP of the tracking method from 31.94% to 32.69%, a net increase of 0.75%. This indicates that the net increase generated by the MVA component is significantly higher than that of MPM.

[0101] Table 4. Evaluation of the effectiveness of the motion pattern mining component and motion-visual adapter in this invention.

[0102]

[0103] like Figure 5 As shown, to demonstrate the roles of MPM and MVA in implicit motion modes, four visual comparison examples are presented, corresponding to four satellite video scenarios.

[0104] By comparison, it was found that MPM is effective, although it may introduce semantic bias, leading to false positives. Combining MPM with MVA can effectively mitigate potential semantic bias. For example, in the third case (moving waves, third row), some false positives generated by MPM are mainly due to semantic bias, while MVA can help MPM suppress this semantic bias. Figure 5 The 3D visualization of the motion modes in the third row of (d) and (e) precisely verifies this observation.

[0105] Furthermore, the results revealed that, when combined with MVA, MPM can effectively detect small moving targets, even those not labeled in the test set. For example, in the second example (moving vehicle, second row), MPM only detected one moving object. However, when integrated with MVA, it detected other moving targets besides those labeled as ground truth on the ground. This additional target was not a false detection, but a genuine small moving target in the video, even though it wasn't considered in the test set. One of the most compelling explanations is that MVA helps MPM effectively reduce the semantic gap between the motion-visual modalities, promoting semantic consistency. Figure 5 The results in (d) and (e) in the second row also confirm this observation.

[0106] To intuitively assess the accuracy of the focus location, in Figure 6 Five sets of feature heatmaps are listed. By comparing with the labeled images, it is easy to see on all heatmaps that the focus areas captured by the feature maps of this invention (in full) are more precise than those captured by the baseline (excluding MPM and MVA). This indicates that both MPM and MVA can effectively help the detector accurately focus on small moving targets.

[0107] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the scope of the claims of the invention.

Claims

1. A cross-modal learning method for motion features of small targets, characterized in that, The detection systems based on this include: motion-inspired cross-modal learning multi-layer recurrent architecture, dynamic-static aggregation network, and detectors; An image sequence is obtained by using multiple consecutive frames of images acquired by a camera; The dynamic-static aggregation network performs primary visual feature extraction on the input image sequence through the third branch of the dynamic flow structure, resulting in feature map X. s ; The motion-inspired cross-modal learning multi-layer recurrent architecture consists of multiple layers of NODE(l,t), where l represents the number of NODE layers and t represents the time step; the input of each NODE(l,t) layer is the output of the dynamic-static aggregation network corresponding to time step t. Each layer of NODE(l,t) specifically includes a motion pattern mining module and a motion-visual adaptation module; the motion pattern mining module mines local motion patterns P based on the output of the dynamic-static aggregation network corresponding to time step t. t The motion-visual adaptation module adapts to local motion patterns P. t Semantic bias and modality consistency adaptation are performed to obtain the hidden state of NODE(l,t). The final motion modal feature H is obtained by concatenating the hidden states of all layers NODE(l,t) into a matrix. s ; Feature map X s With motion modal features H s As supervisory information for training detectors; The feature map X extracted by the dynamic and static aggregation network s As input to the trained detector, the small object detection results are obtained.

2. The cross-modal learning method for motion features of small targets according to claim 1, characterized in that, The motion pattern mining module employs a Long Short-Term Memory (LSTM) network, consisting of one Cell, T nodes (l,t), and a channel attention module. The input to the Cell is the output of the previous layer's node (l,t), and the output of the Cell is the hidden state. The input to each node(l,t) is the output of the Cell and the output of the dynamic-static aggregation network corresponding to time step t. The output of the motion pattern mining module is obtained by concatenating the outputs of T nodes(l,t) and the calculation results of the attention module.

3. The cross-modal learning method for motion features of small targets according to claim 2, characterized in that, The expression for the output of the motion pattern mining module is: Among them, P t This represents the local motion pattern at time step t, i.e., the output of the motion pattern mining module. f is the result obtained by concatenating the outputs of T node(l,t). concat (·) denotes the matrix concatenation operation, f maxpool (·) indicates max pooling. This represents channel-level multiplication, where 'a' is the result of the channel attention module, α is the ReLU function, σ is the Sigmoid function, and z is the feature after global average pooling: z = [z1, z2, ..., z2]. c ] T , z c This represents the spatial features encoded on a single channel, where c represents the number of channels.

4. The cross-modal learning method for motion features of small targets according to claim 3, characterized in that, The output of each node(l,t) is either a forward motion flow or a reverse motion flow.

5. The cross-modal learning method for motion features of small targets according to claim 4, characterized in that, The implementation process of the motion-visual adaptation module is as follows: A1. Projection: The output of the Cell is projected using max pooling. Downsampling to P t Consistent dimensions, resulting in and P t and Projected onto the first common semantic space; R t =Q t U t Among them, Q t and U t It is the projected representation located in the first public semantic space. and Let b represent two projection matrices respectively. Q and b H These represent two bias terms, R. t For motion-visual relationships; A2. Aggregation, specifically: (same as above) Projecting onto the second semantic space yields the visually associated feature N. t Based on the obtained visually dependent motion characteristics; Z t =f mean (R t ||N t ) Among them, Z t It is a visually dependent motion feature, f mean (·) denotes the element-wise average function of the matrix, N t It is a feature of visual association. and b N These represent the projection matrix and the bias term, respectively. A3. Adaptation, specific: according to Z t With Cell output Obtain the hidden state of NODE(l,t). in, It is Z t Visually dependent motion features obtained through upsampling and These are two linear weight matrices, where α and σ are the ReLU and Sigmoid functions, respectively.

Citation Information

Patent Citations

  • Recognition method for finding dynamic target in air in real time

    CN113157800A

  • Multimodal named entity identification method based on entity-level cross-modal interaction

    CN115796182A