Infrared small target detection network with same normal form feature extraction and ODE guide feature fusion

By adopting the same-paradigm feature extraction and ODE-guided feature fusion method in the infrared small object detection network, combined with the CNN, Transformer and Mamba network architecture, the problem of insufficient feature diversity and adaptability in the existing technology is solved, and a more efficient and robust infrared small object detection effect is achieved.

CN120198757AActive Publication Date: 2025-06-24ZHONGBEI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510263728.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The existing infrared small-object detection technology has shortcomings in terms of feature diversity and adaptability, making it difficult to effectively identify and locate small targets, especially in environments with complex backgrounds and severe noise.

Method used

An infrared small object detection network with the same paradigm feature extraction and ODE-guided feature fusion is adopted, combined with the three-branch network architecture of CNN, Transformer and Mamba, local features are extracted through convolutional layers, Transformer extracts global context information, Mamba dynamically adjusts the state space model to eliminate redundant information, and designs a feature fusion module guided by ordinary differential equations to integrate the visual features of different coding strategies.

Benefits of technology

It significantly improves the accuracy and robustness of infrared small object detection, enhances the model's perception of subtleties, and can more effectively deal with background interference and retain target details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198757A_ABST
    Figure CN120198757A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared small target detection network based on same normal form feature extraction and ODE guide feature fusion, and belongs to the technical field of infrared small target detection. Aiming at the problem that a traditional network is insufficient in feature diversity and adaptability, firstly, a three-branch architecture including CNN, Transform and Mama is adopted, wherein the CNN extracts local features, the Transform obtains global information, and the Mama filters out irrelevant information by means of a unique selection mechanism and retains important features of key visual clues in an image; secondly, a feature fusion module inspired by ODE is designed to serve as an information bottleneck to suppress high-frequency noise, and meanwhile target features are enhanced through back-propagation gradient; secondly, abstracting a universal framework for Transform and Mamba, adding a multi-layer perceptron to enhance the nonlinear problem processing capability, and introducing a gating mechanism to dynamically adjust information flow, suppress noise and irrelevant background information; and finally, performing an ablation experiment and a contrast experiment on the public and available SIRST data set, and verifying the effectiveness of the OFSPNet in the aspect of improving the detection performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of infrared small target detection, and specifically relates to an infrared small target detection network for extracting homogeneous paradigm features and fusing ODE-guided features. Background Art

[0002] Infrared small target detection is a key technology for identifying and locating tiny targets in infrared images, and has important application values in the fields of military, security, medical treatment, etc. However, infrared small targets may be extremely small, less than a few pixels, which makes it difficult to identify their shape and texture features. Moreover, infrared signals are prone to attenuation during long-distance transmission, resulting in a decrease in the brightness of the target in the image and making it difficult to distinguish from the background. These characteristics make the infrared small target detection task extremely challenging. Therefore, improving the detection rate of the infrared small target detection task is still an urgent need in practical applications.

[0003] Traditional infrared small target detection methods are model-driven methods, such as methods based on filtering, simulation of the human visual system, and low-rank sparse matrices. The filter-based method highlights small targets by differencing the original image and the filtered background image, but when the background is complex, the detection performance of small targets decreases and the robustness is poor; the human visual system-based method uses a saliency map to distinguish small targets according to the difference between the local range of the target and the background, and is mainly applicable to scenarios where the target brightness is large and different from the surrounding background; the low-rank sparse matrix-based method uses the low-rank characteristics of the background and the sparse characteristics of the target to improve the detection accuracy, but some strong clutter signals are as sparse as the target signals, resulting in a high false alarm rate. Traditional methods generally face problems such as noise sensitivity, parameter dependence, and poor adaptability.

[0004] Benefiting from the development of computer vision in many applications, convolutional neural networks (CNNs) and vision transformers (ViTs) have been proven to be able to effectively handle various visual tasks. The convolutional operation of CNNs can effectively extract local features in images, but the perception range is limited and it is difficult to obtain global information. Different from the standard CNN-based method that processes images pixel by pixel, the vision transformer (ViT) regards an image as a series of patch tokens. At each layer of the network, ViT uses a multi-head self-attention mechanism to process patch tokens according to the relationship between each pair of tokens, so as to be able to construct a global representation of the entire image, but it is insufficient in local feature extraction and detail preservation, and it is difficult to accurately identify small targets with fuzzy texture features.

[0005] In addition, with the emergence of the Mamba model, the architecture based on the Selective State Space Model (SSM) has attracted extensive attention in the field of vision, providing new ideas for infrared small target detection. Mamba enhances the model's ability to model long-term dependencies through a cleverly designed state space structure. The core operations of Mamba can be highly parallelized, enabling Mamba to be more efficient when processing large-scale data. However, Mamba needs to flatten spatial data into one-dimensional tokens, which destroys the natural local two-dimensional dependencies and weakens the model's ability to accurately interpret spatial relationships. Although VMamba introduces a 2D scanning technique to address this issue by scanning the image in the horizontal and vertical directions, it is still difficult to maintain the proximity of originally adjacent tokens in the scanning sequence, which is crucial for effective local representation modeling.

[0006] Although they each have significant advantages, they have deficiencies in feature diversity and adaptability, which limits their performance in dealing with complex or specific scenarios. And currently, the in-depth research and innovative work on these model structures themselves are relatively limited. Summary of the Invention

[0007] Aiming at the problem of the deficiencies of traditional networks in feature diversity and adaptability, the present invention provides an infrared small target detection network for homogeneous paradigm feature extraction and ODE-guided feature fusion.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] An infrared small target detection network for homogeneous paradigm feature extraction and ODE-guided feature fusion, the deep network adopts a three-branch network architecture of CNN, Transformer, and Mamba, optimizes the structures of Mamba and Transformer, retains the MLP module and the residual module in the Transformer structure, does not specify a specific attention module, introduces a gating module, and uses it as a structural block under the same paradigm; at the same time, designs an ordinary differential equation-guided feature fusion module for integrating visual features from different encoding strategies to enhance the visual representation ability of the network model.

[0010] In this network architecture, the convolutional layer CNN focuses on mining the local feature details of images; the self-attention mechanism in Transformer is responsible for extracting the global context information of images, broadening the model's vision; the unique selection mechanism of the Mamba model allows the model to dynamically adjust the parameters of the state space model (SSM) according to the input data, accurately eliminating redundant information and highlighting the key visual features in the images. This multi-dimensional feature extraction method not only greatly improves the accuracy of the detection task but also enhances the model's perception ability of subtle details. And the module structures of CNN, Transformer, and Mamba are optimized. In this design, a multi-layer perceptron (MLP) is added to enhance the model's ability to handle non-linear problems, enabling it to capture more complex feature relationships. At the same time, a gating mechanism is introduced, enabling the model to dynamically adjust the information flow, effectively retaining key features while suppressing noise and irrelevant background information.

[0011] Furthermore, in the structural blocks under the same paradigm, each structural block includes two residual sub-blocks; the first residual sub-block contains an unspecified module and a gating mechanism; the second residual sub-block consists of a two-layer MLP with non-linear activation. The formula for the structural block is as follows:

[0012]

[0013] where, x i represents the input of the i-th layer, represents the output of the first residual block, σ represents SiLu, Y i represents the unspecified module, and this module is specified as an attention module, an SSM module, and a multi-scale convolutional module. The formulas are as follows:

[0014] Y i =Attetnion(x i ·W Q ,x i ·W K ,x i ·W V ) (3)

[0015] Y i =conv 1×1 (concat(conv 1×1 (x i ),conv 3×3 (x i ),conv 5×5 (x i ))) (4)

[0016]

[0017] Among them, Attention represents the standard attention, and its calculation formula is Q, K, and V represent the query, key, and value matrices; j ∈ {1, 2, 3, 4} represents one of the four scanning directions; z j represents the feature sequence output by the extended scan; It means that then the four resulting feature sequences are processed separately by the S6 block; S6 represents the state space model processing the input feature sequence.

[0018] Furthermore, the designed ordinary differential equation-guided feature fusion module is used to integrate visual features from different coding strategies. This module acts as an information bottleneck, effectively suppressing high-frequency noise, and at the same time strengthening the target features through the gradient of backpropagation. Specifically:

[0019] The ordinary differential equation is represented by an equation with the function y(t) as the variable t and its derivative, as follows:

[0020]

[0021] Among them, f(t, y(t)) defines a time-varying function with discrete time-varying time, t = t0, t1....., t0 represents the initial value of the ODE, and the solution at time t1 is expressed as:

[0022]

[0023] Formula (7) is rewritten as:

[0024] y(t + Δt) = y(t) + Δtf(t, y(t)) (8)

[0025] Among them, Δt = t1 - t0, by defining Δtf(t, y(t)) = hg(y t ), y(t + Δt) = y t+1 and y(t) = y t , formula (8) is mapped to the Resblock in ResNet, and formula (8) is expressed in formula (9) as:

[0026] y t+1 = y i + hg(y t ) (9)

[0027] The fourth-order Runge-Kutta method is used to replace the first-order Euler method to reduce the local truncation error of formula (8). The formula is as follows:

[0028]

[0029] Among them, F1, F2, F3, and F4 represent the slopes calculated at different intermediate points; g(y t ) describes the rate of change of y with respect to time t; h represents the step size.

[0030] Furthermore, the deep network is trained using binary cross-entropy loss and cross-linking loss to improve the network's performance in infrared small target detection. The formula is as follows:

[0031] l = Bce(y,) + Iou(y,) (11)

[0032] Among them, Bce(y,) represents the binary cross-entropy loss, and Iou(y,) represents the cross-linking loss.

[0033] Compared with the prior art, the present invention has the following advantages:

[0034] (1) A three-branch network structure integrating convolutional neural network (CNN), Mamba, and Transformer is proposed to fully extract important local features and global context information. Drawing on the U-Net architecture, skip connections are adopted in the decoding stage to effectively restore image details, thereby improving the detection ability of small targets in infrared images.

[0035] (2) Innovations are made to the structures of Mamba, CNN, and Transformer. The structure integrates multi-layer perceptron (MLP) sub-blocks and gating mechanisms to enhance the model's performance and efficiency in processing large-scale datasets and complex tasks.

[0036] (3) A feature fusion module guided by ordinary differential equations (ODE) is proposed to adaptively aggregate visual features from different coding strategies. Enhance features and suppress noise, effectively handle background interference and retain target details. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic diagram of the structural module. (a) is the Mamba module, (b) is the Transformer module; (c) is the same paradigm structure;

[0038] Figure 2 It is the overall architecture diagram of OFSPNet;

[0039] Figure 3 It is the branch structure diagram. (a) is the Transformer Block; (b) is the MambaBlock; (c) is the ConvBlock;

[0040] Figure 4Diagram of the ODE-inspired fusion module; (a) Runge-Kutta Block; (b) Structure of F1, F2, F3, F4;

[0041] Figure 5 Diagram for comparing ROC with the latest method;

[0042] Figure 6 Visual comparison diagram of detection results on the SIRST-5K dataset;

[0043] Figure 7 Visual comparison diagram of detection results on the SIRST-AUG dataset;

[0044] Figure 8 Visual comparison diagram of detection results on the NUAA dataset;

[0045] Figure 9 Diagram of 3D visualization results of different methods on test images. Detailed implementation manners

[0046] To gain a deep understanding of the present invention, we will describe it comprehensively and meticulously. However, the present invention has multiple implementation manners and is not limited to the specific examples listed herein. The presentation of these examples aims to deepen the comprehensive understanding of the disclosed content of the present invention.

[0047] An infrared small target detection network with the same-paradigm feature extraction and ODE-guided feature fusion. The deep network adopts a three-branch network architecture of CNN, Transformer, and Mamba, and optimizes the structures of Mamba and Transformer. As Figure 1 shown in (c), retain the MLP module and the residual module in the Transformer structure, no longer specify a specific attention module, introduce a gating module, and use it as a structural block under the same paradigm; at the same time, design an ordinary differential equation-guided feature fusion module for integrating visual features from different encoding strategies and enhancing the visual representation ability of the network model.

[0048] To more effectively fuse the features output by each branch and reduce the semantic differences, first perform a channel concatenation operation on the features generated by different branches. The aim is to retain the uniqueness and richness of the features captured by each branch. Subsequently, the concatenated features will be passed to the ODE-inspired feature fusion module. This module plays a key role as an information bottleneck, which can not only effectively filter out high-frequency noise but also enhance the target features with the gradient signal of backpropagation, thereby improving the quality and representativeness of the features.

[0049] Analyze the structures of the mamba and Transformer modules, asFigure 1 (b) The structure of the Transformer encoder mainly consists of two parts: the first part is the attention module; the second part includes other key modules such as a multi-layer perceptron (MLP) and residual connections. Different from the Transformer, Mamba designs its structure by combining two basic designs, H3 and gated attention, thus forming the Figure 1 architecture shown in (a). The present invention optimizes the structural design of the Mamba and Transformer blocks, and abstracts a structural block under the same paradigm, such as Figure 1 shown in (c). While retaining the residual connection and MLP structure in the Transformer structure, it incorporates the gated mechanism in the Mamba structure.

[0050] In the structural block under the same paradigm, each structural block includes two residual sub-blocks; the first residual sub-block contains an unspecified module and a gated mechanism; the second residual sub-block consists of a two-layer MLP with non-linear activation. The formula for the structural block is as follows:

[0051]

[0052] where x i represents the input of the i-th layer, represents the output of the first residual block, σ represents SiLu, and Y i represents the unspecified module. By specifying the specific design of this module, different models can be obtained. As Figure 3 shown, this module is specified as an attention module, an SSM module, and a multi-scale convolutional module. The formulas are as follows:

[0053] Y i = Attention(x i ·W Q , x i ·W K , x i ·W V ) (3)

[0054] Y i = conv 1×1 (concat(conv 1×1 (x i ), conv 3×3 (x i ), conv 5×5 (x i ))) (4)

[0055]

[0056] Among them, Attention represents the standard attention, and its calculation formula is Q, K, and V represent the query, key, and value matrices; j ∈ {1, 2, 3, 4} represents one of the four scanning directions; z j represents the feature sequence output by the extended scan; It means that then the four resulting feature sequences are processed separately by the S6 block; S6 represents the state space model to process the input feature sequence.

[0057] Furthermore, the designed ordinary differential equation-guided feature fusion module is used to integrate visual features from different encoding strategies, specifically:

[0058] The ordinary differential equation is represented by an equation with the function y(t) as the variable t and its derivative, as follows:

[0059]

[0060] Among them, f(t, y(t)) defines a time-varying function with discrete time-varying time, t = t0, t1....., t0 represents the initial value of the ODE, and the solution at time t1 is expressed as:

[0061]

[0062] Formula (7) is rewritten as:

[0063] y(t + Δt) = y(t) + Δtf(t, y(t)) (8)

[0064] Among them, Δt = T1 - T0, by defining Δtf(T, y(t)) = hg(y t ), y(t + Δt) = y t+1 and y(t) = y t , formula (8) is mapped to the Resblock in ResNet, and formula (8) is expressed in formula (9) as:

[0065] y t+1 = y i + hg(y t ) (9)

[0066] Using the fourth-order Runge-Kutta method to replace the first-order Euler method to reduce the local truncation error of formula (8), the formula is as follows:

[0067]

[0068] Among them, F1, F2, F3, and F4 represent the slopes calculated at different intermediate points; g(y t ) describes the rate of change of y with time t; h represents the step size.

[0069] The ODE-guided feature fusion module improves the network structure by borrowing ideas from the field of mathematics. Taking Figure 4 the convolutional layer-activation layer-convolutional layer-activation layer in (b) as the basic encoder block to implement Figure 4 the fourth-order Runge-Kutta method for solving ODEs in (a). This method uses the integration process of ordinary differential equations (ODEs) to smoothly process the input data, effectively maintaining the stability of the solution during the time-stepping process, thereby suppressing the amplification of noise caused by numerical errors. At the same time, the dynamic adjustment ability of the ODE model enables it to adaptively distinguish signals and noise and strengthen useful signals. During the backpropagation process, the ODE model can propagate gradients more effectively and enhance target features because the gradients can more accurately reflect the contribution of target features to the final output.

[0070] Furthermore, the deep network is trained using binary cross-entropy loss and cross-linking loss to improve the network's performance in infrared small target detection. The formula is as follows:

[0071] l = Bce(y,) + Iou(y,) (11)

[0072] where Bce(y,) represents binary cross-entropy loss and Iou(y,) represents cross-linking loss.

[0073] To verify the performance of the OFSPNet network, experiments were conducted on the NUAA dataset, SIRST-5K dataset, and SIRST-AUG dataset. As Figure 2 shown, given an input image I, each pixel is classified through the end-to-end processing of the network to distinguish whether it is a target pixel, and finally a segmentation result of the same size as I is output. The ratio of the training set to the test set is 8:2. During training, each image is cropped to 224x224. OFSPNet was compared with existing deep learning-based infrared small target detection algorithms. The algorithm is implemented based on Pytorch, and the batch size is set to 4. The optimizer uses Adams, where the momentum and weight decay coefficients are set to 0.9 and 0.0004 respectively. The initial learning rate is 0.001, and the decay strategy of CosineAnnealingLR is also used. It is trained for 150 epochs on the IRST-5k and ISTR-AUG datasets and 300 epochs on the NUAA dataset. In terms of hardware, a RTX3090 GPU is used for training.

[0074] Loss function

[0075] The binary cross-entropy loss and the cross-linking loss are combined to train the network, aiming to improve its performance in infrared small target detection. The infrared small target detection method based on segmentation can be regarded as a binary classification task, and the binary cross-entropy loss optimizes the classification accuracy of the model. The cross-linking loss improves the localization accuracy and ensures a high overlap between the predicted pixels and the true pixels. This combination not only enhances the detection accuracy but also improves the robustness of the model in complex backgrounds.

[0076] Evaluation Metrics

[0077] In terms of evaluation metrics, classic semantic segmentation evaluation metrics are used, including Precision, Recall, F-measure, and mean Intersection over Union (mIoU). Precision and Recall represent the ratios of correctly classified pixels to all marked targets and predicted targets respectively, and they influence each other. To measure the relationship between them, F-measure is used, which means they are equally important. The formulas are as follows, where TP, FP, and FN represent the numbers of true positives, false positives, and false negatives respectively:

[0078]

[0079] In addition, the ROC curve is used to characterize the dynamic relationship between the false positive rate (FPR) and the true positive rate (TPR). The area under the curve (AUC) is used as a key metric for quantitatively evaluating the ROC, which reflects the ability of the model to distinguish positive and negative class samples at all possible thresholds. A higher AUC value indicates better classification performance of the model. The AUC value of a perfect classifier is 1, while that of a classifier that randomly guesses is close to 0.5. The formulas are as follows:

[0080]

[0081] Comparative Experiments

[0082] 1) Numerical Evaluation: To accurately illustrate the effectiveness of OFSPNet, the present invention uses numerical methods for quantitative evaluation. As can be seen from Table 1, the method proposed in the present invention reaches the maximum values in terms of the two metrics of mIoU and F-measure. In addition, to more vividly show the comparison of AUC, Figure 5 the ROC curves of the method comparison are given. The experimental data show that OFSPNet has strong background suppression ability, can detect targets more accurately, and can segment targets more accurately.

[0083] Table 1. Comparison with the Latest Methods on the NUAA, SIRST-5K, and SIRST-AUG Datasets

[0084]

[0085] 2) Visual Evaluation: AsFigures 6 - 8 As shown, typical infrared small target scenarios are selected from three datasets respectively, and a visual comparison of seven infrared small target detection methods is carried out. In the figure, the detected small targets are magnified by a red dotted rectangle, the green dotted circle indicates missed detection, and the yellow dotted circle indicates false detection. Compared with other methods, the method proposed in the present invention can not only accurately detect the target, but also detect more complete target pixels. To facilitate the observation of clutter in the detection results, Figure 9 The 3D display of 6 scenarios shows that, as can be seen from the peak and the area under the peak, the method of the present invention has a higher prediction confidence in the target and is closer to the ground-truth (GT).

[0086] Ablation experiment

[0087] To verify the rationality of OFSPNet, by controlling variables, each module is ablated, and it is verified that each module has a certain improvement in the performance of the network. Concat refers to the simple channel concatenation of the outputs of the branches, and ODFM is the feature fusion module of the present invention.

[0088] Table 2 Module ablation

[0089]

[0090] Parameter setting:

[0091] For the setting of the loss function, the binary cross-entropy loss and the cross-linking loss are combined to improve the detection accuracy. The binary cross-entropy loss function optimizes the classification accuracy of the model, while the IOU loss focuses on the segmentation accuracy. The attention of the model to the pixel-level classification and the overall segmentation performance can be flexibly adjusted by weighting to better cope with the problem of class imbalance. Among them, α and β are balance coefficients, and the formula is as follows:

[0092] l i = αBce(y, ) + βIou(y, )

[0093] Table 3 Parameter setting

[0094] α∶β mIOU Precision Recall Fmeasure 1∶1 70.44 82.59 83.73 83.15 2∶1 70.55 84.38 81.83 83089 3∶2 73.62 85.87 83.77 84.80 2∶3 66.06 77.30 81.96 79.56

[0095] As can be seen from Table 3, when α∶β = 3∶2, the performance is the best.

[0096] The content not described in detail in the specification of the present invention belongs to the prior art well-known to those skilled in the art. Although the illustrative specific embodiments of the present invention have been described above for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

Claims

1. Infrared small target detection network with same paradigm feature extraction and ODE guided feature fusion, characterized by: The network adopts a three-branch network architecture of CNN, Transformer and Mamba, optimizes the structures of the Mamba and Transformer, retains the MLP module and residual module in the Transformer structure, no longer specifies a specific attention module, introduces a gating module, and uses it as a structural block under the same paradigm; at the same time, a feature fusion module guided by ordinary differential equations is designed to integrate visual features from different encoding strategies and enhance the visual representation ability of the network model.

2. The infrared small target detection network with same paradigm feature extraction and ODE guided feature fusion according to claim 1 is characterized in that: In the structural blocks under the same paradigm, each structural block includes two residual sub-blocks; the first residual sub-block contains an unspecified module and a gating mechanism; the second residual sub-block consists of a two-layer MLP with nonlinear activation. The formula of the structural block is as follows: Among them, x i represents the input of the i-th layer, represents the output of the first residual block, σ represents SiLu, Y i Represents an unspecified module. This module is specified as an attention module, an SSM module, and a multi-scale convolution module. The formula is as follows: Y i =Attention(x i ·W Q ,x i ·W K ,x i ·W V ) (3) Y i =conv 1×1 (Concat(conv 1×1 (x i ),conv 3×3 (x i ),conv 5×5 (x i ))) (4) Among them, Attention represents standard attention, and the calculation formula is Q, K, V represent query, key, and value matrices; j∈{1,2,3,4} represents one of the four scanning directions; z j represents the feature sequence output by the extended scan; It means that the four resulting feature sequences are then processed individually by the S6 block; S6 represents the state-space model processing the input feature sequence.

3. The infrared small target detection network with same paradigm feature extraction and ODE guided feature fusion according to claim 2 is characterized in that: The feature fusion module guided by the ordinary differential equation is designed to integrate visual features from different encoding strategies, specifically: Ordinary differential equations are represented by equations with the function y(t) as the variable t and its derivatives, as follows: Among them, f(t,y(t)) defines a time-varying function with discrete time-varying time, t=t0,t1....., t0 represents the initial value of the ODE, and the solution at time t1 is expressed as: Formula (7) is rewritten as: y(t+Δt)=y(t)+Δtf(t,y(t)) (8) Where Δt = t1-t0, by defining Δtf(t,y(t)) = hg(y t ), y(t+Δt)=y t+1 and y(t)=y t , formula (8) is mapped to Resblock in ResNet, and formula (8) is expressed in formula (9) as: y t+1 =y i +hg(y t ) (9) The fourth-order Runge-Kutta method is used instead of the first-order Euler method to reduce the local truncation error of formula (8), and the formula is as follows: Among them, F1, F2, F3, and F4 represent the slopes calculated at different intermediate points; g(y t ) describes the rate of change of y with time t; h represents the step size.

4. The infrared small target detection network with same paradigm feature extraction and ODE guided feature fusion according to claim 3 is characterized in that: The deep network is trained using binary cross entropy loss and cross-link loss to improve the performance of the network in infrared small target detection. The formula is as follows: l=Bce(y,)+Iou(y,) (11) Among them, Bce(y,) represents the binary cross entropy loss and Iou(y,) represents the cross-linking loss.

Citation Information

Patent Citations

  • Elevator fault diagnosis method based on short sequence time convolutional network

    CN118964959A

  • Multi-scale feature fusion triple branch network method for multi-organ segmentation

    CN119478404A