An infrared bird motion target detection method based on background learning and adaptive motion feature fusion

CN122821083APending Publication Date: 2026-09-25HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610869887.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]目的:为了克服现有技术中存在的不足,解决现有红外鸟类运动目标检测中复杂高亮背景干扰强、检测误报率高、对运动特征利用不足等问题,本发明提供一种基于背景学习和自适应运动特征融合的红外鸟类运动目标检测方法

Benefits of technology

[0121]1.本发明提出的背景显著性区域学习方法与置信度机制能够在复杂环境中有效地区分稳定背景显著性区域与鸟类运动目标,在面对高噪声、高亮背景、复杂结构背景等场景时,具备显著的抗干扰能力,并能有效减少网络误差,使得整体检测性能获得大幅提升;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821083A_ABST
    Figure CN122821083A_ABST
Patent Text Reader

Abstract

The application discloses an infrared bird motion target detection method based on background learning and adaptive motion feature fusion, and belongs to the technical field of infrared image target detection. The method comprises the following steps: acquiring a continuous infrared image sequence containing a bird motion target; performing background suppression processing on the continuous infrared image sequence by using a background saliency region image; inputting the background suppression processed infrared motion target detection network, extracting static features, slow motion features and fast motion features, and fusing the features to obtain an infrared bird motion target; the background saliency region image is continuously updated in the training stage and the detection stage of the infrared motion target detection network; the training of the infrared motion target detection network comprises the following steps: performing background suppression processing on a data set by using the background saliency region image, and training the data set subjected to the background suppression processing. The scheme adopts a combined network framework of background learning and infrared bird motion target detection, and can improve the detection accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of infrared image target detection technology, and specifically relates to an infrared bird moving target detection method based on background learning and adaptive motion feature fusion. Background Technology

[0002] With the increasing demand for all-weather monitoring in power facilities, airport security, and important industrial areas, infrared image target detection technology is widely used in security monitoring, reconnaissance, and industrial-related detection fields due to its advantages such as low light interference, strong penetration, better identification of concealed targets, and all-weather availability.

[0003] Existing methods primarily rely on deep convolutional neural networks or traditional infrared enhancement algorithms to identify moving bird targets. However, substation environments often contain numerous background hotspots, such as conductive metal components, equipment contact terminals, substation insulators, and other heated or reflective areas. These heated areas also exhibit significant brightness advantages in the spatial domain, offering limited brightness differentiation from moving targets like birds, especially under low-contrast conditions such as rain and fog. Furthermore, infrared bird targets are small, typically only a dozen pixels in an image, and their movements are disordered, with frequent contour changes and unstable speeds. Against the backdrop of a substation, they are easily obscured by equipment, leading to significant positional shifts and greatly increasing the difficulty of detection. Summary of the Invention

[0004] Objective: In order to overcome the shortcomings of existing technologies and solve the problems of strong interference from complex and bright backgrounds, high false alarm rate, and insufficient utilization of motion features in existing infrared bird moving target detection, this invention provides an infrared bird moving target detection method based on background learning and adaptive motion feature fusion.

[0005] Technical solution:

[0006] An infrared bird moving target detection method based on background learning and adaptive motion feature fusion includes:

[0007] Acquire a continuous sequence of infrared images containing moving birds;

[0008] Background suppression processing was performed on the acquired continuous infrared image sequence containing moving bird targets using images of salient background regions;

[0009] A continuous infrared image sequence containing moving birds, after background suppression processing, is input into a pre-trained infrared moving target detection network. Static features, slow motion features, and fast motion features are extracted and fused to obtain infrared moving birds.

[0010] The background salient region image is continuously updated during the training and detection phases of the infrared moving target detection network;

[0011] The training of the infrared moving target detection network includes background suppression processing of the dataset using images of salient background regions, and training using the background-suppressed dataset.

[0012] This solution employs a joint network framework combining background learning and infrared bird moving target detection, which can improve detection accuracy and robustness.

[0013] In some embodiments, the construction of the background salient region image includes:

[0014] Collect M consecutive infrared image sequences containing moving bird targets to obtain sample images;

[0015] The local brightness characteristics of each pixel location are calculated in the local neighborhood of the collected sample images, and a preset background containing salient regions is obtained through regional connectivity.

[0016] By using inter-frame difference to remove foreground targets from a preset background and selecting the median brightness of each pixel, an image of the salient background region is obtained.

[0017] In some embodiments, the step of calculating the local brightness characteristics of the local neighborhood of each pixel location based on the acquired sample image, and obtaining a preset background containing salient regions through region connectivity includes:

[0018] The local brightness characteristics within the local neighborhood of each pixel location in the M-frame sample image are calculated using the following formula:

[0019] ;

[0020] in, For the sample image of frame t, let the coordinates be... The pixel brightness at coordinates (i,j) within a local neighborhood centered on the pixel. Coordinates are The local average brightness of the pixels. Coordinates are The local brightness standard deviation of the pixels, Represented by coordinates The local neighborhood range centered on the pixel, Indicates the number of pixels within a local neighborhood;

[0021] The brightness threshold is calculated based on local brightness characteristics, and its expression is:

[0022] ;

[0023] in, The coordinates of the sample image in frame t are The brightness threshold of the pixels, к is an adjustment parameter used to adjust the level of the local brightness threshold, thereby controlling the sensitivity of the determination of salient background areas;

[0024] When the pixel brightness of the t-th frame sample image satisfies When this happens, the pixel is identified as a highlight pixel. The coordinates of the sample image in frame t are Pixel brightness;

[0025] Connectivity analysis and edge structure constraints are performed on all bright pixels to divide the sample image into several spatially independent salient structure regions. The brightness of pixels in the salient structure regions is retained, while the brightness of pixels in the remaining regions is set to 0, thus generating a preset background for each frame of the sample image.

[0026] In some embodiments, the step of using inter-frame difference to remove foreground targets in a preset background and selecting the median brightness of each pixel to obtain an image of the salient background region includes:

[0027] The preset background is subjected to inter-frame differencing to obtain a differencing image. Foreground targets are then removed by accumulating pixel fluctuations. The expression is as follows:

[0028] ;

[0029] in, The instantaneous change intensity of the background is preset for the sample image of frame t. The cumulative change intensity of the preset background. The preset background for the sample image of frame t, The preset background for the sample image of frame t-1;

[0030] when or When a pixel is determined to be a foreground target, it is removed from the preset background.

[0031] The median pixel brightness of each position in the preset background after removing the foreground target from M frames is taken to obtain the image of the salient background region;

[0032] In some embodiments, updating the background salient region image includes:

[0033] Define the image update conditions for background salient regions, including rate constraints and time consistency constraints;

[0034] When both the rate constraint and the time consistency constraint are satisfied, the background salient region image is updated using the background information of the current frame image.

[0035] Specifically, during the training phase of the infrared moving target detection network, the background salient region image is updated based on the background information of the training images labeled with real data; during the detection phase of the infrared moving target detection network, the background salient region image is updated based on the background information of the detection result image of the infrared moving target detection network.

[0036] The core idea of ​​background salient region update is to gradually correct the background salient regions using the background information of the current frame image. During the training phase, the real annotation information provided in the dataset is used to obtain the background salient region information, while during the testing phase, the prediction results of the detection network are relied upon to obtain the background salient region information.

[0037] During the training phase, the background saliency region update process is constrained using the real object annotation information provided by the dataset. The ground truth image after image annotation at time t is obtained. The brightness of the marked area is set to 0 to mask the foreground target. Then, the local average brightness of the ground truth image at time t is calculated. Local brightness standard deviation Constructing a brightness threshold based on brightness characteristics When the pixel brightness meets the requirements When this happens, the pixel is identified as a highlighted pixel. Connectivity analysis and edge structure constraints are performed on all highlighted pixels to generate a ground truth background. .

[0038] During the testing phase, the prediction results of the detection network are used to update the salient regions of the background. The prediction result image at time t is obtained, and the local average brightness and local brightness standard of the predicted image are calculated to construct a brightness threshold. Through threshold judgment and connectivity analysis, the image background of the current frame is generated. .

[0039] In some embodiments, the image update conditions for defining the salient background region include:

[0040] Calculate the local average brightness, local brightness standard deviation, and brightness threshold of the image in the background salient region;

[0041] The state probability matrix is ​​constructed based on the background salient region image, and its expression is:

[0042] ;

[0043] in, The coordinates in the state probability matrix are The state probability value of the pixel. The coordinates of the background salient region in the image are The brightness threshold of the pixels, The coordinates of the background salient region in the image are Pixel brightness;

[0044] A rate constraint mechanism is introduced during the background saliency region update process to prevent short-term anomalies or foreground objects from being incorrectly absorbed into the background saliency region model. This is implemented when a background saliency region image with foreground objects removed is obtained. Then, based on the rate constraint condition defined by the local average brightness of the background salient region image, the rate of change of the background brightness in the region is calculated, and the expression is:

[0045] ;

[0046] in, The coordinates of the sample image in frame t are The rate of change of background brightness in the region of pixels, The coordinates of the image obtained after masking the foreground target in the training image or detection result image of frame t are: The local average brightness of the pixels. The image coordinates of the background salient region in frame t-1 are: The local average brightness of the pixels;

[0047] when When the rate constraint condition is met, ε is the threshold for the change in regional background brightness;

[0048] Furthermore, to effectively distinguish between long-term background devices and short-term bird-like moving targets, the state probability matrix is ​​updated over time to describe the background state probability of a pixel at time t. A larger value indicates that the pixel is more likely to be background. This probability variable is continuously updated using a time-recursive approach and is expressed as:

[0049] ;

[0050] in, The coordinates of the sample image in frame t are The state probability value of the pixel. The coordinates of the sample image in frame t-1 are The state probability value of the pixel. These are state correction parameters, used to correct the probability values ​​of each pixel in the state probability matrix that belong to the background;

[0051] when If η is the state probability threshold, then the structural region is considered to exist in the spatial location for a long time and satisfy the time consistency constraint.

[0052] When the rate constraint and time consistency constraint are satisfied, the background saliency region image is updated using an exponentially weighted average method, expressed as:

[0053] ;

[0054] in, The image coordinates of the background saliency region after the update in frame t are: pixel brightness, The image coordinates of the background salient region in frame t-1 are: The pixel brightness, where α is the learning rate parameter. This is used to control the contribution of the current frame's background information to the update of the image of the salient background region.

[0055] In some embodiments, the background suppression processing step includes:

[0056] Calculate the local brightness characteristics of the updated background salient region image;

[0057] The background deviation of the training image or the detection image relative to any pixel within the salient background region is calculated based on the local brightness characteristics of the image in the salient background region. The formula is as follows:

[0058] ;

[0059] in, This represents the coordinates of the training or detection image relative to the background saliency region. The degree of pixel deviation, The coordinates of the t-th frame of the image to be trained or detected are: pixel brightness, The image coordinates of the background saliency region after the update in frame t are: pixel brightness, The updated background saliency region image coordinates are: The standard deviation of local brightness of the pixels;

[0060] Based on this, the system further generates a background confidence map and constructs a background confidence score according to the background deviation, expressed as:

[0061] ;

[0062] in, Background confidence level;

[0063] After the background confidence model is built, the background confidence map is mapped to a suppression weight map, and the intermediate features are reweighted using the background suppression function during the network feature extraction stage. The expression is as follows:

[0064] ;

[0065] in, The coordinates of the training image or the detection image after background suppression are: pixel brightness, This is a pixel-by-pixel multiplication.

[0066] In some embodiments, the infrared moving target detection network's detection steps for infrared bird moving targets include:

[0067] The continuous infrared image sequence containing moving birds after background suppression is input into the pre-trained infrared moving target detection network backbone in batch processing, and the feature images of moving birds in infrared images are extracted in parallel through convolutional layers and residual structures.

[0068] The extracted infrared bird moving target feature images are input into the stationary feature extraction branch, the slow motion feature extraction branch, and the fast motion feature extraction branch to further extract stationary features, slow motion features, and fast motion features;

[0069] The extracted static features, slow-moving features, and fast-moving features are input into the adaptive weight fusion module for weighted fusion, and infrared bird moving targets are obtained through the decoupling head.

[0070] In some embodiments, the static feature extraction step includes:

[0071] By using multi-scale parity-symmetric filtering convolution kernels, the phase response of the local structure of a stationary target is analyzed at multiple scales and directions. The expression for the parity-symmetric filtering convolution kernel is as follows:

[0072] ;

[0073] in, For infrared bird moving target feature images, the odd-even symmetric filtering convolution kernel is used in coordinates. The values ​​of θ are given by k, where k is the index scale and θ is the direction. The coordinates of the infrared bird moving target feature image are: The even-symmetric component of the pixel, The coordinates of the infrared bird moving target feature image are: The odd symmetric component of the pixel, where i is the imaginary unit;

[0074] The complex domain response is constructed by pairing odd-symmetric and even-symmetric components, and the expression is:

[0075] ;

[0076] in, The coordinates of the infrared bird moving target feature image are: The even-symmetric response of the pixels, The coordinates of the infrared bird moving target feature image are: The odd-symmetric response of the pixels, For convolution operations, The coordinates of the infrared bird moving target feature image input to the static feature extraction branch are: eigenvalues;

[0077] Using the phase consistency model, the local structural features of a stationary target are enhanced, as expressed by:

[0078] ;

[0079] in, The coordinates in the feature map obtained after applying an even-odd symmetric filter convolution kernel to an infrared image of moving birds are: The phase consistency of the pixels, The coordinates in the feature map obtained after applying an even-odd symmetric filter convolution kernel to an infrared bird moving target feature image are: The amplitude of pixels, The coordinates in the feature map obtained after applying an even-odd symmetric filter convolution kernel to an infrared bird moving target feature image are: The phase of the pixel, The coordinates of the infrared bird moving target feature image are: The initial phase of the pixel;

[0080] The initial phase expression is:

[0081] ;

[0082] Where atan2 is the arctangent function, which calculates the phase angle of a two-dimensional vector based on its two components;

[0083] By utilizing a phase difference amplification mechanism, subtle phase changes are enhanced through a nonlinear strategy, making them stand out from a static background. Pixel integration is performed using 1×1 convolution, followed by normalization and activation functions to finally obtain the static features, expressed as:

[0084] ;

[0085] in, It is a static feature. For convolution operations, The coordinates of the infrared bird moving target feature images in frame t and frame t-1 are: The phase difference of the pixels, ReLU is the activation function, and BN is the normalization layer. To multiply element by element, N represents N frames of infrared bird moving target feature images.

[0086] In some embodiments, the slow motion feature extraction step includes:

[0087] A two-dimensional fast Fourier transform and phase enhancement are performed on the input infrared image of moving birds to obtain the enhanced frequency domain features, expressed as follows:

[0088] ;

[0089] in, The coordinates of the infrared bird moving target feature images in frame t and frame t-1 are: The phase difference of the pixels, It is a phase difference amplifier. The coordinates of the infrared bird moving target feature image in frame t are... The phase of the pixel, The coordinates of the infrared bird moving target feature image in frame t+1 are: The phase of the pixel, This refers to the pixel features with frequency domain coordinates (u,v) after frequency domain enhancement of an infrared image of moving birds. The pixel features with frequency domain coordinates (u,v) after the two-dimensional fast Fourier transform of an infrared image of moving birds are represented by the pixel features, where j is the imaginary unit. The pixels of the infrared bird moving target feature image in frame t are represented by spatial domain coordinates. The phase is transformed to the frequency domain coordinates (u,v); the phase difference in the time dimension is calculated and amplified using a phase amplifier, and the amplified phase difference is reconstructed back to the frequency domain to obtain the enhanced frequency domain features;

[0090] To further enhance the ability to perceive subtle boundary changes, the local phase gradient is calculated to capture local structural changes in slow-moving targets, yielding local structural features, expressed as:

[0091] ;

[0092] in, The coordinates of the infrared bird moving target feature image in frame t are... The phase gradient of the pixel, The local structural features of pixels with frequency domain coordinates (u,v) in an infrared image of moving birds;

[0093] The enhanced frequency domain features and local structural features are fused to obtain slow motion response features. Then, through convolution, normalization, and activation functions, the slow motion frequency domain features are further extracted, as expressed in the following expression:

[0094] ;

[0095] in, The feature value at coordinates (u, v) is the slow motion frequency domain feature map of an infrared image of a moving bird. This is an element-wise addition;

[0096] Slow motion characteristics are obtained through inverse fast Fourier transform.

[0097] In some embodiments, the rapid motion feature extraction step includes:

[0098] The input sequence of infrared bird motion target feature images is stacked along the time dimension to form a three-dimensional tensor. A one-dimensional fast Fourier transform is then performed on the time dimension to obtain the frequency map of the pixels in time. A frequency energy map is constructed by focusing on the strong responses generated in the high-frequency regions of the frequency map, and the expression is:

[0099] ;

[0100] in, The three-dimensional tensor formed by stacking infrared bird motion target feature image sequences along frame number time t is located at coordinates... The value at that location, for The index f obtained after performing a one-dimensional fast Fourier transform on t is located at coordinates... Frequency response at that point, for The coordinates in the cumulative energy map of the high-frequency region are: The value at that location, , For high-frequency starting indices, frequencies close to the Nyquist zone are typically chosen. ρ is the proportionality constant. The maximum frequency is N, where N represents N frames of infrared bird moving target feature images;

[0101] To further utilize frequency domain structure information, the frequency energy map is used as a mask to guide the frequency domain feature analysis of the image to be detected; first, a two-dimensional fast Fourier transform is performed on each frame to obtain... Mapping the time-frequency plot to frequency domain guided weights, and then normalizing and... Element-wise multiplication is performed, taking advantage of the high energy distribution near the direction of movement in terms of time frequency of fast-moving targets, thereby enhancing the spatial frequency domain characteristics of bird-like moving targets. The expression is as follows:

[0102] ;

[0103] in, Let be the enhanced spatial frequency domain features of the t-th frame, and Norm be the normalization operation;

[0104] The enhanced spatial frequency domain features are further strengthened by using inter-frame difference. Pixel aggregation of the difference image yields the fast motion frequency domain features, expressed as:

[0105] ;

[0106] in, The value at frequency domain coordinates (u, v) is taken from the fast motion frequency domain feature map of an infrared bird moving target feature image. Let be the value at frequency domain coordinates (u, v) in the frequency domain residual feature map of the infrared bird moving target feature image of frame t. For convolution operations with a kernel size of 3×3, The value at frequency domain coordinates (u,v) is taken in the enhanced spatial frequency domain feature map of the infrared bird moving target feature image of frame t+1.

[0107] Rapid motion characteristics are obtained through inverse fast Fourier transform.

[0108] In some embodiments, the step of inputting the extracted stationary features, slow-moving features, and fast-moving features into an adaptive weight fusion module for weighted fusion, and obtaining the infrared bird motion target through the decoupling head includes:

[0109] Based on the differences in motion energy distribution in the infrared image after background suppression, the weights of the stationary feature extraction branch, the slow motion feature extraction branch, and the fast motion feature extraction branch are determined to obtain the fusion weights that satisfy the normalization constraint. The steps include:

[0110] For an input N-frame continuous infrared image containing a moving bird target after background suppression, optical flow is used to calculate motion energy to measure motion amplitude, expressed as:

[0111] ;

[0112] Where E represents motion energy, and N represents N frames of infrared images of moving birds. The infrared bird moving target feature images of frame t and frame t-1 are located at coordinates... The optical flow amplitude is calculated at each pixel, where H is the height of the infrared bird moving target feature image and W is the width of the infrared bird moving target feature image.

[0113] The static features, slow motion features, and fast motion features are concatenated with motion energy. A multilayer perceptron is used to obtain the weights of the static feature extraction branch, slow motion feature extraction branch, and fast motion feature extraction branch through nonlinear transformation, and then normalized. The expression is:

[0114] ;

[0115] in, Let be the weight of the i-th branch, i = 1, 2, 3; Concat is the connection operation; MLP is the multilayer perceptron; and Softmax is the normalization operation. It is a static feature. Characterized by slow motion. Characterized by rapid motion;

[0116] To ensure the stability of the distribution of the fused features, the fused feature map is obtained through channel normalization and lightweight convolution, expressed as:

[0117] ;

[0118] in, To fuse feature maps, Extract branch weights for static features. Extracting branch weights for slow motion features. Extracting branch weights for fast motion features;

[0119] The fused feature map is used to obtain infrared bird motion targets through YOLOX's decoupling head and non-maximum suppression.

[0120] Beneficial effects: The infrared bird moving target detection method based on background learning and adaptive motion feature fusion provided by this invention has the following advantages:

[0121] 1. The background saliency region learning method and confidence mechanism proposed in this invention can effectively distinguish stable background saliency regions from bird moving targets in complex environments. It has significant anti-interference ability when facing high noise, high brightness backgrounds, complex structure backgrounds, etc., and can effectively reduce network errors, thereby greatly improving the overall detection performance.

[0122] 2. The adaptive motion feature fusion infrared moving target detection network designed in this invention captures the salience of bird moving targets from multiple complementary feature perspectives through three different motion feature extraction branches and an adaptive feature fusion module, effectively improving the accuracy and stability of target detection, enabling the network to maintain optimal feature expression ability in different environments, and achieving highly robust target detection performance. Attached Figure Description

[0123] Figure 1 This is a flowchart of an infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to an embodiment of the present invention;

[0124] Figure 2 This is a schematic diagram of the overall network structure of the infrared moving target detection network according to an embodiment of the present invention;

[0125] Figure 3 This is a schematic diagram of the static feature extraction branch structure according to an embodiment of the present invention;

[0126] Figure 4 This is a schematic diagram of the slow motion feature extraction branch structure according to an embodiment of the present invention;

[0127] Figure 5 This is a schematic diagram of the branch structure for rapid motion feature extraction according to an embodiment of the present invention;

[0128] Figure 6 This is a schematic diagram of the adaptive weight fusion module structure according to an embodiment of the present invention. Detailed Implementation

[0129] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0130] The present invention will be further described below with reference to specific embodiments.

[0131] This embodiment provides an infrared bird moving target detection method based on background learning and adaptive motion feature fusion, such as... Figure 1 As shown, it includes:

[0132] Acquire a continuous sequence of infrared images containing moving birds;

[0133] Background suppression processing was performed on the acquired continuous infrared image sequence containing moving bird targets using images of salient background regions;

[0134] A continuous infrared image sequence containing moving birds, after background suppression processing, is input into a pre-trained infrared moving target detection network. Static features, slow motion features, and fast motion features are extracted and fused to obtain infrared moving birds.

[0135] The background salient region image is continuously updated during the training and detection phases of the infrared moving target detection network;

[0136] The training of the infrared moving target detection network includes background suppression processing of the dataset using images of salient background regions, and training using the background-suppressed dataset.

[0137] The construction of the background salient region image includes:

[0138] Collect M consecutive infrared image sequences containing moving bird targets to obtain sample images;

[0139] The local brightness characteristics of each pixel location are calculated in the local neighborhood of the collected sample images, and a preset background containing salient regions is obtained through regional connectivity.

[0140] By using inter-frame difference to remove foreground targets from a preset background and selecting the median brightness of each pixel, an image of the salient background region is obtained.

[0141] The step of calculating the local brightness characteristics of the local neighborhood of each pixel location based on the acquired sample image, and obtaining the preset background containing salient regions through region connectivity includes:

[0142] The local brightness characteristics within the local neighborhood of each pixel location in the M-frame sample image are calculated using the following formula:

[0143] ;

[0144] in, For the sample image of frame t, let the coordinates be... The pixel brightness at coordinates (i,j) within a local neighborhood centered on the pixel. Coordinates are The local average brightness of the pixels. Coordinates are The local brightness standard deviation of the pixels, Represented by coordinates The local neighborhood range centered on the pixel, Indicates the number of pixels within a local neighborhood;

[0145] The brightness threshold is calculated based on the above local brightness characteristics, and its expression is:

[0146] ;

[0147] in, The coordinates of the sample image in frame t are The brightness threshold of the pixels, к is an adjustment parameter used to adjust the level of the local brightness threshold, thereby controlling the sensitivity of the determination of salient background areas;

[0148] When the pixel brightness of the t-th frame sample image satisfies When this happens, the pixel is identified as a highlight pixel. The coordinates of the sample image in frame t are Pixel brightness;

[0149] Connectivity analysis and edge structure constraints are performed on all bright pixels to divide the sample image into several spatially independent salient structure regions. The brightness of pixels in the salient structure regions is retained, while the brightness of pixels in the remaining regions is set to 0, thus generating a preset background for each frame of the sample image.

[0150] The step of removing foreground targets from a preset background using inter-frame difference and selecting the median brightness of each pixel to obtain an image of the salient background region includes:

[0151] The preset background is subjected to inter-frame differencing to obtain a differencing image. Foreground targets are then removed by accumulating pixel fluctuations. The expression is as follows:

[0152] ;

[0153] in, The instantaneous change intensity of the background is preset for the sample image of frame t. The cumulative change intensity of the preset background. The preset background for the sample image of frame t, The preset background for the sample image of frame t-1;

[0154] when or When a pixel is determined to be a foreground target, it is removed from the preset background.

[0155] The median pixel brightness of each position in the preset background after removing the foreground target from M frames is taken to obtain the image of the salient background region.

[0156] The update of the background salient region image includes:

[0157] Define the image update conditions for background salient regions, including rate constraints and time consistency constraints;

[0158] When both the rate constraint and the time consistency constraint are satisfied, the background salient region image is updated using the background information of the current frame image.

[0159] Specifically, during the training phase of the infrared moving target detection network, the background salient region image is updated based on the background information of the training images labeled with real data; during the detection phase of the infrared moving target detection network, the background salient region image is updated based on the background information of the detection result image of the infrared moving target detection network.

[0160] To achieve stable and continuous background salient region updates in complex infrared scenes, it is necessary to update the salient regions frame by frame during image processing. The core idea of ​​background salient region update is to gradually correct the background salient regions using the background information of the current frame image. During the training phase, the real-world annotation information provided in the dataset is used to obtain the background salient region information, while during the testing phase, the prediction results of the detection network are relied upon to obtain the background salient region information.

[0161] During the training phase, the background saliency region update process is constrained using the real object annotation information provided by the dataset. The ground truth image after image annotation at time t is obtained. The brightness of the marked area is set to 0 to mask the foreground target. Then, the local average brightness of the ground truth image at time t is calculated. Local brightness standard deviation Constructing a brightness threshold based on brightness characteristics When the pixel brightness meets the requirements When this happens, the pixel is identified as a highlighted pixel. Connectivity analysis and edge structure constraints are performed on all highlighted pixels to generate a ground truth background. .

[0162] During the testing phase, the prediction results of the detection network are used to update the salient regions of the background. The prediction result image at time t is obtained, and the local average brightness and local brightness standard of the predicted image are calculated to construct a brightness threshold. Through threshold judgment and connectivity analysis, the image background of the current frame is generated. .

[0163] The image update conditions for defining the salient background region include:

[0164] Calculate the local average brightness, local brightness standard deviation, and brightness threshold of the image in the background salient region;

[0165] The state probability matrix is ​​constructed based on the background salient region image, and its expression is:

[0166] ;

[0167] in, The coordinates in the state probability matrix are The state probability value of the pixel. The coordinates of the background salient region in the image are The brightness threshold of the pixels, The coordinates of the background salient region in the image are Pixel brightness.

[0168] Given the slow-changing physical characteristics of the substation background over time, a rate constraint mechanism is introduced during the background saliency region update process to prevent short-term anomalies or foreground targets from being incorrectly absorbed into the background saliency region model. This is done when a background image with foreground targets removed is obtained. Next, calculate the rate of change of background brightness in the region. Define the rate constraint condition as follows:

[0169] ;

[0170] in, The coordinates of the sample image in frame t are The rate of change of background brightness in the region of pixels, The coordinates of the image obtained by masking the foreground target region based on the detection box in the training image or detection result image of frame t are: The local average brightness of the pixels. The image coordinates of the background salient region in frame t-1 are: The local average brightness of the pixels;

[0171] when When the rate constraint condition is met, ε is the threshold for the change in regional background brightness;

[0172] Furthermore, to effectively distinguish between the long-term persistent background of devices and the short-term appearance of bird movement targets, a state probability matrix is ​​introduced to describe the background state probability of a pixel at time t. A larger value indicates that the pixel is more likely to be background. This probability variable is continuously updated through a time-recursive approach, expressed as:

[0173] ;

[0174] in, The coordinates of the sample image in frame t are The state probability value of the pixel. The coordinates of the sample image in frame t-1 are The state probability value of the pixel. These are state correction parameters, used to correct the probability values ​​of each pixel in the state probability matrix that belong to the background;

[0175] when If η is the state probability threshold, then the structural region is considered to exist in the spatial location for a long time and satisfy the time consistency constraint.

[0176] When both the rate constraint and the time consistency constraint are satisfied, the background saliency region image is updated using an exponentially weighted average method, expressed as:

[0177] ;

[0178] in, The image coordinates of the background saliency region after the update in frame t are: pixel brightness, The image coordinates of the background salient region in frame t-1 are: The pixel brightness, where α is the learning rate parameter. This is used to control the contribution of the current frame's background information to the update of the image of the salient background region.

[0179] In order to better preserve the detailed information of bird moving targets in the image, instead of directly removing the salient background regions from the image, we use a method of suppressing the salient background regions to weaken the interference of complex background on bird moving target detection.

[0180] The background suppression processing steps include:

[0181] Calculate the local brightness characteristics of the updated background salient region image;

[0182] The background deviation of the training image or the detection image relative to any pixel within the background salient region is calculated based on the local brightness characteristics of the updated background salient region image. The formula is as follows:

[0183] ;

[0184] in, This represents the coordinates of the training or detection image relative to the background saliency region. The degree of pixel deviation, The coordinates of the t-th frame of the image to be trained or detected are: pixel brightness, The image coordinates of the background saliency region after the update in frame t are: pixel brightness, The updated background saliency region image coordinates are: The standard deviation of local brightness of the pixels;

[0185] Based on this, the system further generates a background confidence map and constructs a background confidence score according to the background deviation, expressed as:

[0186] ;

[0187] in, Background confidence level;

[0188] After the background confidence model is built, the background confidence map is mapped to a suppression weight map, and the intermediate features are reweighted using the background suppression function during the network feature extraction stage. The expression is as follows:

[0189] ;

[0190] in, The coordinates of the training image or the detection image after background suppression are: pixel brightness, This is a pixel-by-pixel multiplication.

[0191] The infrared moving target detection network's detection steps for infrared bird moving targets include:

[0192] Using ResNet18 with the last pooling layer and all fully connected layers removed as the encoder backbone, a series of continuous infrared images containing moving birds after background suppression are input into the pre-trained infrared moving target detection network backbone in batch form. Each image is treated as an independent sample, and the feature images of moving birds in infrared images are extracted in parallel through convolutional layers and residual structures.

[0193] The extracted infrared bird moving target feature images are input into the stationary feature extraction branch, the slow motion feature extraction branch, and the fast motion feature extraction branch to further extract stationary features, slow motion features, and fast motion features;

[0194] The extracted static features, slow-moving features, and fast-moving features are input into the adaptive weight fusion module for weighted fusion. The infrared bird motion target is then obtained through a decoupled head. Figure 2 .

[0195] The static feature extraction step includes:

[0196] like Figure 3 As shown, the phase response of the local structure of a stationary target is analyzed at multiple scales and directions using a multi-scale parity-symmetric filtering convolution kernel. The expression for the parity-symmetric filtering convolution kernel is:

[0197] ;

[0198] in, For infrared bird moving target feature images, the odd-even symmetric filtering convolution kernel is used in coordinates. The values ​​of θ are given by k, where k is the index scale and θ is the direction. The coordinates of the infrared bird moving target feature image are: The even-symmetric component of the pixel, The coordinates of the infrared bird moving target feature image are: The odd symmetric component of the pixel, where i is the imaginary unit;

[0199] Even-symmetric filtering convolution kernels can capture symmetry patterns in local constructions, while odd-symmetric filtering convolution kernels are sensitive to gradient changes. They construct complex domain responses through odd-even pairing, as expressed in the following expression:

[0200] ;

[0201] in, The coordinates of the infrared bird moving target feature image are: The even-symmetric response of the pixels, The coordinates of the infrared bird moving target feature image are: The odd-symmetric response of the pixels, For convolution operations, The coordinates of the infrared bird moving target feature image input to the static feature extraction branch are: eigenvalues;

[0202] Using the phase consistency model, the local structural features of a stationary target are enhanced, as expressed by:

[0203] ;

[0204] in, The coordinates in the feature map obtained after applying an even-odd symmetric filter convolution kernel to an infrared image of moving birds are: The phase consistency of the pixels, The coordinates in the feature map obtained after applying an even-odd symmetric filter convolution kernel to an infrared bird moving target feature image are: The amplitude of pixels, The coordinates in the feature map obtained after applying an even-odd symmetric filter convolution kernel to an infrared bird moving target feature image are: The phase of the pixel, The coordinates of the infrared bird moving target feature image are: The initial phase of the pixel;

[0205] The initial phase expression is:

[0206] ;

[0207] Where atan2 is the arctangent function, which calculates the phase angle of a two-dimensional vector based on its two components;

[0208] By utilizing a phase difference amplification mechanism, subtle phase changes are enhanced through a nonlinear strategy, making them stand out from a static background. Pixel integration is performed using 1×1 convolution, followed by normalization and activation functions to finally obtain the static features, expressed as:

[0209] ;

[0210] in, It is a static feature. For convolution operations, The coordinates of the infrared bird moving target feature images in frame t and frame t-1 are: The phase difference of the pixels, ReLU is the activation function, and BN is the normalization layer. To multiply element by element, N represents N frames of infrared bird moving target feature images.

[0211] like Figure 4 As shown, the slow motion feature extraction step includes:

[0212] A two-dimensional fast Fourier transform and phase enhancement are performed on the input infrared image of moving birds to obtain the enhanced frequency domain features, expressed as follows:

[0213] ;

[0214] in, The coordinates of the infrared bird moving target feature images in frame t and frame t-1 are: The phase difference of the pixels, It is a phase difference amplifier. The coordinates of the infrared bird moving target feature image in frame t are... The phase of the pixel, The coordinates of the infrared bird moving target feature image in frame t+1 are: The phase of the pixel, This refers to the pixel features with frequency domain coordinates (u,v) after frequency domain enhancement of an infrared image of moving birds. The pixel features with frequency domain coordinates (u,v) after the two-dimensional fast Fourier transform of an infrared image of moving birds are represented by the pixel features, where j is the imaginary unit. The pixels of the infrared bird moving target feature image in frame t are represented by spatial domain coordinates. The phase is transformed to the frequency domain coordinates (u,v); the phase difference in the time dimension is calculated and amplified using a phase amplifier, and the amplified phase difference is reconstructed back to the frequency domain to obtain the enhanced frequency domain features;

[0215] To further enhance the ability to perceive subtle boundary changes, the local phase gradient is calculated to capture local structural changes in slow-moving targets, yielding local structural features, expressed as:

[0216] ;

[0217] in, The coordinates of the infrared bird moving target feature image in frame t are... The phase gradient of the pixel, The local structural features of pixels with frequency domain coordinates (u,v) in an infrared image of moving birds;

[0218] The enhanced frequency domain features and local structural features are fused to obtain slow motion response features. Then, through convolution, normalization, and activation functions, the slow motion frequency domain features are further extracted, as expressed in the following expression:

[0219] ;

[0220] in, The feature value at coordinates (u, v) is the slow motion frequency domain feature map of an infrared image of a moving bird. This is an element-wise addition;

[0221] Slow motion characteristics are obtained through inverse fast Fourier transform.

[0222] like Figure 5 As shown, the rapid motion feature extraction step includes:

[0223] The input sequence of infrared bird motion target feature images is stacked along the time dimension to form a three-dimensional tensor. A one-dimensional fast Fourier transform is then performed on the time dimension to obtain the frequency map of the pixels in time. A frequency energy map is constructed by focusing on the strong responses generated in the high-frequency regions of the frequency map, and the expression is:

[0224] ;

[0225] in, The three-dimensional tensor formed by stacking infrared bird motion target feature image sequences along frame number time t is located at coordinates... The value at that location, for The index f obtained after performing a one-dimensional fast Fourier transform on t is located at coordinates... Frequency response at that point, for The coordinates in the cumulative energy map of the high-frequency region are: The value at that location, , For high-frequency starting index, ρ is the proportionality constant. The maximum frequency is N, which represents N frames of infrared bird moving target feature images. In this embodiment, N=5 is preferred.

[0226] To further utilize frequency domain structure information, the frequency energy map is used as a mask to guide the frequency domain feature analysis of the image to be detected; first, a two-dimensional fast Fourier transform is performed on each frame to obtain... Mapping the time-frequency plot to frequency domain guided weights, and then normalizing and... Element-wise multiplication is performed, taking advantage of the high energy distribution near the direction of movement in terms of time frequency of fast-moving targets, thereby enhancing the spatial frequency domain characteristics of bird-like moving targets. The expression is as follows:

[0227] ;

[0228] in, Let be the enhanced spatial frequency domain features of the t-th frame, and Norm be the normalization operation;

[0229] The enhanced spatial frequency domain features are further strengthened by using inter-frame difference. Pixel aggregation of the difference image yields the fast motion frequency domain features, expressed as:

[0230] ;

[0231] in, The value at frequency domain coordinates (u, v) is taken from the fast motion frequency domain feature map of an infrared bird moving target feature image. Let be the value at frequency domain coordinates (u, v) in the frequency domain residual feature map of the infrared bird moving target feature image of frame t. For convolution operations with a kernel size of 3×3, The value at frequency domain coordinates (u,v) is taken in the enhanced spatial frequency domain feature map of the infrared bird moving target feature image of frame t+1.

[0232] Finally, the rapid motion characteristics are obtained through inverse fast Fourier transform.

[0233] like Figure 6 As shown, based on the differences in motion energy distribution in the infrared image after background suppression, the weights of the stationary feature extraction branch, the slow motion feature extraction branch, and the fast motion feature extraction branch are determined to obtain the fusion weights that satisfy the normalization constraint. The steps include:

[0234] For an input N-frame continuous infrared image containing a moving bird target after background suppression, optical flow is used to calculate motion energy to measure motion amplitude, expressed as:

[0235] ;

[0236] Where E represents motion energy, and N represents N frames of infrared images of moving birds. The infrared bird moving target feature images of frame t and frame t-1 are located at coordinates... The optical flow amplitude is calculated at each pixel, where H is the height of the infrared bird moving target feature image and W is the width of the infrared bird moving target feature image; preferably, N=5.

[0237] The static features, slow motion features, and fast motion features are concatenated with motion energy and input into a lightweight multilayer perceptron. The weights of the static feature extraction branch, slow motion feature extraction branch, and fast motion feature extraction branch are obtained through nonlinear transformation and normalized. The expression is as follows:

[0238] ;

[0239] in, Let be the weight of the i-th branch, i = 1, 2, 3; Concat is the connection operation; MLP is the multilayer perceptron; and Softmax is the normalization operation. It is a static feature. Characterized by slow motion. Characterized by rapid motion;

[0240] The feature maps of the three branches are linearly fused based on weights. To ensure the stability of the distribution of the fused features, channel normalization and lightweight convolution are performed to finally obtain the fused feature map, expressed as:

[0241] ;

[0242] in, To fuse feature maps, Extract branch weights for static features. Extracting branch weights for slow motion features. Extracting branch weights for fast motion features;

[0243] The fused feature map is used to obtain infrared bird motion targets through YOLOX's decoupling head and non-maximum suppression.

[0244] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0245] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0246] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0247] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0248] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for detecting moving infrared birds based on background learning and adaptive motion feature fusion, characterized in that, include: Acquire a continuous sequence of infrared images containing moving birds; Background suppression processing was performed on the acquired continuous infrared image sequence containing moving bird targets using images of salient background regions; A continuous infrared image sequence containing moving birds, after background suppression processing, is input into a pre-trained infrared moving target detection network. Static features, slow motion features, and fast motion features are extracted and fused to obtain infrared moving birds. The background salient region image is continuously updated during the training and detection phases of the infrared moving target detection network; The training of the infrared moving target detection network includes background suppression processing of the dataset using images of salient background regions, and training using the background-suppressed dataset.

2. The infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to claim 1, characterized in that, The construction of the background salient region image includes: Collect M consecutive infrared image sequences containing moving bird targets to obtain sample images; The local brightness characteristics of each pixel location are calculated in the local neighborhood of the collected sample images, and a preset background containing salient regions is obtained through regional connectivity. By using inter-frame difference to remove foreground targets from a preset background and selecting the median brightness of each pixel, an image of the salient background region is obtained.

3. The infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to claim 2, characterized in that, The step of calculating the local brightness characteristics of the local neighborhood of each pixel location based on the acquired sample image, and obtaining the preset background containing salient regions through region connectivity includes: The local brightness characteristics within the local neighborhood of each pixel location in the M-frame sample image are calculated using the following formula: ; in, For the sample image of frame t, let the coordinates be... The pixel brightness at coordinates (i,j) within a local neighborhood centered on the pixel. Coordinates are The local average brightness of the pixels. Coordinates are The local brightness standard deviation of the pixels, Represented by coordinates The local neighborhood range centered on the pixel, Indicates the number of pixels within a local neighborhood; The brightness threshold is calculated based on local brightness characteristics, and its expression is: ; in, The coordinates of the sample image in frame t are The brightness threshold of the pixel, where к is the adjustment parameter used to adjust the local brightness threshold; When the pixel brightness of the t-th frame sample image satisfies When this happens, the pixel is identified as a highlight pixel. The coordinates of the sample image in frame t are Pixel brightness; Connectivity analysis and edge structure constraints are performed on all bright pixels to divide the sample image into several spatially independent salient structure regions. The brightness of pixels in the salient structure regions is retained, while the brightness of pixels in the remaining regions is set to 0, thus generating a preset background for each frame of the sample image.

4. The infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to claim 3, characterized in that, The step of removing foreground targets from a preset background using inter-frame difference and selecting the median brightness of each pixel to obtain an image of the salient background region includes: The preset background is subjected to inter-frame differencing to obtain a differencing image. Foreground targets are then removed by accumulating pixel fluctuations. The expression is as follows: ; in, The instantaneous change intensity of the background is preset for the sample image of frame t. The cumulative change intensity of the preset background. The preset background for the sample image of frame t, The preset background for the sample image of frame t-1; when or When a pixel is determined to be a foreground target, it is removed from the preset background. The median pixel brightness of each position in the preset background after removing the foreground target from M frames is taken to obtain the image of the salient background region.

5. The infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to claim 2, characterized in that, The update of the background salient region image includes: Define the image update conditions for background salient regions, including rate constraints and time consistency constraints; When both the rate constraint and the time consistency constraint are satisfied, the background salient region image is updated using the background information of the current frame image. Specifically, during the training phase of the infrared moving target detection network, the background salient region image is updated based on the background information of the training images labeled with real tags; during the detection phase of the infrared moving target detection network, the background salient region image is updated based on the background information of the detection result image of the infrared moving target detection network.

6. The infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to claim 5, characterized in that, The image update conditions for defining the salient background region include: Calculate the local average brightness, local brightness standard deviation, and brightness threshold of the image in the background salient region; The state probability matrix is ​​constructed based on the background salient region image, and its expression is: ; in, The coordinates in the state probability matrix are The state probability value of the pixel. The coordinates of the background salient region in the image are The brightness threshold of the pixels, The coordinates of the background salient region in the image are Pixel brightness; The rate constraint is defined based on the local average brightness of the image in the salient background region, and the expression is: ; in, The coordinates of the sample image in frame t are The rate of change of background brightness in the region of pixels, The coordinates of the image obtained after masking the foreground target in the training image or detection result image of frame t are: The local average brightness of the pixels. The image coordinates of the background salient region in frame t-1 are: The local average brightness of the pixels; when When the rate constraint condition is met, ε is the threshold for the change in regional background brightness; Based on the state probability matrix, the time consistency constraint is defined as follows: ; in, The coordinates of the sample image in frame t are The state probability value of the pixel. The coordinates of the sample image in frame t-1 are The state probability value of the pixel. These are state correction parameters, used to correct the probability values ​​of each pixel in the state probability matrix that belong to the background; when When the time consistency constraint is satisfied, η is the state probability threshold; And / or, when both the rate constraint and the time consistency constraint are satisfied, the expression for updating the image of the background salient region is: ; in, The image coordinates of the background saliency region after the update in frame t are: pixel brightness, The image coordinates of the background salient region in frame t-1 are: The pixel brightness is α, which is the learning rate parameter used to control the contribution of the current frame's background information to the update of the image of the salient background region.

7. The infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to claim 2, characterized in that, The background suppression processing steps include: Calculate the local brightness characteristics of the updated background salient region image; The background deviation of the training image or the detection image relative to any pixel within the background salient region is calculated based on the local brightness characteristics of the updated background salient region image. The formula is as follows: ; in, This represents the coordinates of the training or detection image relative to the background saliency region. The degree of pixel deviation, The coordinates of the t-th frame of the image to be trained or detected are: pixel brightness, The image coordinates of the background saliency region after the update in frame t are: pixel brightness, The updated background saliency region image coordinates are: The standard deviation of local brightness of the pixels; The background confidence score is constructed based on the background deviation, and the expression is: ; in, Background confidence level; The background confidence map is mapped to a suppression weight map, and the intermediate features are reweighted using a background suppression function during the network feature extraction stage. The expression is as follows: ; in, The coordinates of the training image or the detection image after background suppression are: pixel brightness, This is a pixel-by-pixel multiplication.

8. The infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to claim 1, characterized in that, The infrared moving target detection network's detection steps for infrared bird moving targets include: Infrared bird motion target feature images are extracted from a continuous infrared image sequence containing bird motion targets after background suppression; The extracted infrared bird moving target feature images are input into the stationary feature extraction branch, the slow motion feature extraction branch, and the fast motion feature extraction branch, respectively, to extract stationary features, slow motion features, and fast motion features. The extracted static features, slow-moving features, and fast-moving features are input into the adaptive weight fusion module for weighted fusion, and infrared bird moving targets are obtained through the decoupling head.

9. The infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to claim 8, characterized in that, The static feature extraction step includes: The phase response of the local structure of a stationary target is analyzed using a multi-scale parity-symmetric filtering convolution kernel. The expression for the parity-symmetric filtering convolution kernel is as follows: ; in, For infrared bird moving target feature images, the odd-even symmetric filtering convolution kernel is used in coordinates. The values ​​of θ are given by k, where k is the index scale and θ is the direction. The coordinates of the infrared bird moving target feature image are: The even-symmetric component of the pixel, The coordinates of the infrared bird moving target feature image are: The odd symmetric component of the pixel, where i is the imaginary unit; The complex domain response is constructed based on the odd-symmetric and even-symmetric components, and its expression is: ; in, The coordinates of the infrared bird moving target feature image are: The even-symmetric response of the pixels, The coordinates of the infrared bird moving target feature image are: The odd-symmetric response of the pixels, For convolution operations, The coordinates of the infrared bird moving target feature image input to the static feature extraction branch are: eigenvalues; Using the phase consistency model, the local structural features of a stationary target are enhanced, as expressed by: ; in, The coordinates in the feature map obtained after applying an even-odd symmetric filter convolution kernel to an infrared image of moving birds are: The phase consistency of the pixels, The coordinates in the feature map obtained after applying an even-odd symmetric filter convolution kernel to an infrared bird moving target feature image are: The amplitude of the pixels, The coordinates in the feature map obtained after applying an even-odd symmetric filter convolution kernel to an infrared bird moving target feature image are: The phase of the pixel, The coordinates of the infrared bird moving target feature image are: The initial phase of the pixel; The initial phase expression is: ; Where atan2 is the arctangent function; Pixel integration using 1×1 convolution yields static features, expressed as follows: ; in, It is a static feature. For convolution operations, The coordinates of the infrared bird moving target feature images in frame t and frame t-1 are: The phase difference of the pixels, ReLU is the activation function, and BN is the normalization layer. To multiply element by element, N represents N frames of infrared images of moving bird targets; And / or, the slow motion feature extraction step includes: A two-dimensional fast Fourier transform and phase enhancement are performed on the input infrared image of moving birds to obtain the enhanced frequency domain features, expressed as follows: ; in, The coordinates of the infrared bird moving target feature images in frame t and frame t-1 are: The phase difference of the pixels, It is a phase difference amplifier. The coordinates of the infrared bird moving target feature image in frame t are... The phase of the pixel, The coordinates of the infrared bird moving target feature image in frame t+1 are: The phase of the pixel, This refers to the pixel features with frequency domain coordinates (u,v) after frequency domain enhancement of an infrared image of moving birds. The pixel features with frequency domain coordinates (u,v) after the two-dimensional fast Fourier transform of an infrared image of moving birds are represented by the pixel features, where j is the imaginary unit. The pixels of the infrared bird moving target feature image in frame t are represented by spatial domain coordinates. The phase transformed to frequency domain coordinates (u,v); Calculate the local phase gradient to capture local structural changes in slow-moving targets, and obtain local structural features, expressed as: ; in, The coordinates of the infrared bird moving target feature image in frame t are... The phase gradient of the pixel, The local structural features of pixels with frequency domain coordinates (u,v) in an infrared image of moving birds; The enhanced frequency domain features and local structural features are fused together to further extract the frequency domain features of slow motion, as expressed in the following expression: ; in, The feature value at coordinates (u, v) is the slow motion frequency domain feature map of an infrared image of a moving bird. This is an element-wise addition; Slow motion characteristics are obtained through inverse fast Fourier transform; And / or, the rapid motion feature extraction step includes: The input sequence of infrared bird moving target feature images is transformed into a frequency map over time. A frequency energy map is then constructed based on the frequency map, expressed as: ; in, The three-dimensional tensor formed by stacking infrared bird motion target feature image sequences along frame number time t is located at coordinates... The value at that location, for The index f obtained after performing a one-dimensional fast Fourier transform on t is located at coordinates... Frequency response at that point, for The coordinates in the cumulative energy map of the high-frequency region are: The value at that location, , This is the high-frequency starting index, and ρ is the proportionality coefficient. The maximum frequency is N, where N represents N frames of infrared bird moving target feature images; Mapping the frequency energy map to frequency domain guided weights enhances the spatial frequency domain characteristics of bird-moving targets; the expression is as follows: ; in, Let be the enhanced spatial frequency domain features of the t-th frame, and Norm be the normalization operation; The enhanced spatial frequency domain features are further strengthened by using inter-frame difference. Pixel aggregation of the difference image yields the fast motion frequency domain features, expressed as: ; in, The value at frequency domain coordinates (u, v) is taken from the fast motion frequency domain feature map of an infrared bird moving target feature image. Let be the value at frequency domain coordinates (u, v) in the frequency domain residual feature map of the infrared bird moving target feature image of frame t. For convolution operations with a kernel size of 3×3, The value at frequency domain coordinates (u,v) is taken in the enhanced spatial frequency domain feature map of the infrared bird moving target feature image of frame t+1. Rapid motion characteristics are obtained through inverse fast Fourier transform.

10. The infrared bird moving target detection method based on background learning and adaptive motion feature fusion according to claim 8, characterized in that, The extracted stationary features, slow-moving features, and fast-moving features are input into the adaptive weight fusion module for weighted fusion. The infrared bird motion targets obtained through the decoupling head include: For an input N-frame continuous infrared image containing a moving bird target after background suppression, optical flow is used to calculate motion energy to measure motion amplitude, expressed as: ; Where E represents motion energy, and N represents N frames of infrared images of moving birds. The infrared bird moving target feature images of frame t and frame t-1 are located at coordinates... The optical flow amplitude is calculated at each pixel, where H is the height of the infrared bird moving target feature image and W is the width of the infrared bird moving target feature image. The static features, slow motion features, and fast motion features are concatenated with motion energy. A multilayer perceptron is then used to obtain the weights of the static feature extraction branch, the slow motion feature extraction branch, and the fast motion feature extraction branch, which are then normalized. The expression is as follows: ; in, Let be the weight of the i-th branch, where i = 1, 2, 3, Concat is the connection operation, MLP is the multilayer perceptron, and Softmax is the normalization operation. It is a static feature. Characterized by slow motion. Characterized by rapid motion; The fused feature map is obtained after channel normalization and lightweight convolution, expressed as: ; in, To fuse feature maps, For convolution operations, Norm is for normalization. Extract branch weights for static features. Extracting branch weights for slow motion features. Extract branch weights for fast motion features; The fused feature map is used to obtain infrared bird motion targets through YOLOX's decoupling head and non-maximum suppression.