An airport bird repelling detection method based on multi-sensor fusion

Through multi-sensor fusion and deep learning models, the airport bird repelling system is realized, and the problems of limited detection range of a single sensor and insufficient fusion of multi-source data are solved, improving detection accuracy and bird repelling success rate.

CN120009877BActive Publication Date: 2025-08-01JIANGSU JINGWEI ZHILIAN AVIATION TECH CO LTD

Patent Information

Application Number
CN202510507429.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-01
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing airport bird reaping system relies on a single sensor, has a limited detection range, is susceptible to environmental interference, lacks multi-source data fusion, lacks intelligent decision-making mechanisms, and bird reaping measures lack initiative and flexibility, making it difficult to achieve real-time monitoring and accurate response.

Method used

A multi-sensor fusion method is adopted, including dynamic vision sensors, millimeter-wave radar, infrared thermal imaging sensors and microphone arrays, combined with the improved YOLOv8 model and Transformer model, accurate identification of bird targets, behavioral pattern analysis and dynamic threat assessment, through time synchronization and spatial registration, multimodal feature fusion, dynamic adjustment of weights to calculate comprehensive confidence, and dynamic selection of bird repelling measures.

Benefits of technology

It has achieved accurate and reliable bird monitoring all-weather, improved detection accuracy, shortened response time, improved bird repelling success rate, and improved system intelligence level, solving the problems of inaccurate detection, lag in response and poor bird repelling effects in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120009877B_ABST
    Figure CN120009877B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of airport bird repelling and detection, and in particular to an airport bird repelling and detection method based on multi-sensor fusion. Through the collaborative work of a dynamic vision sensor, a millimeter-wave radar, an infrared thermal imaging sensor, and a microphone array, combined with an improved object detection model and a multi-modal feature fusion algorithm, it realizes the accurate identification of bird targets, the analysis of behavior patterns, and the dynamic threat assessment, thereby significantly improving the real-time performance, accuracy, and intelligent level of airport bird repelling and detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of airport bird repelling and detection, and particularly to an airport bird repelling and detection method based on multi-sensor fusion. Background Art

[0002] With the rapid development of the air transportation industry and the continuous expansion of airport scale, airport safety issues have attracted increasing attention. Among them, bird strikes on aircraft have become one of the important factors threatening flight safety. According to statistics, bird strike incidents can not only cause serious damage to aircraft, but also trigger major safety accidents, bringing huge economic losses and safety hazards to airlines, airport operators and passengers. Therefore, how to effectively detect and drive away birds has become a key issue in airport safety management.

[0003] There are still many deficiencies in the existing technologies for airport bird repelling and detection, which are specifically manifested as follows:

[0004] 1) The current airport bird repelling system mainly relies on a single type of sensor (such as radar or video surveillance) for bird detection. Although these sensors have certain advantages in specific scenarios, their detection ranges are limited and they are easily interfered by environmental factors (such as weather, lighting conditions, etc.). For example, the performance of millimeter-wave radar deteriorates under complex meteorological conditions, and it is difficult for a single sensor to comprehensively capture multi-dimensional information such as the position, speed, and attitude of birds, resulting in inaccurate and incomplete detection results.

[0005] 2) The airport bird repelling system usually involves data collection from multiple sensors (such as radar, infrared), but there are obvious shortcomings in the integration and processing of multi-source heterogeneous data in the existing technologies. The data formats generated by different sensors are diverse, and the data collection frequencies and accuracies are also different, resulting in a serious data island phenomenon. The existing technologies often lack effective data fusion algorithms and cannot achieve the efficient integration and collaborative analysis of multi-source data, thus affecting the overall performance of the bird repelling system.

[0006] 3) Traditional bird repelling methods mostly rely on manual observation and empirical judgment, with low efficiency and difficulty in coping with complex flight environments. In recent years, although some automated monitoring means (such as rule-based early warning systems) have been introduced, due to the lack of support from advanced intelligent algorithms, these systems have large delays in data processing and analysis, and it is difficult to achieve real-time monitoring and rapid response. In addition, the existing technologies have insufficient intelligence levels in predicting bird behavior patterns and evaluating threat levels, and cannot formulate precise bird repelling strategies according to the dynamic trajectories of birds and environmental changes;

[0007] 4) Current bird repellent measures (such as playing the sounds of natural enemies, using laser bird repellers, etc.) are mostly passive responses, lacking initiative and flexibility. These measures often start when the flying birds are already approaching the dangerous area, and fail to make full use of early detection data for preventive intervention. At the same time, the selection and implementation of bird repellent measures usually rely on fixed rules and lack a dynamic adjustment mechanism, resulting in poor bird repellent effects. Summary of the Invention

[0008] The present invention discloses an airport bird repellent detection method based on multi-sensor fusion, which solves problems in the prior art such as limited detection ability of a single sensor, insufficient multi-source data fusion, and lack of an intelligent decision-making mechanism. Through the collaborative work of a dynamic vision sensor, a millimeter-wave radar, an infrared thermal imaging sensor, and a microphone array, combined with an improved object detection model and a multi-modal feature fusion algorithm, accurate identification of flying bird targets, analysis of behavior patterns, and dynamic threat assessment are realized, thereby significantly improving the real-time performance, accuracy, and intelligence level of the airport bird repellent system.

[0009] In order to achieve the purpose of the present invention, the technical solution adopted is: an airport bird repellent detection method based on multi-sensor fusion, including:

[0010] S1. Deploy a camera equipped with a dynamic vision sensor, a millimeter-wave radar, an infrared thermal imaging sensor, and a microphone array in the target area;

[0011] S2. Perform time synchronization and spatial registration on the data collected by the camera equipped with a dynamic vision sensor, the millimeter-wave radar, the infrared thermal imaging sensor, and the microphone array respectively;

[0012] S3. Input the visual data processed in step S2 into an improved YOLOv8 model to identify and detect flying bird targets, and output the position and attitude information of the flying birds; the millimeter-wave radar data includes the distance, speed, and movement trajectory of the flying birds; the infrared data outputs infrared features to assist in confirming the presence of the flying birds; extract the sound features of the flying birds' calls through the microphone array to supplement the confirmation of the types and behavior patterns of the flying birds;

[0013] S4. Based on the position and attitude information of the flying birds output by the improved YOLOv8 model, extract spatial features; based on the distance, speed, and trajectory information of the flying birds, extract dynamic features; based on the infrared features of the infrared thermal imaging sensor, extract thermal features; based on the sound features of the flying birds' calls, extract behavior pattern features; input the extracted features into a Transformer model to generate a joint feature vector, and dynamically adjust the weight to calculate the comprehensive confidence according to the environmental perception, and judge whether to start the bird repellent measures according to a preset threshold;

[0014] S5. Predict the threat level according to the position and trajectory of the flying birds, and dynamically select bird repellent actions.

[0015] As an optimization solution of the present invention, in step S1, a camera equipped with a dynamic vision sensor perceives the dynamic behavior of the bird in real time by capturing the brightness changes caused by the movement of the bird, generates sparse event stream data, and obtains the morphological features of the bird to assist target recognition.

[0016] As an optimization solution of the present invention, in step S2, the camera, millimeter-wave radar, infrared thermal imaging sensor, and microphone array generate an initial timestamp when collecting data, calibrate the timestamp using the global clock source NTP protocol, measure and compensate for the time offset, and interpolate the data based on the sampling frequency to map it to a unified time axis.

[0017] As an optimization solution of the present invention, the improved YOLOv8 model is as follows: DCNv3 is used to replace the standard convolutional layer in Backbone; the dynamic pyramid network is used to replace the feature pyramid network contained in Neck; and the adaptive bounding box regression is used to replace the bounding box regression in Head.

[0018] As an optimization solution of the present invention, the visual data processed in step S2 is input into the improved YOLOv8 model to identify and detect the flying bird target, and output the position and posture information of the flying bird, including:

[0019] 1) DCNv3 allows the convolution kernel to adaptively adjust the sampling position according to the input content by dynamically learning the offset and modulation factor. The specific formula is:

[0020] ;

[0021] Where: y(p) is the value of the output feature map at position p, x is the input feature map, K is the convolution kernel size, w k is the convolution kernel weight, Δp k is the offset of dynamic learning, Δm k is the modulation factor;

[0022] 2) The dynamic pyramid network dynamically weights the multi-scale feature maps output by Backbone to generate preliminary multi-scale feature maps; for each pair of feature maps C i and C j , calculate the dynamic weight α between them ij , the formula is:

[0023] α ij = Softmax(Conv light (C i ⊕C j ));

[0024] Where: C i and Cj is the input feature map; C i ⊕C j represents feature concatenation; Conv light is a lightweight convolution for generating weights; Softmax normalization ensures that the sum of the weights is 1; According to the generated dynamic weight α ij , weighted fusion of multi-scale feature maps is performed to generate a preliminary multi-scale feature map, and the weighted fusion formula is:

[0025] ;

[0026] In the formula: P i is the fused feature map, and Resize is to adjust C j to the same resolution as P i ;

[0027] 3) Use adaptive bounding box regression to accurately locate the position of the flying bird, and finally output the position and pose information of the flying bird.

[0028] As an optimization scheme of the present invention, the flying bird call data is used to supplement and confirm the species and behavior patterns of the flying bird, including:

[0029] (1) Extract Mel Frequency Cepstral Coefficients MFCC from the collected flying bird call data;

[0030] (2) Train an LSTM classifier, with the input being the MFCC sequence, and output the class probability through the classification layer:

[0031] ;

[0032] Among them: X is the MFCC feature sequence, W c1 is the classification layer weight matrix, b c1 is the classification layer bias vector, and p(Class=c|X) represents the probability that the class is c under the condition of X.

[0033] As an optimization scheme of the present invention, in step S4, dynamically adjusting the weight to calculate the comprehensive confidence through environmental perception is specifically:

[0034] ;

[0035] Among them: C final is the comprehensive confidence, α i is the dynamic weight of the i-th modality; W c is the learnable weight matrix, b c is the bias term, H joint is the joint feature vector, and the Sigmoid function is an activation function.

[0036] As an optimized solution of the present invention, in step S5,

[0037] The threat level formula is: ;

[0038] Where: ThreatLevel is the danger level value, C fina is the comprehensive confidence level, d min is the shortest Euclidean distance from the bird flight trajectory to the runway, and σ is the attenuation coefficient.

[0039] As an optimized solution of the present invention, the bird repelling actions include low-intensity acoustic wave bird repelling, medium-intensity laser bird repelling, and high-intensity UAV interception.

[0040] The present invention has positive effects: 1) The present invention overcomes the limitations of a single sensor in extreme weather (rain / fog / night) through the four-modal fusion of a dynamic vision sensor, a millimeter-wave radar, an infrared thermal imager, and a microphone array. The detection accuracy of the four-sensor fusion scheme reaches 89.5% in foggy nights, achieving all-weather accurate and reliable monitoring;

[0041] 2) The improved YOLOv8 model of the present invention combines DCNv3 convolution and a dynamic pyramid network, enabling the mAP of bird detection to reach 93.1%, a 36.5% improvement compared to traditional methods. Adaptive bounding box regression reduces the small target positioning error by 42%, especially suitable for fast-moving birds at a long distance;

[0042] 3) Based on the multi-modal feature fusion of Transformer and the adjustment of environmental perception weights in the present invention, the data fusion system breaks through the technical bottleneck of data islands in traditional methods. Through the spatio-temporal alignment mechanism and the dynamic weighted fusion algorithm, the deep integration of multi-source heterogeneous data from the physical layer to the decision layer is achieved, shortening the average response time of the system to 85 ms and increasing the bird repelling success rate to 94%;

[0043] 4) Through the multi-modal sensor fusion and deep learning model optimization in the present invention, precise bird repelling strategies are formulated according to the dynamic trajectory of birds and environmental changes, changing the passive response mode of traditional bird repelling measures, constructing an intelligent airport bird repelling method, and effectively solving the problems of performance degradation in extreme weather, insufficient multi-source data fusion, and response lag in traditional methods. It provides a reliable technical guarantee for airport safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0045] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The method of the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0047] As Figure 1 shown, the present invention discloses an airport bird repelling detection method based on multi-sensor fusion, including the following steps:

[0048] S1. Deploy a camera equipped with a dynamic vision sensor, a millimeter-wave radar, an infrared thermal imaging sensor, and a microphone array in the target area;

[0049] The camera equipped with a dynamic vision sensor captures the brightness changes caused by the movement of flying birds, real-time perceives the dynamic behavior of flying birds, and generates sparse event stream data; the camera also obtains the morphological characteristics of flying birds to assist in target recognition, and obtains visual data through the camera equipped with a dynamic vision sensor. The dynamic vision sensor (DVS) can capture the events of brightness changes in the scene. Deploy a camera equipped with a dynamic vision sensor in the target area to real-time perceive the dynamic behavior of flying birds. The high temporal resolution characteristic of the dynamic vision sensor enables it to operate stably under complex lighting conditions (such as strong light or weak light environment). The camera is also used to obtain the morphological characteristics of flying birds (such as wing shape, body length, etc.) to assist in target recognition and classification.

[0050] The millimeter-wave radar calculates the distance and speed of the target by emitting electromagnetic waves and receiving echo signals. The millimeter-wave radar can work stably under complex weather conditions (such as rain, fog, etc.), and provides reliable data for the precise positioning and trajectory prediction of flying birds.

[0051] The infrared thermal imaging sensor is used to capture the heat distribution information of flying birds, assist in confirming the presence of flying birds, and distinguish flying birds from other non-biological targets (such as drones, bird models, etc.). The infrared thermal imaging sensor generates a thermal imaging map by detecting the infrared radiation intensity emitted by the target. At night or under low light conditions, the infrared thermal imaging sensor can effectively supplement the deficiencies of the camera and ensure all-weather monitoring capabilities.

[0052] The microphone array is used to capture the sound signals of flying birds' calls, extract sound features to confirm the species and behavior patterns of flying birds. The microphone array simultaneously collects sound signals through multiple microphones, uses beamforming technology to enhance the sound signals in the target direction, and extracts features such as the Mel-frequency cepstral coefficients (MFCC), call interval, and duration of flying birds' calls, for analyzing the species and behavior patterns of flying birds (such as foraging, warning, etc.).

[0053] Multiple sensors work collaboratively. The dynamic vision sensor, millimeter-wave radar, infrared thermal imaging sensor, and microphone array are each responsible for different data acquisition tasks, but their data is complementary. The dynamic vision sensor and infrared thermal imaging sensor are respectively suitable for monitoring needs during the day and at night, while the millimeter-wave radar can operate under complex weather conditions to ensure the system has all-weather monitoring capabilities. By fusing the data collected by each sensor, more comprehensive bird information can be generated to support subsequent intelligent decision-making (such as whether to initiate bird repelling measures).

[0054] S2. Perform time synchronization and spatial registration on the data collected by the camera equipped with a dynamic vision sensor, millimeter-wave radar, infrared thermal imaging sensor, and microphone array respectively.

[0055] In step S2, when the camera, millimeter-wave radar, infrared thermal imaging sensor, and microphone array collect data, they generate initial timestamps. Use the Network Time Protocol (NTP) of the global clock source to calibrate the timestamps, measure and compensate for the time offset, and perform interpolation processing on the data based on the sampling frequency, so as to map it onto a unified time axis.

[0056] 1. Initial timestamp generation;

[0057] Each sensor generates a local timestamp t local,i when collecting data, indicating the time when the sensor collects a certain data point. Since the clocks of each sensor may have deviations, these timestamps need to be calibrated and aligned. Among them: t local,i = t sensor,i , t sensor,i is the local time of the i-th sensor; i represents different sensors (camera, millimeter-wave radar, etc.).

[0058] To ensure the alignment of the timestamps of all sensors, use the Network Time Protocol (NTP) as the global clock source to calibrate the timestamps. The specific steps are as follows:

[0059] (1) Measure the time offset: Each sensor synchronizes its time with the global clock source (NTP server) and calculates its time offset Δt i : Δt i = t NTP - t local,i , where t NTP is the global time obtained from the NTP server; t local,i is the local time of the sensor.

[0060] (2) Compensate for the time offset: Adjust the local timestamp of the sensor to the global timestamp: t global,i = t local,i + Δt i , where, tglobal,i is the calibrated global timestamp.

[0061] (3) Construct a unified timeline: Define a unified time step ΔT and construct a timeline T unified . T unified = {t0, t0 + ΔT, t0 + 2ΔT, …, t end}, where: t0 is the start time and t end is the end time.

[0062] (4) Data interpolation: Perform linear interpolation on the data of each sensor on the unified timeline to ensure that all data points are aligned.

[0063] Calibrate the timestamp through the NTP protocol to ensure that all sensor data is based on the same time reference. By interpolating the data, map data with different sampling frequencies to the unified timeline, facilitating subsequent multi-modal feature fusion and analysis. Linear interpolation can effectively handle the time deviation problem of sensor data.

[0064] S3. Input the visual data processed in step S2 into the improved YOLOv8 model to identify and detect bird targets, and output the position and pose information of the birds; the millimeter-wave radar data includes the distance, speed, and movement trajectory of the birds; the infrared data outputs infrared features for assisting in confirming the presence of the birds; by extracting the sound features of the bird calls, it is used to supplement and confirm the species and behavior patterns of the birds.

[0065] Among them: the improved YOLOv8 model is: Replace the standard convolutional layer in the Backbone with DCNv3; replace the Feature Pyramid Network included in the Neck with a dynamic pyramid network; replace the bounding box regression in the Head with adaptive bounding box regression.

[0066] Input the visual data processed in step S2 into the improved YOLOv8 model to identify and detect bird targets, and output the position and pose information of the birds, including:

[0067] 1) DCNv3 allows the convolutional kernel to adaptively adjust the sampling position according to the input content by dynamically learning the offset and modulation factor. The specific formula is: [[ID=3^2]]

[0068] ;

[0069] Among them: y(p) is the value of the output feature map at position p, x is the input feature map, K is the convolutional kernel size, w k is the weight of the k-th convolutional kernel, Δp k is the dynamically learned offset, and Δm k is the modulation factor.

[0070] When dealing with object detection tasks, the standard convolutional layer has difficulty capturing complex geometric transformations (such as rotation, deformation, etc.), especially when the morphological changes of targets such as flying birds are large. The Deformable Convolution (DCN) can adaptively adjust the receptive field of the convolutional kernel by introducing offsets, so as to better adapt to the complex pose changes of flying bird targets (such as different wing spreading angles). It can improve the detection ability of flying bird targets, especially in long-distance scenarios.

[0071] 2) The dynamic pyramid network performs dynamic weight fusion on the multi-scale feature maps {C3, C4, C5} output by the Backbone to generate preliminary multi-scale feature maps {P3, P4, P5}; for each pair of feature maps C i and C j , calculate the dynamic weight α ij , and the formula is:

[0072] α ij = Softmax(Conv light (C i ⊕C j ));

[0073] Among them: C i and C j are the input feature maps; C i ⊕C j represents feature concatenation; Conv light is a lightweight convolution used to generate weights; Softmax normalization ensures that the sum of the weights is 1; according to the generated dynamic weight α ij , perform weighted fusion on the feature maps to generate preliminary multi-scale feature maps {P3, P4, P5}, and the weighted fusion formula is:

[0074] ;

[0075] Among them: P i is the fused feature map, and Resize is to adjust C j to the same resolution as P i .

[0076] The Neck part of YOLOv8 usually uses the Feature Pyramid Network (FPN) to fuse multi-scale features. However, the static fusion method of FPN may not fully exploit the relationship between features of different scales. The multi-scale feature maps output by the Backbone are {C3, C4, C5}, where each feature map represents feature information at different resolutions. For example, C3 corresponds to lower-level features, with higher resolution and more detailed information; C4 and C5 correspond to medium and high-level features respectively, with decreasing resolution but richer semantic information. The Dynamic Pyramid Network can adaptively adjust the weights of each feature map according to the input data by introducing a dynamic weight mechanism, thus better capturing the relationship between features of different scales. Since the Dynamic Pyramid Network can fuse multi-scale features more precisely, the model performs better when dealing with complex scenarios (such as large variations in the poses of flying birds and different target sizes). Especially in small object detection, the dynamic weight mechanism helps to enhance the importance of low-resolution feature maps, thereby improving the overall detection accuracy. The Dynamic Pyramid Network enhances the robustness of the model to various environmental conditions (such as changes in lighting and weather effects) by dynamically adjusting the importance of feature maps of different scales.

[0077] 3) Use adaptive bounding box regression to accurately locate the position of the flying bird, and combine the classification head to determine the presence and category of the flying bird, and finally output the position and pose information of the flying bird.

[0078] The overall process of adaptive bounding box regression:

[0079] Input: Feature maps {P3, P4, P5}: Multi-scale feature maps generated by the Dynamic Pyramid Network. Candidate regions (Anchor Boxes): Initially estimated target positions.

[0080] Output: Accurate bounding box position (x, y, w, h) of the flying bird, class label of the flying bird.

[0081] Step1: Generate initial candidate boxes;

[0082] Use predefined Anchor Boxes to generate candidate boxes on the multi-scale feature maps. Each candidate box is represented by initial coordinates: (x init , y init , w init , h init )

[0083] where: x init , y init are the horizontal and vertical coordinates of the center point of the initial candidate box; w init , h init are the width and height of the initial candidate box.

[0084] Step 2: Adaptive bounding box regression; for each candidate box, its position is refined through a regression branch. The coordinates of the adjusted bounding box are represented as: (x, y, w, h);

[0085] Regression formula: x = x init + t x • w init ; y = y init + t y • h init ; ; ;

[0086] where: t x , t y represent the offsets of the center point of the bounding box relative to the initial candidate box, which are the values predicted by the regression branch and are used to fine-tune the position of the center point, and ensure that the width and height are always positive values.

[0087] To enhance the adaptability of the model to objects of different scales, adaptive weights are introduced to adjust the output of the regression branch: t x = β x • f x (P i ); t y = β y • f y (P i ); t w = β w • f w (P i ); t h = β h • f h (P i ); where: f x (P i ), f y (P i ), f w (P i ), f h (P i ) are the features extracted from the feature map P i ; β x , β y , β w , β h are learnable adaptive weights used to dynamically adjust the regression accuracy in different directions and scales. β x controls the weight of the horizontal direction (x-axis) offset, β y controls the weight of the vertical direction (y-axis) offset, β w controls the weight of the width adjustment, βh Control the weight of height adjustment. f x (P i ) Features extracted from the feature map P i are used to predict the offset in the horizontal direction (x-axis); f y (P i ) Features extracted from the feature map P i are used to predict the offset in the horizontal direction (y-axis); f w (P i ) is the scaling factor used to predict the width, f h (P i ) Features extracted from the feature map P i are used to predict the scaling factor of the height.

[0088] Step3: Classification head prediction; Through the classification branch, predict whether each candidate box contains a bird and the specific category of the bird.

[0089] The output of the classification head is expressed as: p = Softmax(W c2 • F(P i ) + b c2);

[0090] where: F(P i ) are the features extracted from the feature map P i ; W c2 and b c2 are the weight matrix and bias vector of the classification branch; p is a probability distribution representing the probability that the candidate box belongs to each category.

[0091] Step4: Output result; The final output includes: Bounding box position: (x, y, w, h); Class label: arg max(p). arg max(p) represents finding the index (or category) that maximizes the probability value in the given probability distribution p.

[0092] By introducing adaptive weights, the model can dynamically adjust the regression parameters according to objects of different scales and shapes, thereby significantly improving the positioning accuracy of the bounding box. Combining multi-scale feature maps and adaptive mechanisms, the model performs more robustly when dealing with complex backgrounds or small objects.

[0093] Processing flow of bird call data;

[0094] (1) Extract Mel Frequency Cepstral Coefficients (MFCC);

[0095] MFCC can capture the spectral characteristics of sound signals and is particularly suitable for the analysis of complex acoustic signals such as bird calls.

[0096] Steps: 1) Preprocessing: Frame the collected bird call data (usually each frame has a length of 20 - 40 ms), and apply a Hamming Window to reduce boundary effects.

[0097] 2) Fast Fourier Transform (FFT): Apply FFT to each frame of the signal to convert the time-domain signal into a frequency-domain signal.

[0098] 3) Mel filter bank: Use a set of triangular filters to weight the frequency-domain signal and map it to the Mel scale.

[0099] 4) Discrete Cosine Transform (DCT): Apply DCT to the logarithmic energy output by the Mel filter bank to obtain MFCC features.

[0100] The MFCC sequence is: X = [x1, x2, …, x T ; where: T is the time step, and x t is the MFCC feature vector of the t-th frame.

[0101] (2) Train an LSTM classifier;

[0102] Input the final hidden state of the LSTM into a fully connected layer to calculate the class probability:

[0103]

[0104] where: X is the MFCC feature sequence, Wc1 is the weight matrix of the classification layer, bc1 is the bias vector of the classification layer, and p(Class = c|X) represents the probability that the class is c under the condition of X.

[0105] (3) Judge the behavior through the time-domain pattern of the voiceprint;

[0106] The behavior patterns of bird calls (such as foraging, warning, etc.) can be further analyzed through the time features of the calls (such as intervals, durations). Extract the following time-domain features from the call data: Call interval: The time difference between adjacent calls; Duration: The duration of a single call; Call frequency: The number of calls per unit time. Use SVM to classify the time-domain features to judge the behavior pattern of the bird.

[0107] S4. Based on the position and pose information output by the improved YOLOv8 model, extract spatial features; based on the distance, speed, and trajectory information of the bird, extract dynamic features; based on the infrared features of the infrared sensor, extract thermal features; based on the sound features of the bird call, extract behavior pattern features; input the extracted features into the Transformer model to generate a joint feature vector, dynamically adjust the weights through environmental perception to calculate the comprehensive confidence, and judge whether to start bird repelling measures according to a preset threshold. Specifically:

[0108] 1. Feature extraction;

[0109] 1) Spatial feature extraction;

[0110] Input: The position (x, y, w, h) of the flying bird and the pose information (wing angle, head orientation, etc.) output by the improved YOLOv8 model.

[0111] Feature representation: Position encoding: Convert the bounding box coordinates into a vector in the normalized Cartesian coordinate system: ; where: W and H are the width and height of the image; f spatial represents the spatial feature vector of the flying bird target output from the improved YOLOv8 model.

[0112] Pose feature: Extract the wing angle θ and head direction Φ of the flying bird through key point detection, and construct a pose vector (f pose ): f pose = [sinθ, cosθ, sinΦ, cosΦ];

[0113] Integrated spatial feature (F spatial ): F spatial = MLP(f spatial ⊕f pose ); where ⊕ represents vector concatenation, and MLP is a multi-layer perceptron.

[0114] 2) Dynamic feature extraction (based on millimeter-wave radar data);

[0115] Input: The distance d, speed v, and trajectory point sequence of the flying bird . T represents the length of the sequence. (x t , y t ) is the specific position at time t.

[0116] Feature representation: Instantaneous dynamic feature (f dynamic ): f dynamic = [d, v, dx / dt, dy / dt]; dx / dt and dy / dt represent the speed components of the flying bird in the horizontal and vertical directions.

[0117] Trajectory time series feature: Apply a sliding window to the trajectory point sequence to extract acceleration (a t ) and trajectory curvature (κ t ): Trajectory curvature κ t represents the degree of bending of the flying bird's motion trajectory at time t, which is used to measure the degree of bending or turning amplitude of the trajectory at a certain point, and the trajectory curvature of the flying bird can help judge whether it is changing its flight direction, such as bypassing an obstacle or adjusting its flight path.

[0118] , ;

[0119] Construct the time - series feature vector f trajectory ∈ . represents the time - series feature matrix of the flying - bird movement trajectory extracted from millimeter - wave radar data; T is the time step, and 4 is the feature dimension of each time step. The feature dimension of each time step includes: horizontal position x t ; vertical position y t ; instantaneous velocity v t ; acceleration a t . Comprehensive dynamic feature (F dynamic ): F dynamic = LSTM(f dynamic ⊕f trajectory ).

[0120] 3) Thermal feature extraction (based on infrared sensor data);

[0121] Input: The heat - distribution matrix H of the infrared thermal image ∈ . M: represents the height direction (number of rows) of the infrared thermal image, i.e., the vertical resolution (number of pixels). N: represents the width direction (number of columns) of the infrared thermal image, i.e., the horizontal resolution (number of pixels). H is a two - dimensional matrix, and each element H i,j represents the infrared radiation intensity value of the pixel at the i - th row and j - th column in the image (usually positively correlated with temperature).

[0122] Feature representation includes: regional heat intensity: Perform ROI extraction on the flying - bird area and calculate the average heat intensity (μ heat ): ; K is the total number of pixels within the ROI.

[0123] Thermal distribution gradient: Calculate the heat - map gradient through the Sobel operator: G x = H×S x , G y = H×S y ; where S x and S y are Sobel operators used to calculate the gradients in the horizontal and vertical directions respectively; S y = S x T ; G x and G y are the heat - gradient matrices in the horizontal and vertical directions respectively (reflecting the intensity and direction of temperature change).

[0124] Construct the heat - gradient feature (f gradient ): f gradient = [vec(Gx ), vec(G y )]. Among them: vec(•): Flattens a matrix into a vector.

[0125] Comprehensive thermal feature: F thermal = Conv1D(μ heat ⊕ f gradient ); Conv1D: One-dimensional convolutional layer.

[0126] 4) Behavioral pattern feature extraction (based on microphone array data);

[0127] Input: MFCC feature sequence X ∈ and time-domain features (call interval Δt, duration τ). T is the time step, and D is the feature dimension.

[0128] Voiceprint feature: The voiceprint encoding is extracted through a pre-trained LSTM classifier: h audio = LSTM(X); Use the time-domain features to train an SVM classifier, and output the behavioral probability distribution p behavior ∈ R C (C is the number of behavioral categories).

[0129] Comprehensive behavioral feature: F behavior = h audio ⊕ p behavior .

[0130] 2. Multimodal feature fusion (Transformer model);

[0131] Input: Four types of feature vectors F spatial , F dynamic , F thermal , F behavior .

[0132] Feature alignment: Map the features to a unified feature dimension d through linear projection model : Z i = W i F i + b i (i ∈ {spatial, dynamic, thermal, behavior});

[0133] Among them: F i represents the original feature vector of the i-th modality, W i is the linear projection matrix, b i is the bias vector, and Z i is the aligned feature vector.

[0134] Positional encoding: Add temporal positional encoding P t(Due to the sequential information in radar and audio data): ; where: is the feature vector after adding sequential information.

[0135] Transformer encoder:

[0136] Multi-head self-attention:

[0137] After concatenating the multi-head outputs, we get H attn ∈ . Where: Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the query matrix, and Softmax is the normalized attention weight. QK T is the similarity calculation, which measures the matching degree between the query and the key through the dot product.

[0138] Feed-forward network: H joint = FFN(H attn ) = ReLU(W1H attn + b1)W2 + b2; where: H attn is the output vector of the self-attention mechanism, W1, b1 are the weight matrix and bias of the first layer, W2, b2 are the weight matrix and bias of the second layer, and H joint is the combined feature vector output by the FFN; H attn is the output vector of the self-attention mechanism.

[0139] 3. Environmental perception dynamic weight adjustment;

[0140] Environmental parameters: light intensity l, weather condition w (such as rain and fog level), time t (day and night indicator).

[0141] Weight generation includes:

[0142] Environmental encoding: mapping environmental parameters to a vector: e env = Embedding(l, w, t);

[0143] Dynamic weight calculation: ;

[0144] where i ∈ {spatial, dynamic, thermal, behavior}, ∑α i = 1. e env is the environmental encoding vector, W i is the i-th learnable weight matrix, α i is the dynamic weight of the i-th modality, and j is the temporary index during summation.

[0145] 4. Comprehensive confidence calculation and decision-making;

[0146] Comprehensive confidence level: ;

[0147] Where: C final is the comprehensive confidence level. If C final ≥γ, activate the bird repelling measures; where: γ is a preset threshold (such as 0.7). W c is a learnable weight matrix, and b c is the bias term.

[0148] S5. Predict the threat level based on the position and trajectory of the flying bird, and dynamically select the bird repelling action. Specifically, it includes:

[0149] 1. Input the real-time position (x, y) of the flying bird and the trajectory prediction points {(x t+1 , y t+1 ), …, x t+k , y t+k}; the coordinates of the airport target area (runway range: {(x runway , y runway )});

[0150] Threat level formula: ;

[0151] Where: ThreatLevel is the danger level value, C fina is the comprehensive confidence level, d min is the shortest Euclidean distance from the flying bird trajectory to the runway, indicating the degree of proximity of the flying bird to the target area; σ is the attenuation coefficient, controlling the sensitivity of the distance effect. The level division table is shown in Table 1.

[0152] ;

[0153] Table 1 Level division table

[0154]

[0155] 2. Dynamically select the bird repelling action;

[0156] Acoustic bird repelling (low intensity):

[0157] Applicable conditions: 0.3 ≤ ThreatLevel < 0.5;

[0158] Acoustic bird repelling parameters: acoustic wave frequency f = 5 - 10 kHz, acoustic wave duration T = 10 s;

[0159] Laser bird repelling (medium intensity):

[0160] Applicable conditions: 0.5 ≤ ThreatLevel < 0.7;

[0161] Parameters: Green laser (wavelength 532nm), scanning rate 2Hz;

[0162] 3) UAV interception (high intensity):

[0163] Applicable conditions: ThreatLevel≥0.7;

[0164] Parameters: UAV speed 10m / s, interception radius 20m;

[0165] Adaptive parameter adjustment: If the same bird triggers the bird repelling action multiple times but does not leave, enhance the bird repelling intensity: σ new =σ old •0.9 (narrow the sensitive area), where σ old is the old sensitivity value, and σ new is the new sensitivity value.

[0166] 1. Table 2 compares the performance of different sensor combinations in the airport bird repelling scenario. The experimental data shows that the four-modal fusion scheme based on dynamic vision sensors, millimeter-wave radars, infrared thermal imagings, and microphone arrays demonstrates significant performance improvement compared to traditional single / double sensor systems.

[0167] Table 2: Target detection accuracy in different environments (%)

[0168]

[0169] It can be concluded from Table 2 that the accuracy of multi-sensor fusion in complex environments (such as at night and in foggy days) is significantly higher than that of single sensors, with an average increase of 45%.

[0170] 2. Bird repelling response time and success rate;

[0171] Test scenario: Simulate 100 bird intrusion events.

[0172] Average response time: 85ms (traditional method > 300ms).

[0173] Bird repelling success rate: 94% (traditional method is 72%).

[0174] 3. Table 3 is a comparison of the performance indicators of the present invention and traditional methods;

[0175] Table 3 Comparison of overall performance indicators

[0176]

[0177] It can be concluded from Table 3 that the present invention has achieved breakthroughs in the three core indicators of detection accuracy, false alarm rate, and real-time performance. Its multi-modal fusion architecture and algorithm optimization provide solutions for airport bird repelling.

[0178] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An airport bird repelling detection method based on multi-sensor fusion, characterized in that: Including: S1. Deploy cameras equipped with dynamic vision sensors, millimeter-wave radars, infrared thermal imaging sensors, and microphone arrays in the target area; S2. Perform time synchronization and spatial registration on the data collected by the cameras equipped with dynamic vision sensors, millimeter-wave radars, infrared thermal imaging sensors, and microphone arrays respectively; S3. Input the processed visual data in step S2 into the improved YOLOv8 model to identify and detect bird targets, and output the position and attitude information of the birds; the millimeter-wave radar data includes the distance, speed, and movement trajectory of the birds; the infrared data outputs infrared features to assist in confirming the presence of the birds; extract the sound features of the bird calls through the microphone array to supplement the confirmation of the species and behavior patterns of the birds; S4. Based on the position and attitude information of the birds output by the improved YOLOv8 model, extract the spatial feature F spatial ; Based on the distance, speed and trajectory information of the birds, extract the dynamic feature F dynamic ; Based on the infrared features of the infrared thermal imaging sensor, extract the thermal feature F thermal ; Based on the sound features of the birds' calls, extract the behavior pattern feature F behavior ; Input the extracted features into the Transformer model to generate a joint feature vector, dynamically adjust the weights through environmental perception to calculate the comprehensive confidence, and judge whether to initiate bird repelling measures according to the preset threshold; S5. Predict the threat level based on the position and trajectory of the birds, and dynamically select bird repelling actions; In step S1, the camera equipped with a dynamic vision sensor captures the brightness changes caused by the movement of the birds, perceives the dynamic behavior of the birds in real time, and generates sparse event stream data. At the same time, the camera obtains the morphological features of the birds to assist in target recognition; The improved YOLOv8 model is: replace the standard convolutional layer in the Backbone with DCNv3; Replace the Feature Pyramid Network included in the Neck with a dynamic pyramid network; replace the bounding box regression in the Head with adaptive bounding box regression; Dynamically adjust the weight calculation to obtain the comprehensive confidence through environmental perception, specifically: Environmental parameters include light intensity l , weather conditions w , time t ; Weight generation includes: mapping environmental parameters to a vector: e env =Embedding( l, w, t ); Dynamic weight calculation: ; where: ∑α i = 1; e env is the environmental coding vector, W i is the i-th learnable weight matrix, α i is the dynamic weight of the i-th modality, j is the temporary index during summation, i ∈ {spatial, dynamic, thermal, behavior}; ; Wherein: C final is the comprehensive confidence level, W c is the learnable weight matrix, b c is the bias term, H joint is the joint feature vector.

2. The airport bird repelling detection method based on multi-sensor fusion according to claim 1, characterized in that: In step S2, the camera, millimeter-wave radar, infrared thermal imaging sensor, and microphone array generate initial timestamps when collecting data, calibrate the timestamps using the global clock source NTP protocol, measure and compensate for the time offset, and perform interpolation processing on the data based on the sampling frequency, so as to map to a unified time axis.

3. The airport bird repelling detection method based on multi-sensor fusion according to claim 2, wherein: Input the processed visual data in step S2 into the improved YOLOv8 model to identify and detect bird targets, and output the position and attitude information of the birds, including: 1) DCNv3 allows the convolutional kernel to adaptively adjust the sampling position according to the input content by dynamically learning the offset and modulation factor. The specific formula is: ; where: y(p) is the value of the output feature map at position p, x is the input feature map, K is the convolution kernel size, w k is the convolution kernel weight, Δ p k is the dynamically learned offset, Δ m k is the modulation factor; 2) The Dynamic Pyramid Network performs dynamic weight fusion on the multi-scale feature maps output by the Backbone to generate preliminary multi-scale feature maps; for each pair of feature maps C i and C j , calculate the dynamic weight α ij , and the formula is: α ij = Softmax(Conv light ( C i ⊕ C j )); Wherein: C i and C j are the input feature maps; C i ⊕ C j represents feature concatenation; Conv light is a lightweight convolution for generating weights; Softmax normalization ensures that the sum of the weights is 1; According to the generated dynamic weights α ij , the multi-scale feature maps are weighted and fused to generate a preliminary multi-scale feature map, and the weighted fusion formula is: ; In the formula: P i is the fused feature map, and Resize means to C j adjust it to be the same as P i the same resolution; 3) Use adaptive bounding box regression to accurately locate the position of the birds, and finally output the position and attitude information of the birds.

4. The airport bird repelling detection method based on multi-sensor fusion according to claim 3, wherein: The bird call data is used to supplement the confirmation of the species and behavior patterns of the birds, including: (1) Extract the Mel Frequency Cepstral Coefficients (MFCC) from the collected bird call data; (2) Train an LSTM classifier, with the input being the MFCC sequence, and output the class probability p through the classification layer; ; Where: X is the MFCC feature sequence, W c1 is the weight matrix of the classification layer, b c1 is the bias vector of the classification layer, and p(Class=c|X) represents the probability that the class is c under the condition of X.

5. The airport bird repelling detection method based on multi-sensor fusion according to claim 4, wherein: In step S5, The threat level formula is as follows: ; where: ThreatLevel is the numerical value of the danger level, d min is the shortest Euclidean distance from the bird flight path to the runway, σ is the attenuation coefficient.

6. The airport bird repelling detection method based on multi-sensor fusion according to claim 5, wherein: The bird repelling actions include low-intensity sound wave bird repelling, medium-intensity laser bird repelling, and high-intensity drone interception.

Citation Information

Patent Citations

  • Bird repelling method and system for power transmission line

    CN118585958A

  • Natural environment bird monitoring method based on multi-modal fusion deep learning and computer device

    CN119027775A

Cited By

  • Power distribution network directional bird repelling control method and system based on bird behavior recognition

    CN121488937A