Low-altitude small target detection method based on multivariate information fusion

By fusing acoustic, infrared and visible light information in low-altitude object detection technology, combined with deep learning models and intelligent decision-making fusion technology, the problem of insufficient small object detection accuracy in complex environments is solved, and efficient and accurate low-altitude small object detection is achieved.

CN119989253APending Publication Date: 2025-05-13NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411951137.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing low-altitude object detection technology lacks the accuracy of small object detection in complex environments and does not fully utilize the complementarity of multimodal information, resulting in insufficient robustness and real-timeness.

Method used

The low-altitude small-object detection method based on multivariate information fusion is adopted, and the detection robustness and accuracy are improved by fusing acoustic, infrared and visible information, combined with deep learning models and intelligent decision-making fusion technology. Specific steps include multi-stage audio signal separation and noise reduction processing, infrared and visible image preprocessing, multi-scale convolution improved convolution recurrent neural network for acoustic object detection, fused visible and infrared dual-branch network for visual object detection, and decision-making fusion through improved DS evidence theory.

Benefits of technology

It realizes efficient and accurate detection of low-altitude small targets in complex environments, significantly improves the robustness and accuracy of detection, and meets application needs such as low-altitude area safety monitoring and drone management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989253A_ABST
    Figure CN119989253A_ABST
Patent Text Reader

Abstract

The invention discloses a low-altitude small target detection method based on multivariate information fusion. The method comprises the following steps: carrying out separation and noise reduction processing on an audio signal through principal component analysis and fast independent component analysis, and carrying out efficient identification on acoustic features in combination with a convolutional recurrent neural network of multi-scale convolutional optimization so as to realize accurate detection of an acoustic target; differentiated preprocessing is carried out on image data, guided filtering is carried out on an infrared image to achieve high-fidelity noise reduction, image enhancement is completed on a visible light image through Gamma correction, and detection of a visual target is achieved in combination with an infrared and visible light double-branch improved YOLOv8 backbone network; the improved DS evidence theory based on real-time reliability and dynamic weight is utilized to carry out decision-making layer fusion on acoustic and visual detection results, and the detection precision and robustness of the method for low-altitude small targets in a complex environment are comprehensively improved. The method provides an innovative solution with theoretical and practical values for a small target detection and defense technology in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of small target detection, and in particular to a low-altitude small target detection method based on multi-information fusion. Background Art

[0002] In recent years, with the rapid development of technologies such as low-altitude drones and small aircraft, they have been widely used in logistics and transportation, public security monitoring, and military reconnaissance. However, low-altitude small targets usually have the characteristics of small size, strong concealment, and high flight speed, which makes their detection and positioning in complex environments face severe challenges. Existing detection methods mostly rely on a single type of sensor, such as infrared imaging, visible light cameras, or acoustic sensors. However, although infrared imaging can work in low-light or night conditions, it is often difficult to effectively distinguish targets when the background temperature difference is small; visible light cameras perform better during the day and in well-lit environments, but are easily disturbed in insufficient light or complex backgrounds; although acoustic sensors can capture the sound characteristics of the target, they show insufficient stability in environments with complex background noise or severe signal reflection.

[0003] Although some studies have attempted to improve the robustness of low-altitude target detection by multimodal information fusion, they often fail to fully exploit the complementarity of different modal information, resulting in limited fusion effects. In addition, traditional methods often lack real-time and dynamic adaptability, and are prone to reduced detection accuracy in complex and changing environments.

[0004] In response to the above problems, there is an urgent need for a method that can fully utilize the characteristics of multimodal information to achieve high-precision detection of low-altitude small targets in complex environments, so as to meet the application requirements of low-altitude area security monitoring and drone management. Summary of the invention

[0005] The purpose of the present invention is to provide a low-altitude small target detection method based on multi-information fusion to address the problems of insufficient detection accuracy of small targets in complex environments and insufficient utilization of multimodal information in existing low-altitude target detection technologies. By fusing acoustic, infrared and visible light information, combining deep learning models and intelligent decision-making fusion technology, the robustness and accuracy of detection are improved, and efficient and accurate detection of low-altitude small targets in complex environments can be achieved.

[0006] The technical solution to achieve the purpose of the present invention is: a low-altitude small target detection method based on multi-information fusion, the method comprising the following steps:

[0007] Step 1: Separate and reduce the noise of the audio signal in the detection area in a multi-stage manner, and extract key acoustic features;

[0008] Step 2, filtering and preprocessing the infrared image of the detection area, and performing brightness correction on the visible light image of the detection area;

[0009] Step 3: Use the convolutional recurrent neural network improved based on multi-scale convolution to identify the acoustic features of the audio signal and complete acoustic target detection;

[0010] Step 4: Improve the YOLOv8 network by building a network structure that integrates visible light and infrared branches, and then use the improved network to achieve visual target detection using multimodal information;

[0011] Step 5: Use the improved DS evidence theory based on real-time reliability and dynamic weight to make decision fusion between the acoustic target detection results and the visual target detection results to obtain the low-altitude small target detection results.

[0012] Furthermore, the multi-stage method in step 1 is to use principal component analysis and fast independent component analysis in sequence to separate and reduce noise on the audio signal.

[0013] Furthermore, step 1 specifically includes:

[0014] Step 1-1, establishing an acoustic model of the input signal;

[0015] For multi-channel audio signals, the audio signal of each channel is represented as a time series x i (t), where x i (t) represents the time series of the i-th channel at time t, i = 1, 2, ..., N, N is the total number of channels. The time series data of each channel is organized to obtain a matrix signal representation X:

[0016]

[0017] Where T is the total number of time points, x i (t j ) is the time series of the i-th channel at the j-th time point, j = 1, 2, ..., T;

[0018] Step 1-2, reduce the dimension of the input signal by principal component analysis, namely PCA method, which specifically includes:

[0019] (1) Perform central processing on the input signal data. The processing formula is:

[0020]

[0021] In the formula, For x i (t j ) The corresponding centralized processing result;

[0022] This gives the centralized matrix X′:

[0023]

[0024] (2) Calculate the covariance matrix ∑, whose dimension is N×N, which is used to describe the correlation between audio signals of different channels:

[0025]

[0026] (3) Perform eigenvalue decomposition on the covariance matrix ∑ to obtain the eigenvalue λ i and the eigenvector v l ;

[0027] (4) Select the first k eigenvectors v1, v2, ..., v with the largest eigenvalues k , construct the transformation matrix V k , and project the matrix X′ to the new principal component space X pca :

[0028] X pca =X′V k

[0029] Steps 1-3, perform FastICA separation on the reduced-dimensional signal, specifically including:

[0030] (1) Perform eigenvalue decomposition of the covariance matrix of the data after dimensionality reduction in step 1-2 to obtain the eigenvalue Λ and eigenvector W, and then further obtain the whitened signal

[0031]

[0032] (2) Continuously iterate the following formula to optimize w until convergence, and obtain the estimated matrix S of the independent components;

[0033]

[0034] Among them, w (n+1) 、w (n) are the weight vectors at the n+1th iteration and the nth iteration respectively; f(*) is a nonlinear function; λ is a regularization term;

[0035] (3) Select the first k independent components in the estimated matrix S, thereby separating the independent audio signal;

[0036] Steps 1-4, extract key acoustic features of the audio signal.

[0037] Furthermore, in steps 1-4, the key acoustic features of the audio signal are extracted by Mel-frequency cepstral coefficients MFCC.

[0038] Furthermore, in step 2, the infrared image is filtered and preprocessed, specifically: the infrared image is filtered and denoised by guided filtering.

[0039] Furthermore, in step 2, the improved Gamma correction is used to adjust the brightness of the visible light image. The improved Gamma correction formula is:

[0040]

[0041]

[0042] Where L is the normalized input pixel value; E is the corrected pixel value; R slope is the height ratio of the straight line and the curve at the point of tangency; x start is the horizontal coordinate of the tangent point; γ is the brightness adjustment index.

[0043] Furthermore, in step 3, the acoustic features of the audio signal are identified by using a convolutional recurrent neural network improved based on multi-scale convolution to complete acoustic target detection, which specifically includes:

[0044] Step 3-1, establish a CRNN model based on multi-scale convolution improvement: adopt a parallel time-frequency joint convolution structure, realize feature extraction in the time domain and frequency domain through a parallel architecture of four groups of three-layer convolution kernels, and use 1×n and n×1 decomposition convolution to replace the traditional n×n convolution operation; specifically:

[0045] The first set of parallel CNN layers consists of 1×1 convolutional layers followed by batch normalization and rectified linear unit activation functions;

[0046] The remaining three groups of parallel CNN layers are composed of three-layer CNN structures, including 1×1 convolutional layer, 1×n convolutional layer for extracting frequency domain features, and n×1 convolutional layer for extracting time domain features, and each convolutional layer is followed by BN and ReLU operations, where n=3, 5, and 7;

[0047] In the network output part, the results of the four parallel CNNs are concatenated into a one-dimensional vector, and the representative features are extracted through the maximum pooling layer;

[0048] Step 3-2: Input the key acoustic features extracted in step 1 into the improved CRNN model in step 3-2. The model outputs the detection probability P(target) for evaluating the existence of low-altitude small targets:

[0049] P(target)=softmax(W o ·h T +b o )

[0050]

[0051] Where W o is the weight matrix of the output layer, with a size of d h ×d c , d h is the dimension of the hidden layer, d c is the number of target categories; h T is the hidden state vector of the last time step T, generated by the recurrent layer GRU, representing the final compressed representation of the temporal features, with a dimension of d h ; b o The bias vector of the output layer, dimension d c , offset adjustment is performed on the scores of each category; z i It is only used to represent function variables and has no actual meaning; the detection probability P(target) is used as the acoustic target detection result.

[0052] Furthermore, in step 4, the YOLOv8 network is improved by constructing a network structure that integrates visible light and infrared dual branches, and then the improved network is used to realize visual target detection using multimodal information, specifically including:

[0053] Step 4-1: Use the visible light image and infrared image as inputs of the improved network to form a multimodal input pair (X vis ,X ir ), representing the feature tensors of visible light image and infrared image respectively;

[0054] Step 4-2, use the basic module of the YOLOv8 network to extract the preliminary features of the two modalities in step 4-1:

[0055] F vis =f vis (X vis )

[0056] F ir =f ir (X ir )

[0057] In the formula, f vis and f ir There are two feature extraction branches, F vis 、F ir X vis ,X ir Corresponding preliminary features;

[0058] Step 4-3, modality fusion through attention interaction mechanism, includes:

[0059] (1) Adjust the channel weight of each branch feature to generate a weighted feature representation:

[0060] F′ vis =σ(MLP(GAP(F vis )))·F vis

[0061] F′ ir =σ(MLP(GAP(F ir )))·F ir

[0062] Among them, GAP represents global average pooling, MLP represents multi-layer perceptron, σ is the Sigmoid function, which is used to generate the weight matrix; F′ vis , F′ ir F vis 、F ir The corresponding weighted features;

[0063] (2) Enhance features through cross-modal interaction to form fusion features F fuse :

[0064] F fuse =Concat(F′ vis ,F′ ir )+Attention(F′ vis ,F′ ir )

[0065] In the formula, Attention represents the attention mapping between modalities, and Concat represents the convolution between modalities;

[0066] (3) The fusion feature F fuse Input to the detection head of the YOLOv8 network:

[0067] D = f det (F fuse )

[0068] In the formula, f det It represents the detection module, which outputs the target bounding box and classification probability. D is the visual target detection result, including the target category class, location coordinates (x, y, w, h) and confidence score.

[0069] Furthermore, the improved DS evidence theory based on real-time reliability and dynamic weight described in step 5 specifically includes:

[0070] Step 5-1, real-time reliability calculation

[0071] Set a fixed time window l, collect the classification results of various types of target detection methods at different times, and construct the classification probability matrix P i,t :

[0072]

[0073] In the formula, Indicates that the classification result of the i-th type of target detection method at time t is the k-th category θ k The detection probability of

[0074] The introduction of Logistic model to detect reliability i,t To model:

[0075]

[0076] In the formula, N is the total number of classification categories;

[0077] Step 5-2, dynamic weight calculation

[0078] (1) Calculate the initial weight w of each type of target detection method i :

[0079]

[0080] Among them, a i is the historical classification accuracy of the i-th target detection method, M is the total number of types of target detection methods;

[0081] (2) Introducing discrete factor D i,t Describe the dispersion of classification scores at a certain moment:

[0082]

[0083] Discrete factor D i,t The larger it is, the stronger the sensor's classification ability is for the current target, and a higher weight should be given to it.

[0084] in, is the mean of the scores of different categories; σ i,t is the standard deviation of the scores for different categories;

[0085] (3) Combined with the initial weight w i and discrete factor D i,t , the dynamic weight calculation formula of each type of target detection method at time t is as follows:

[0086]

[0087] Where W i,t is the dynamic weight of the i-th target detection method;

[0088] (4) Introduce a weight penalty mechanism based on the difference in classification confidence between “detected targets” and “non-detected targets”:

[0089]

[0090] In the formula, Represents the category θ k The average classification probability of the i-th target detection method; Represents the j-th target detection method and all detection methods for category θ k Deviation from consistency;

[0091] Step 5-3, DS evidence fusion

[0092] (1) At time t, based on the DS evidence theory, combined with the classification results, reliability and dynamic weights of various types of target detection methods, the pair Ω = {θ1, ...θ k The mass function m θ,i :

[0093]

[0094] in, is the normalization factor, r i,t is the reliability factor of the i-th target detection method, W i,t is the dynamic weight of the i-th target detection method in the current environment, i = 1, 2, ..., k;

[0095] (2) The basic reliability of the sensors is integrated to calculate the final classification probability. The calculation formula is as follows:

[0096]

[0097] In the formula, Represents the category θ n The overall confidence level, Represents the confidence provided by different target detection methods; r i,t Reliability index representing different target detection methods; Indicates the likelihood of joint support between different hypotheses;

[0098] (3) According to the fused classification probability, select the category with the highest confidence as the final classification result c:

[0099]

[0100] Furthermore, in step 5, the acoustic target detection result and the visual target detection result of step 3 are respectively used as different types of target detection methods, and input into the improved DS evidence theory based on real-time reliability and dynamic weight to obtain the low-altitude small target detection result.

[0101] Compared with the prior art, the present invention has the following significant advantages:

[0102] (1) In acoustic target detection, the quality of acoustic feature extraction and the accuracy of acoustic target detection are greatly improved through multi-stage separation and noise reduction processing (PCA and FastICA methods) and key feature extraction (MFCC).

[0103] (2) The proposed multi-scale convolution improved convolutional recurrent neural network effectively combines the feature extraction capabilities of the time domain and frequency domain, significantly improving the ability to recognize acoustic targets, and is particularly suitable for the detection of low-altitude small targets.

[0104] (3) Through improved gamma correction and guided filtering preprocessing technology, the quality of infrared and visible light images is significantly enhanced, so that the visual detection module still has a high detection accuracy in an environment with large changes in lighting conditions.

[0105] (4) By introducing the attention interaction mechanism and cross-modal feature enhancement, the modal fusion effect of infrared and visible light images is improved, the complementarity between multimodal data is fully explored, and the robustness of visual target detection is greatly improved.

[0106] (5) Combining real-time reliability and environmental dynamic weight, and applying the improved DS evidence theory to fuse the decision information of multiple sensors, this method can more comprehensively characterize the impact of complex environment on low-altitude small target detection. The fusion algorithm effectively alleviates the fusion paradox problem in traditional algorithms and significantly improves the accuracy and reliability of the fusion results, meeting the high-precision requirements of low-altitude small target detection tasks.

[0107] The present invention is further described in detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0108] Figure 1 It is a structural schematic diagram of the low-altitude small target detection method based on multi-information fusion of the present invention.

[0109] Figure 2 It is a schematic diagram of the improved CRNN model in the method of the present invention.

[0110] Figure 3 It is a schematic diagram of the improved YOLOv8 in the method of the present invention. DETAILED DESCRIPTION

[0111] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0112] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0113] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0114] In one embodiment, in combination Figure 1 , provides a low-altitude small target detection method based on multi-information fusion, the method comprising:

[0115] Step 1: Separate and reduce the noise of the audio signal in a multi-stage manner to extract key acoustic features;

[0116] Here, the audio signal is an audio signal of the area to be detected collected by, but not limited to, a microphone array;

[0117] Step 2, filter preprocessing the infrared image and perform brightness correction on the visible light image;

[0118] Here, the infrared image is an infrared image of the area to be inspected collected by, but not limited to, an infrared camera;

[0119] Here, the visible light image is a visible light image of the area to be detected collected by, but not limited to, a visible light camera;

[0120] Step 3: Use the convolutional recurrent neural network improved based on multi-scale convolution to identify the acoustic features of the audio signal and complete acoustic target detection;

[0121] Step 4: Improve the YOLOv8 network by building a network structure that integrates visible light and infrared branches, and then use the improved network to achieve visual target detection using multimodal information;

[0122] Step 5: Use the improved DS evidence theory based on real-time reliability and dynamic weight to make decision fusion between the acoustic target detection results and the visual target detection results to obtain the low-altitude small target detection results.

[0123] Furthermore, in one embodiment, the multi-stage method in step 1 is to use principal component analysis (PCA) and fast independent component analysis (FastICA) in sequence to separate and reduce noise on the audio signal.

[0124] Here, in some embodiments, step 1 specifically includes:

[0125] Step 1-1, establishing an acoustic model of the input signal;

[0126] For multi-channel audio signals, the audio signal of each channel is represented as a time series x i (t), where x i (t) represents the time series of the i-th channel at time t, i = 1, 2, ..., N, N is the total number of channels, and the time series data of each channel are organized to obtain a matrix signal representation X:

[0127]

[0128] Where T is the total number of time points, x i (t j ) is the time series of the i-th channel at the j-th time point, j = 1, 2, ..., T; the matrix X is a T×N data matrix, which contains the audio signals of all time points and channels;

[0129] Step 1-2, reduce the dimension of the input signal by principal component analysis, namely PCA method, which specifically includes:

[0130] (1) Perform central processing on the input signal data. The processing formula is:

[0131]

[0132] In the formula, For x i (t j ) The corresponding centralized processing result;

[0133] This gives the centralized matrix X′:

[0134]

[0135] (2) Calculate the covariance matrix ∑, whose dimension is N×N, which is used to describe the correlation between audio signals of different channels:

[0136]

[0137] (3) Perform eigenvalue decomposition on the covariance matrix ∑ to obtain the eigenvalue λ i and the eigenvector vl ;

[0138] (4) Select the first k eigenvectors v1, v2, ..., v with the largest eigenvalues k , construct the transformation matrix V k , and project the matrix X′ to the new principal component space X pca :X pca =X′V k

[0139] This step effectively reduces the dimensionality of the multi-channel audio signal to k principal components, retains most of the data change information, and effectively reduces noise and redundancy;

[0140] Steps 1-3, perform FastICA separation on the reduced-dimensional signal, specifically including:

[0141] (1) Perform eigenvalue decomposition of the covariance matrix of the data after dimensionality reduction in step 1-2 to obtain the eigenvalue Λ and eigenvector W, and then further obtain the whitened signal

[0142]

[0143] The core idea of ​​FastICA is to separate independent components by maximizing the non-Gaussianity of the signal. In order to make the separated components statistically independent, negative entropy is used as the optimization objective. Specifically, FastICA maximizes independence through the following iterative optimization algorithm;

[0144] (2) Continuously iterate the following formula to optimize w until convergence, and obtain the estimated matrix S of the independent components;

[0145]

[0146] Among them, w (n+1) 、w (n) are the weight vectors at the n+1th iteration and the nth iteration respectively; f(*) is a nonlinear function; λ is a regularization term;

[0147] (3) Select the first k independent components in the estimated matrix S, thereby separating the independent audio signal;

[0148] Steps 1-4, extract key acoustic features of the audio signal.

[0149] Preferably, in steps 1-4, key acoustic features of the audio signal are extracted by Mel-frequency cepstral coefficients MFCC.

[0150] For audio signals, the key features are extracted through Mel Frequency Cepstral Coefficients (MFCC). First, the input signal x(t) is enhanced in high frequency and background noise is filtered out:

[0151] x′(t)=x(t)-αx(t-1)

[0152] Where α is the pre-emphasis coefficient, which is usually 0.97. Then, the signal is divided into multiple frames and a Hamming window w(n) is added to reduce the boundary effect:

[0153] x w (n) = x′(n)·w(n)

[0154]

[0155] Then, the windowed signal is subjected to a fast Fourier transform (FFT) to obtain the spectrum:

[0156]

[0157] Next, the spectral energy is mapped to the logarithmic Mel scale through the Mel filter bank:

[0158]

[0159] Among them, H m (k) is the weight of the Mel filter. Finally, the Mel energy is subjected to discrete cosine transform (DCT) to obtain MFCC:

[0160]

[0161] The extracted MFCC features C = {c1, c2, ..., c N} is the main characteristic representation of the audio signal.

[0162] Preferably, in some embodiments, in step 2, filtering preprocessing is performed on the infrared image, specifically: filtering and denoising the infrared image by guided filtering.

[0163] Taking the guidance image I as a reference, assume that the output image q is a local linear transformation of I:

[0164]

[0165] Among them, a k and b k is the coefficient to be solved; w k is a window centered at pixel k. To preserve edge characteristics, the optimization objective is constructed by minimizing the difference between the input images p and q:

[0166]

[0167] Among them, ε is the regularization parameter. After minimizing the optimization objective, the solution is obtained:

[0168]

[0169] Among them, μ k and They are window w k The mean and variance of I in ; is the mean of the input image p. When , the filtering realizes mean smoothing in the smooth area; when When it is large, the filter output approximately maintains the input characteristics, thereby retaining edge information.

[0170] Preferably, in some embodiments, in step 2, an improved Gamma correction is used to adjust the brightness of the visible light image, and the improved Gamma correction formula is:

[0171]

[0172] Where L is the normalized input pixel value; E is the corrected pixel value; R slope is the height ratio of the straight line and the curve at the point of tangency; x start is the horizontal coordinate of the tangent point; γ is the brightness adjustment index.

[0173] In order to verify the effect of the improved Gamma correction in brightness adjustment, the brightness distribution and visual quality of the visible light image before and after the improvement are compared. A visible light image with complex lighting conditions is selected, including too dark areas and highlight areas.

[0174] Contrast indicators: average pixel brightness (APL), contrast enhancement ratio (CEB), brightness uniformity (LU).

[0175] method APL CEB LU Image observation effect Traditional Gamma Correction 0.45 1.2 0.68 Loss of detail in dark areas and overflow in highlights Improved Gamma Correction 0.62 1.8 0.85 Dark details are significantly enhanced, and there is no overflow in the highlights

[0176] The improved Gamma correction algorithm dynamically adjusts x start and R slope , adapting to the image processing needs under different brightness conditions, can effectively improve the overall brightness and contrast of the image, while retaining more dark and highlight detail information. Especially under complex lighting conditions, the improved method shows better visual effects and brightness uniformity, meeting the detection needs in low-light environments.

[0177] Furthermore, in one embodiment, the step 3 uses a convolutional recurrent neural network based on multi-scale convolution to identify the acoustic features of the audio signal and complete acoustic target detection, which specifically includes:

[0178] Step 3-1, establish a CRNN model based on multi-scale convolution improvement, such as Figure 2As shown in the figure, a parallel time-frequency joint convolution structure is adopted to realize feature extraction in the time domain and frequency domain through a parallel architecture of four groups of three-layer convolution kernels. 1×n and n×1 decomposition convolutions are used to replace the traditional n×n convolution operation, which significantly reduces the computational cost while ensuring the comprehensiveness of feature extraction. Specifically:

[0179] The first set of parallel CNN layers consists of 1×1 convolutional layers, followed by batch normalization (BN) and rectified linear unit (ReLU) activation functions;

[0180] The remaining three groups of parallel CNN layers are composed of three-layer CNN structures, including 1×1 convolutional layer, 1×n convolutional layer for extracting frequency domain features, and n×1 convolutional layer for extracting time domain features, and each convolutional layer is followed by BN and ReLU operations, where n=3, 5, and 7; the filter size of the final convolutional layer is set to 32;

[0181] In the network output part, the results of the four parallel CNNs are concatenated into a one-dimensional vector, and the representative features are extracted through a maximum pooling layer (size 128);

[0182] Step 3-2: Input the key acoustic features extracted in step 1 into the improved CRNN model in step 3-2. The model outputs the detection probability P(target) for evaluating the existence of low-altitude small targets:

[0183] P(target)=softmax(W o ·h T +b o )

[0184]

[0185] Where W o is the weight matrix of the output layer, with a size of d h ×d c , d h is the dimension of the hidden layer, d c is the number of target categories; h T is the hidden state vector of the last time step T, generated by the recurrent layer GRU, representing the final compressed representation of the temporal features, with a dimension of d h ; b o The bias vector of the output layer, dimension d c , offset adjustment is performed on the scores of each category; z i It is only used to represent function variables and has no actual meaning; the detection probability P(target) is used as the acoustic target detection result.

[0186] Furthermore, in one embodiment, in step 4, the YOLOv8 network is improved by constructing a network structure that integrates visible light and infrared dual branches, and then the improved network is used to realize visual target detection using multimodal information, combined with Figure 3 , including:

[0187] Step 4-1: Use the visible light image and infrared image as inputs of the improved network to form a multimodal input pair (X vis ,X ir ), representing the feature tensors of visible light image and infrared image respectively;

[0188] Step 4-2, use the basic module of the YOLOv8 network to extract the preliminary features of the two modalities in step 4-1:

[0189] F vis =f vis (X vis )

[0190] F ir =f ir (X ir )

[0191] In the formula, f vis and f ir There are two feature extraction branches, F vis 、F ir X vis ,X ir Corresponding preliminary features;

[0192] Step 4-3, modality fusion through attention interaction mechanism, includes:

[0193] (1) Adjust the channel weight of each branch feature to generate a weighted feature representation:

[0194] F′ vis =σ(MLP(GAP(F vis )))·F vis

[0195] F′ ir =σ(MLP(GAP(F ir )))·F ir

[0196] Among them, GAP represents global average pooling, MLP represents multi-layer perceptron, σ is the Sigmoid function, which is used to generate the weight matrix; F′ vis , F′ ir F vis 、F ir The corresponding weighted features;

[0197] (2) Enhance features through cross-modal interaction to form fusion features F fuse :

[0198] F fuse =Concat(F′ vis ,F′ ir )+Attention(F′ vis ,F′ ir )

[0199] In the formula, Attention represents the attention mapping between modalities, which is used to highlight complementary information; Concat represents convolution between modalities;

[0200] (3) The fusion feature F fuse Input to the detection head of the YOLOv8 network:

[0201] D = f det (F fuse )

[0202] In the formula, f det It represents the detection module, which outputs the target bounding box and classification probability. D is the visual target detection result, including the target category class, location coordinates (x, y, w, h) and confidence score.

[0203] Furthermore, in one embodiment, the improved DS evidence theory based on real-time reliability and dynamic weight in step 5 specifically includes:

[0204] Step 5-1, real-time reliability calculation

[0205] Set a fixed time window l, collect the classification results of various types of target detection methods at different times, and construct the classification probability matrix P i,t :

[0206]

[0207] In the formula, Indicates that the classification result of the i-th type of target detection method at time t is the k-th category θ k The detection probability of

[0208] Here, if a binary judgment method is used, the categories only include two categories, such as the presence of a small target and the absence of a small target (drone);

[0209] Since the same target is observed, the classification results within the time window should be consistent in theory, that is, However, in actual applications, classification results may fluctuate due to external factors.

[0210] In order to quantitatively evaluate the impact of this fluctuation on reliability, the Logistic model is introduced to detect reliability r i,t To model:

[0211]

[0212] In the formula, N is the total number of classification categories;

[0213] When the classification results fluctuate less, the reliability r i,t When the fluctuation is large, the reliability is close to 0. This method effectively excludes data that is greatly disturbed by the environment and reduces the adverse effects of unstable detection methods on the fusion results.

[0214] Step 5-2, dynamic weight calculation

[0215] The weights of decision information of different types of target detection methods reflect their relative importance to the task. In target classification, the weight is closely related to the classification ability of the target detection method. In historical data, the target detection method with good classification results indicates that it has strong classification ability and the assigned weight should be larger.

[0216] (1) Calculate the initial weight w of each type of target detection method i :

[0217]

[0218] Among them, a i is the historical classification accuracy of the i-th target detection method, M is the total number of types of target detection methods;

[0219] (2) By analyzing the historical classification data, it can be found that when the target detection method works in a suitable environment, its score for the correct category is significantly higher than that for the other category, which is manifested as a large gap between the scores of the two categories. Based on this characteristic, the discrete factor D is introduced. i,t Describe the dispersion of classification scores at a certain moment:

[0220]

[0221] Discrete factor D i,t The larger it is, the stronger the sensor's classification ability is for the current target, and a higher weight should be given to it.

[0222] in, is the mean of the scores of different categories; σ i,t is the standard deviation of the scores for different categories;

[0223] (3) Combined with the initial weight w i and discrete factor D i,t, the dynamic weight calculation formula of each type of target detection method at time t is as follows:

[0224]

[0225] Where W i,t is the dynamic weight of the i-th target detection method;

[0226] Here, by combining historical classification results with the classification capabilities in real-time environments, the dynamic weights can accurately reflect the classification capabilities of sensors in different environments;

[0227] (4) However, if target detection is a binary classification problem, the reliability of the target detection method may continue to increase when it continuously detects that the target does not exist. This may result in a failure to correctly detect the drone when it actually exists due to the failure of a target detection method, thus affecting the final result of the fusion detection.

[0228] In order to solve this problem, a weight penalty mechanism based on the difference in classification confidence between "detected target" and "non-detected target" is introduced:

[0229]

[0230] In the formula, Represents the category θ k The average classification probability of the i-th target detection method; Represents the j-th target detection method and all detection methods for category θ k Deviation from consistency;

[0231] Here, if it is a binary classification, the mechanism is expressed as:

[0232]

[0233] When the classification probability of a "detected target" is significantly higher than that of a "non-detected target", the penalty mechanism will reduce the dynamic weight of the sensor, thereby reducing its impact on the fusion result. This method effectively evaluates the current classification performance of the sensor through discrete factors and adjusts the weights in real time, significantly enhancing the robustness and adaptability of the fusion algorithm in complex target detection tasks.

[0234] Step 5-3, DS evidence fusion

[0235] (1) At time t, based on the DS evidence theory, combined with the classification results, reliability and dynamic weights of various types of target detection methods, the pair Ω = {θ1,… k The mass function m θ,i :

[0236]

[0237] in, is the normalization factor, r i,t is the reliability factor of the i-th target detection method, W i,t is the dynamic weight of the i-th target detection method in the current environment, i = 1, 2, ..., k, P(θ) is the detection probability;

[0238] Indicates invalid or contradictory information and does not participate in the hypothesis analysis.

[0239] Represents a set of hypotheses, which is a specific object of possibility analysis.

[0240] θ = P(θ): represents the normalized probability of the hypothesis, which is used to normalize the confidence of the final decision layer.

[0241] The three states are the products of different stages in processing evidential reasoning and play a role in the entire evidence theory framework.

[0242] (2) The basic reliability of the sensors is integrated to calculate the final classification probability. The calculation formula is as follows:

[0243]

[0244] In the formula, Represents the category θ n The overall confidence level, Represents the confidence provided by different target detection methods; r i,t Reliability index representing different target detection methods; Indicates the likelihood of joint support between different hypotheses;

[0245] (3) According to the fused classification probability, select the category with the highest confidence as the final classification result c:

[0246]

[0247] Then, in step 5, the acoustic target detection result and the visual target detection result of step 3 are respectively used as different types of target detection methods, and input into the improved DS evidence theory based on real-time reliability and dynamic weight to obtain the low-altitude small target detection result.

[0248] Here, by combining real-time reliability and environmental dynamic weights, and applying the improved DS evidence theory to fuse the decision information of multiple sensors, this method can more comprehensively characterize the impact of complex environments on low-altitude small target detection. The fusion algorithm effectively alleviates the fusion paradox problem in traditional algorithms, and significantly improves the accuracy and reliability of the fusion results, meeting the high-precision requirements of low-altitude small target detection tasks. The method comprehensively improves the detection accuracy and robustness of low-altitude small targets in complex environments. This method provides an innovative solution with theoretical and practical value for small target detection and defense technology in complex scenarios.

[0249] In one embodiment, a low-altitude small target detection system based on multivariate information fusion is provided, the system comprising:

[0250] The first module is used to separate and reduce the noise of the audio signal of the detection area in a multi-stage manner and extract key acoustic features;

[0251] The second module is used to perform filtering preprocessing on the infrared image of the detection area and perform brightness correction on the visible light image of the detection area;

[0252] The third module is used to identify the acoustic features of audio signals using a convolutional recurrent neural network improved based on multi-scale convolution to complete acoustic target detection;

[0253] The fourth module is used to improve the YOLOv8 network by building a network structure that integrates visible light and infrared dual branches, and then use the improved network to achieve visual target detection using multimodal information;

[0254] The fifth module is used to adopt the improved DS evidence theory based on real-time reliability and dynamic weight to make decision fusion on the acoustic target detection results and the visual target detection results to obtain the low-altitude small target detection results.

[0255] For the specific definition of the low-altitude small target detection system based on multivariate information fusion, please refer to the definition of the low-altitude small target detection method based on multivariate information fusion mentioned above, which will not be repeated here. Each module in the above-mentioned low-altitude small target detection system based on multivariate information fusion can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0256] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following is achieved:

[0257] Step 1: Separate and reduce the noise of the audio signal in the detection area in a multi-stage manner, and extract key acoustic features;

[0258] Step 2, filtering and preprocessing the infrared image of the detection area, and performing brightness correction on the visible light image of the detection area;

[0259] Step 3: Use the convolutional recurrent neural network improved based on multi-scale convolution to identify the acoustic features of the audio signal and complete acoustic target detection;

[0260] Step 4: Improve the YOLOv8 network by building a network structure that integrates visible light and infrared branches, and then use the improved network to achieve visual target detection using multimodal information;

[0261] Step 5: Use the improved DS evidence theory based on real-time reliability and dynamic weight to make decision fusion between the acoustic target detection results and the visual target detection results to obtain the low-altitude small target detection results.

[0262] For the specific limitations of each step, please refer to the limitations of the low-altitude small target detection method based on multi-information fusion mentioned above, which will not be repeated here.

[0263] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the computer program implements:

[0264] Step 1: Separate and reduce the noise of the audio signal in the detection area in a multi-stage manner, and extract key acoustic features;

[0265] Step 2, filtering and preprocessing the infrared image of the detection area, and performing brightness correction on the visible light image of the detection area;

[0266] Step 3: Use the convolutional recurrent neural network improved based on multi-scale convolution to identify the acoustic features of the audio signal and complete acoustic target detection;

[0267] Step 4: Improve the YOLOv8 network by building a network structure that integrates visible light and infrared branches, and then use the improved network to achieve visual target detection using multimodal information;

[0268] Step 5: Use the improved DS evidence theory based on real-time reliability and dynamic weight to make decision fusion between the acoustic target detection results and the visual target detection results to obtain the low-altitude small target detection results.

[0269] For the specific limitations of each step, please refer to the limitations of the low-altitude small target detection method based on multi-information fusion mentioned above, which will not be repeated here.

[0270] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A low-altitude small target detection method based on multivariate information fusion, characterized in that: The method comprises the following steps: Step 1: Separate and reduce the noise of the audio signal in the detection area in a multi-stage manner, and extract key acoustic features; Step 2, filtering and preprocessing the infrared image of the detection area, and performing brightness correction on the visible light image of the detection area; Step 3: Use the convolutional recurrent neural network improved based on multi-scale convolution to identify the acoustic features of the audio signal and complete acoustic target detection; Step 4: Improve the YOLOv8 network by building a network structure that integrates visible light and infrared branches, and then use the improved network to achieve visual target detection using multimodal information; Step 5: Use the improved DS evidence theory based on real-time reliability and dynamic weight to make decision fusion between the acoustic target detection results and the visual target detection results to obtain the low-altitude small target detection results.

2. The low-altitude small target detection method based on multivariate information fusion according to claim 1 is characterized in that: The multi-stage method described in step 1 is to use principal component analysis and fast independent component analysis in sequence to separate and reduce noise on the audio signal.

3. The low-altitude small target detection method based on multivariate information fusion according to claim 2 is characterized in that: Step 1 specifically includes: Step 1-1, establishing an acoustic model of the input signal; For multi-channel audio signals, the audio signal of each channel is represented as a time series x i (t), where x i (t) represents the time series of the i-th channel at time t, i = 1, 2, ..., N, N is the total number of channels, and the time series data of each channel are organized to obtain a matrix signal representation X: Where T is the total number of time points, x i (t j ) is the time series of the i-th channel at the j-th time point, j = 1, 2, ..., T; Step 1-2, reduce the dimension of the input signal by principal component analysis, namely PCA method, which specifically includes: (1) Perform central processing on the input signal data. The processing formula is: In the formula, For x i (t j ) The corresponding centralized processing result; This gives the centralized matrix X′: (2) Calculate the covariance matrix v, whose dimension is N×N, which is used to describe the correlation between audio signals of different channels: (3) Perform eigenvalue decomposition on the covariance matrix ∑ to obtain the eigenvalue λ i and the eigenvector v l ; (4) Select the first k eigenvectors v1, v2, ..., v with the largest eigenvalues k , construct the transformation matrix V k , and project the matrix X′ to the new principal component space X pca : X pca =X'V k Steps 1-3, perform FastICA separation on the reduced-dimensional signal, specifically including: (1) Perform eigenvalue decomposition of the covariance matrix of the data after dimensionality reduction in step 1-2 to obtain the eigenvalue Λ and eigenvector W, and then further obtain the whitened signal (2) Continuously iterate the following formula to optimize w until convergence, and obtain the estimated matrix S of the independent components; Among them, w (n+1) 、w (n) are the weight vectors at the n+1th iteration and the nth iteration respectively; f(*) is a nonlinear function; λ is a regularization term; (3) Select the first k independent components in the estimated matrix S, thereby separating the independent audio signal; Steps 1-4, extract key acoustic features of the audio signal.

4. The low-altitude small target detection method based on multivariate information fusion according to claim 1 is characterized in that: In steps 1-4, the key acoustic features of the audio signal are extracted through the Mel-frequency cepstral coefficients MFCC.

5. The low-altitude small target detection method based on multivariate information fusion according to claim 1 is characterized in that: In step 2, the infrared image is filtered and preprocessed, specifically: the infrared image is filtered and denoised by guided filtering.

6. The low-altitude small target detection method based on multivariate information fusion according to claim 1 is characterized in that: In step 2, the improved Gamma correction is used to adjust the brightness of the visible light image. The improved Gamma correction formula is: Where L is the normalized input pixel value; E is the corrected pixel value; R slope is the height ratio of the straight line and the curve at the point of tangency; x start is the horizontal coordinate of the tangent point; γ is the brightness adjustment index.

7. The low-altitude small target detection method based on multivariate information fusion according to claim 1 is characterized in that: Step 3 uses a convolutional recurrent neural network based on multi-scale convolution to identify the acoustic features of the audio signal and complete acoustic target detection, which specifically includes: Step 3-1, establish a CRNN model based on multi-scale convolution improvement: adopt a parallel time-frequency joint convolution structure, realize feature extraction in the time domain and frequency domain through a parallel architecture of four groups of three-layer convolution kernels, and use 1×n and n×1 decomposition convolution to replace the traditional n×n convolution operation; specifically: The first set of parallel CNN layers consists of 1×1 convolutional layers followed by batch normalization and rectified linear unit activation functions; The remaining three groups of parallel CNN layers are composed of three-layer CNN structures, including 1×1 convolutional layer, 1×n convolutional layer for extracting frequency domain features, and n×1 convolutional layer for extracting time domain features, and each convolutional layer is followed by BN and ReLU operations, where n=3, 5, and 7; In the network output part, the results of the four parallel CNNs are concatenated into a one-dimensional vector, and the representative features are extracted through the maximum pooling layer; Step 3-2: Input the key acoustic features extracted in step 1 into the improved CRNN model in step 3-2. The model outputs the detection probability P(target) for evaluating the existence of low-altitude small targets: P(target)=softmax(W o ·h T +b o ) Where W o is the weight matrix of the output layer, with a size of d h ×d c , d h is the dimension of the hidden layer, d c is the number of target categories; h T is the hidden state vector of the last time step T, generated by the recurrent layer GRU, representing the final compressed representation of the temporal features, with a dimension of d h ; b o The bias vector of the output layer, dimension d c , offset adjustment is performed on the scores of each category; z i It is only used to represent function variables and has no actual meaning; the detection probability P(target) is used as the acoustic target detection result.

8. The low-altitude small target detection method based on multivariate information fusion according to claim 1 is characterized in that: In step 4, the YOLOv8 network is improved by building a network structure that integrates visible light and infrared dual branches, and then the improved network is used to achieve visual target detection using multimodal information, specifically including: Step 4-1: Use the visible light image and infrared image as inputs of the improved network to form a multimodal input pair (X vis ,X ir ), representing the feature tensors of visible light image and infrared image respectively; Step 4-2, use the basic module of the YOLOv8 network to extract the preliminary features of the two modalities in step 4-1: F vis =f vis (X vis ) F ir =f ir (X ir ) In the formula, f vis and f ir There are two feature extraction branches, F vis 、F ir X vis ,X ir Corresponding preliminary features; Step 4-3, modality fusion through attention interaction mechanism, includes: (1) Adjust the channel weight of each branch feature to generate a weighted feature representation: F′ vis σ(MLP(GAP(F vis )))·F vis F′ ir σ(MLP(GAP(F ir )))·F ir Among them, GAP represents global average pooling, MLP represents multi-layer perceptron, σ is the Sigmoid function, which is used to generate the weight matrix; F v ' is 、F i ' r F vis 、F ir The corresponding weighted features; (2) Enhance features through cross-modal interaction to form fusion features F fuse : F fuse =Concat(F v ′ is ,F i ′ r )+Attention(F v ′ is ,F i ′ r ) In the formula, Attention represents the attention mapping between modalities, and Concat represents the convolution between modalities; (3) The fusion feature F fuse Input to the detection head of the YOLOv8 network: D=f det (F fuse ) In the formula, f det It represents the detection module, which outputs the target bounding box and classification probability. D is the visual target detection result, including the target category class, location coordinates (x, y, w, h) and confidence score.

9. The low-altitude small target detection method based on multivariate information fusion according to claim 1 is characterized in that: The improved DS evidence theory based on real-time reliability and dynamic weight described in step 5 specifically includes: Step 5-1, real-time reliability calculation Set a fixed time window l, collect the classification results of various types of target detection methods at different times, and construct the classification probability matrix P i,t : In the formula, Indicates that the classification result of the i-th type of target detection method at time t is the k-th category θ k The detection probability of The introduction of Logistic model to detect reliability i,t To model: In the formula, N is the total number of classification categories; Step 5-2, dynamic weight calculation (1) Calculate the initial weight w of each type of target detection method i : Among them, a i is the historical classification accuracy of the i-th target detection method, M is the total number of types of target detection methods; (2) Introducing discrete factor D i,t Describe the dispersion of classification scores at a certain moment: Discrete factor D i,t The larger it is, the stronger the sensor's classification ability is for the current target, and a higher weight should be given to it. in, is the mean of the scores of different categories; σ i,t is the standard deviation of the scores for different categories; (3) Combined with the initial weight w i and discrete factor D i,t , the dynamic weight calculation formula of each type of target detection method at time t is as follows: Where W i,t is the dynamic weight of the i-th target detection method; (4) Introduce a weight penalty mechanism based on the difference in classification confidence between "detected target" and "non-detected target": In the formula, Represents the category θ k The average classification probability of the i-th target detection method; Represents the j-th target detection method and all detection methods for category θ k Deviation from consistency; Step 5-3, DS evidence fusion (1) At time t, based on the DS evidence theory, combined with the classification results, reliability and dynamic weights of various types of target detection methods, the pair Ω = {θ1,…θ k The mass function m θ,i : in, is the normalization factor, r i,t is the reliability factor of the i-th target detection method, W i,t is the dynamic weight of the i-th target detection method in the current environment, i = 1, 2, ..., k, P(θ) is the detection probability; (2) The basic reliability of the sensors is integrated to calculate the final classification probability. The calculation formula is as follows: In the formula, Represents the category θ n The overall confidence level, Represents the confidence provided by different target detection methods; r i,t Reliability index representing different target detection methods; Indicates the likelihood of joint support between different hypotheses; (3) According to the fused classification probability, select the category with the highest confidence as the final classification result c:

10. The low-altitude small target detection method based on multi-information fusion according to claim 9 is characterized in that: In step 5, the acoustic target detection result and the visual target detection result of step 3 are respectively used as different types of target detection methods, and input into the improved DS evidence theory based on real-time reliability and dynamic weight to obtain the low-altitude small target detection result.