A method and system for detecting shallowly buried explosives based on multi-modal fusion
By using a multi-modal fusion method of magnetic and acoustic signals, the limitations of single-mode detection in shallow-buried explosive detection have been overcome, enabling efficient and stable detection in complex environments and improving the equipment's identification accuracy and anti-interference capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-15
AI Technical Summary
Existing shallow-buried explosive detection technologies suffer from several drawbacks. Single-mode detection methods are susceptible to background magnetic interference, soil attenuation, and scattering. Multi-mode fusion methods struggle to achieve effective registration and deep feature correlation, resulting in complex equipment, high power consumption, and insufficient identification stability.
A multimodal fusion method combining magnetic and acoustic signals is adopted. Through synchronous acquisition, preprocessing, time alignment and data slicing, feature extraction and standardization, a multimodal feature fusion model is constructed. The intermodal correlation weights are calculated using a multi-head attention mechanism, and target discrimination is performed through a lightweight classification structure.
It improves the accuracy of shallow-buried explosive detection in complex environments, reduces equipment complexity and power consumption, and enhances the system's anti-interference capability and identification stability.
Smart Images

Figure CN121878869B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of explosive detection technology, specifically relating to a method and system for detecting shallow-buried explosives based on multimodal fusion. Background Technology
[0002] Existing shallow-buried explosive detection technologies have evolved from single-physical-mode approaches to detection systems that integrate multiple sensor information. Early methods, such as electromagnetic induction and optical / infrared imaging, relied on electromagnetic or light wave propagation and had high detection efficiency for explosives near metal or the surface. However, their signals were severely attenuated in high-loss or opaque soils, making it difficult to reliably detect targets buried at greater depths or with low metal content. To improve underground penetration capabilities, existing technologies have introduced ground-penetrating radar (GPR) and biochemical detection techniques: GPR images by detecting discontinuities in the underground dielectric constant, while biochemical detection identifies explosive residues through the diffusion of explosive vapors or microbial reactions. These methods extend the detection range to non-metallic explosives, but they are highly sensitive to soil moisture, temperature, and diffusion stability, resulting in poor output stability. Electrical impedance tomography (EIT) and magnetometer detection technologies attempt to identify explosives by utilizing differences in underground conductivity and magnetic anomalies generated by ferromagnetic components in the explosive, respectively. However, EIT is susceptible to interference from soil heterogeneity, magnetometer detection is susceptible to environmental magnetic noise, and the detection effect is limited by the distance between the sensor and the target. Acoustic detection technology uses the acoustic impedance difference between shallowly buried explosives and the surrounding soil to identify structures. Under certain conditions, it has good environmental adaptability, but it is still limited by strong scattering and attenuation in complex soil media.
[0003] To mitigate the limitations of single-modal detection methods, existing technologies have proposed multimodal fusion detection schemes, such as electromagnetic induction-ground penetrating radar fusion and acoustic-ground penetrating radar fusion. These schemes reduce false alarms and improve robustness by fusing different physical cues. However, most existing multimodal systems suffer from problems such as complex equipment structures, high system power consumption, and insufficient recognition stability in complex environments. Summary of the Invention
[0004] To address the problems of existing single magnetic anomaly detection methods being susceptible to background magnetic interference, single acoustic detection methods being susceptible to soil attenuation and scattering, and existing multimodal fusion methods struggling to achieve effective registration and deep feature correlation, this invention provides a method and system for detecting shallowly buried explosives based on multimodal fusion. The technical solution is as follows:
[0005] A method for detecting shallowly buried explosives based on multimodal fusion includes the following steps:
[0006] S1. Synchronous acquisition of magnetic and acoustic signals;
[0007] S2. Preprocess the magnetic and acoustic signals;
[0008] S3. Time alignment and data slicing of magnetic and acoustic signals; the magnetic and acoustic signals are segmented according to a unified time window, the length of each time window is a pre-set fixed time period, and the magnetic signal segments and acoustic signal segments within the same time window are matched accordingly.
[0009] S4. Extract magnetic and acoustic signal features and standardize them;
[0010] S5. Construct a multimodal feature fusion model;
[0011] The magnetic signal features are used as query vectors and the acoustic spectrum features are used as key vectors. These are input into a multimodal attention structure. The intermodal correlation weights are calculated through a multi-head attention mechanism, and the acoustic features are reconstructed by weighting. This achieves the fusion of magnetic and acoustic signals in a unified representation space, resulting in a fused feature representation.
[0012] S6. Perform target identification and output the results;
[0013] The fused features are input into a lightweight classification structure, and the probability of target presence is output through multi-layer linear transformation and nonlinear mapping. Based on a set threshold, it is determined whether there is a shallowly buried explosive target in the current detection area, and the discrimination result is output.
[0014] Preferably, in step S2, after the magnetic signal is preprocessed, time alignment and background difference processing based on cross-correlation are performed to suppress environmental magnetic field interference; after the acoustic signal is preprocessed, the resonance response characteristics of the target structure are first obtained by frequency sweep excitation, and spatial positioning and distance gating processing are performed in combination with the FMCW signal.
[0015] Preferably, the magnetic signal processing flow is as follows:
[0016] Step 1: The TMR magnetic sensor acquires the magnetic field signal. The acquired analog signal is then amplified and low-pass filtered in multiple stages. Subsequently, the signal is converted into a narrowband measurement signal and subjected to analog-to-digital conversion to obtain a digital magnetic signal. ;
[0017] Let the first The gain of the stage amplifier is The overall magnification factor is:
[0018] ;
[0019] in This indicates the amplifier stage number. The output signals of each stage in the amplification process are introduced into a low-pass filter to reduce and suppress high-frequency noise.
[0020] Step 2: After Hampel outlier removal and digital filtering and smoothing, the digital signal is processed using cross-correlation-based time alignment and background reference differential processing. The current measured magnetic signal is... The background reference signal is Then the cross-correlation function between the two is:
[0021] ;
[0022] It is a cross-correlation function, representing the time delay. The degree of correlation between two signals under the condition of one sampling point Represents a discrete-time index. It is a discrete delay index, representing the offset of the signal in the sampling sequence;
[0023] when When, it indicates that the background reference signal is generated relative to the magnetic signal. Time offset of each sampling point; by finding the cross-correlation function The location of the maximum value allows us to obtain the optimal offset between the two signals.
[0024] ;
[0025] Time alignment is performed, followed by background differencing.
[0026] ;
[0027] in These are the amplitude matching coefficients, used to compensate for amplitude deviations under different sampling conditions; This indicates that the magnetic anomaly signal after background subtraction processing retains the local magnetic anomaly components caused by underground targets, thereby enhancing the salience of target features;
[0028] Step 3: After completing cross-correlation alignment and background subtraction, perform uniform time alignment and slicing. After obtaining magnetic signal segments of uniform length, perform data augmentation on each segment.
[0029] Step 4: After enhancement processing, the current magnetic signal segment is normalized and time-domain statistical features are extracted to describe the distribution characteristics and trends of the magnetic signal as a whole. Magnetic signal features are extracted from the frequency domain. After completing the extraction of time-domain and frequency-domain information, the system splices and fuses the features of each part to form a unified magnetic feature expression.
[0030] Preferably, Gaussian noise superposition, random amplitude scaling, and signal polarity reversal are used for signal enhancement; wherein, Gaussian noise superposition is used to simulate electronic device noise and slight environmental disturbances; random amplitude scaling is used to simulate signal strength changes caused by different detection distances, burial depths, or attitude changes; and signal polarity reversal is used to simulate polarity reversal caused by changes in sensor orientation or magnetic field orientation.
[0031] After enhancement processing, the current magnetic signal segment is normalized, resulting in the normalized magnetic signal time series. Treated as temporal feature vectors It is represented as:
[0032] ;
[0033] in, and These represent the mean and standard deviation of the current magnetic signal segment, respectively. A small constant introduced to prevent the denominator from being zero; the normalized time-domain sequence. It can itself serve as one of the time-domain representations of magnetic signals, used to preserve fine-grained information about how the original waveform changes over time;
[0034] Time-domain statistical feature extraction is used to describe the overall distribution characteristics and trends of magnetic signals. Time-domain statistical feature signals include mean, standard deviation, maximum value, minimum value, median, skewness, and kurtosis. The statistical feature vector composed of these statistical quantities is denoted as... .
[0035] Preferably, a fast Fourier transform is performed on the normalized magnetic signal segment to obtain its spectral representation:
[0036] ;
[0037] in, Representing the frequency index, the first 128 dimensions of the spectral amplitude are taken as the main frequency domain features, and the corresponding frequency domain features can be expressed as:
[0038] ;
[0039] After extracting information from the time and frequency domains, the data is concatenated and fused to form a unified magnetic feature representation, resulting in the final magnetic feature vector. Represented as:
[0040] ;
[0041] in, This represents the normalized time-domain sequence features. Represents the time-domain statistical eigenvector. This represents the frequency domain eigenvector obtained by the Fast Fourier Transform and then logarithmically compressed.
[0042] The preferred acoustic signal processing flow is as follows:
[0043] Step 1: Resonance Spectrum Feature Extraction;
[0044] Directional sound source transmits frequency sweep excitation signal :
[0045] ;
[0046] in, For time variables, The signal amplitude, This is the starting frequency of the resonant sweep. The resonant sweep slope is used to control the linear change of signal frequency over time.
[0047] Step 2: FMCW distance positioning and distance gating;
[0048] Introducing a range positioning mechanism based on FMCW (linear frequency modulated continuous wave), assuming the transmitted signal is:
[0049] ;
[0050] in For time variables, The signal amplitude, This is the starting frequency of the FMCW. The FMCW frequency modulation slope is used to obtain the beat frequency after the echo signal is demodulated. Then the target distance satisfies:
[0051] ;
[0052] in Given the speed of sound, the echo time interval corresponding to the target can be determined based on the mapping relationship between distance and echo time.
[0053] Let the distance gate interval be The corresponding time window is:
[0054] ;
[0055] The collected echo signal Convert to complete audio data sorted by time;
[0056] Step 3: Time slicing of the acoustic signal. The gated acoustic signal is divided into multiple time segments according to a fixed time window length, so that each time segment is consistent with the corresponding magnetic signal segment in the time dimension, thereby constructing a unified data sample.
[0057] Preferably, a short-time Fourier transform is performed on each audio signal sample to obtain the time-frequency representation of the audio signal:
[0058] ;
[0059] in, For time frame indexing, Angular frequency, Indicates the time of the sound signal With angular frequency The following time-frequency representation, This is a sliding window function used to extract signals within a specific time period. This represents the complex exponential basis function, used to implement frequency domain transformation;
[0060] After obtaining the time-frequency representation, to highlight the resonant frequency characteristics in the acoustic signal, a Mel filter bank is introduced to perform nonlinear frequency mapping on the spectrum and extract the Mel spectral features. ;
[0061] in Represents the Mel frequency dimension. Represents a time frame. ,in Represents the real number field. For frequency dimension, Represent a set of real numbers 3D matrix This refers to the number of time frames.
[0062] To enhance dynamic range adaptability, Perform logarithmic amplitude compression:
[0063] ;
[0064] After obtaining the Mel spectral features, data augmentation is performed on the spectral features to improve the model's robustness in complex environments. This includes random noise perturbation, local frequency perturbation, and temporal occlusion to simulate acoustic variations under different detection distances and environmental conditions. The augmented spectral features are then standardized to obtain the final acoustic feature vector. .
[0065] Preferably, in step S5, the magnetic and acoustic features are projected onto a representation space of the same dimension using a linear mapping function, thereby obtaining a feature representation of a unified dimension. This leads to the multimodal feature fusion stage, where a cross-attention mechanism is used to model the correlation between the magnetic and acoustic modal features. The steps are as follows:
[0066] Using magnetic features as the query vector and acoustic features as the key and value vectors, dynamic attention to acoustic modal information from magnetic modes is achieved. Definition:
[0067] ;
[0068] in Represents the query vector. Represents the key vector. Represents a value vector. , , The weight matrix is a learnable matrix;
[0069] The correlation weights between multimodal features are calculated based on the attention mechanism, and the expression is as follows:
[0070] ;
[0071] in This represents the similarity matrix between the query vector and the key vector. For feature dimension, This is a normalization function used to convert relevance into attention weights;
[0072] Through the aforementioned cross-attention mechanism, the correlation between magnetic and acoustic modes can be adaptively learned, thereby obtaining a fused multimodal feature representation. .
[0073] Preferably, in step S6, during the target discrimination stage, a multilayer perceptron structure is used to implement the classifier. This classifier consists of two layers of linear transformation and a nonlinear activation function, and the classification output is expressed as:
[0074] ;
[0075] in This represents the predicted probability of the target category. , This is the weight matrix. , For bias terms, It is a non-linear activation function. Used to convert the output into a probability distribution.
[0076] During the model training phase, the loss function is optimized and updated to obtain stable discrimination performance;
[0077] The loss function is expressed in the form of cross-entropy:
[0078] ;
[0079] Indicates the true label, This means that the model predicts the probability by minimizing the loss function, and finally outputs the probability prediction result of the presence of shallowly buried explosives.
[0080] A shallow-buried explosive detection system based on multimodal fusion includes a magnetic signal acquisition unit, an acoustic signal acquisition unit, a processing unit, and an output unit;
[0081] The magnetic signal acquisition unit includes a TMR magnetic sensor, which is used to collect magnetic anomaly signals caused by underground targets. The collected analog signals are successively amplified and low-pass filtered, and then converted into narrowband measurement signals and analog-to-digital conversion to obtain digital magnetic signals, which are then transmitted to the processing unit for processing.
[0082] The acoustic signal acquisition unit includes a directional sound source and a microphone array. The directional sound source is used to emit frequency-sweeping sound waves to the ground. The sound waves propagate in the underground medium and couple with the underground target structure to generate a vibration response, forming an echo signal. The echo signal is received by the microphone array and transmitted to the processing unit for processing.
[0083] The processing unit is used to receive signals collected by the magnetic sensor and microphone, and to perform multimodal fusion and target recognition processing internally;
[0084] The output unit outputs the detection results.
[0085] Compared with existing technologies, the beneficial effects are as follows:
[0086] 1) In terms of magnetic signal processing, to address the issue of magnetic signals being susceptible to interference from geomagnetic anomalies and surface metallic debris in complex environments, a high-sensitivity magnetic sensor based on the tunnel magnetoresistance effect is employed for weak magnetic anomaly detection, and a magnetic anomaly detection and background suppression mechanism is constructed. By standardizing the magnetic signal and extracting frequency domain features, joint time-domain and frequency-domain representation is achieved to improve the identifiability of the magnetic anomaly signal. Simultaneously, to address the interference caused by slow background magnetic field drift and time misalignment during motion, a cross-correlation-based time alignment and background reference difference suppression mechanism is introduced to reduce the impact of environmental magnetic fields and system drift on the target anomaly signal.
[0087] 2) In terms of acoustic signal processing, addressing the issues of attenuation, scattering, and multipath reflection affecting acoustic signal propagation in complex geological media, a swept-frequency acoustic wave excitation method is first employed to transmit continuously varying frequency acoustic signals to the target area, thereby stimulating the vibration response of the underground target structure. The echo signals generated by the swept-frequency signal are acquired through multi-channel microphones, and frame-by-frame processing and spectral analysis are performed to extract features such as spectral energy distribution, main peak frequency position, and frequency band energy ratio, thus characterizing the frequency response characteristics of the target structure. Simultaneously, to improve spatial resolution in complex environments, an FMCW (linear frequency modulated) continuous wave signal is transmitted, and its echo is subjected to beat demodulation and range spectrum construction to determine the target's spatial location. Spatial gating is implemented based on the mapping relationship between distance and echo time, and the gated swept-frequency echo signal undergoes time-frequency analysis and spectral enhancement processing to reduce the impact of non-target reflections and environmental interference on resonance feature extraction.
[0088] 3) In terms of multimodal fusion, considering the differences between magnetic and acoustic signals in sampling rate, data dimension and feature distribution, the two types of signals are processed in a unified time slice to achieve time axis alignment; a multimodal feature fusion mechanism is constructed at the feature layer to map magnetic and acoustic features to a unified representation space, and the correlation model between modes is realized through a cross-attention structure, thereby achieving effective fusion of multimodal features. Attached Figure Description
[0089] Figure 1 A schematic diagram of the process for a magnetoacoustic multimodal shallow-buried explosive detection method;
[0090] Figure 2 A schematic diagram of the deployment of a magnetoacoustic multimodal shallow-buried explosive detection system;
[0091] Figure 3 This is a schematic diagram of the magnetic signal processing flow.
[0092] Figure 4 This is a schematic diagram of the acoustic signal processing flow. Detailed Implementation
[0093] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0094] A method for detecting shallowly buried explosives based on multimodal fusion includes the following steps:
[0095] S1. Synchronous acquisition of magnetic and acoustic signals;
[0096] A highly sensitive magnetic sensor based on the tunnel magnetoresistance effect is deployed above the area to be measured to detect local magnetic field disturbances generated by underground targets. The analog voltage signal output by the magnetic sensor is amplified sequentially by a multi-stage operational amplifier circuit, and high-frequency noise is suppressed by a low-pass filter before being input into an analog-to-digital converter for digital processing. Simultaneously, a directional acoustic wave transmitter is set up in the area to be measured to emit a swept-frequency acoustic excitation signal to stimulate the resonance characteristics of the underground target structure, and an FMCW linear frequency modulated continuous wave signal for target distance identification. Acoustic echo signals reflected from the ground surface are collected using a microphone array.
[0097] S2. Preprocess the magnetic and acoustic signals;
[0098] After preprocessing, the magnetic signal undergoes time alignment and background difference processing based on cross-correlation to suppress environmental magnetic field interference. After preprocessing, the acoustic signal is subjected to frequency sweep excitation to obtain the resonant response characteristics of the target structure, and then combined with the FMCW signal for spatial positioning and distance gating.
[0099] S3. Time alignment and data slicing of magnetic and acoustic signals; the magnetic and acoustic signals are segmented according to a unified time window, the length of each time window is a pre-set fixed time period, and the magnetic signal segments and acoustic signal segments within the same time window are matched accordingly.
[0100] The magnetic signal processing flow is as follows:
[0101] Step 1: The TMR magnetic sensor acquires the magnetic field signal. The acquired analog signal is then amplified and low-pass filtered in multiple stages. Subsequently, the signal is converted into a narrowband measurement signal and subjected to analog-to-digital conversion to obtain a digital magnetic signal. ;
[0102] Let the first The gain of the stage amplifier is The overall magnification factor is:
[0103] ;
[0104] in This indicates the amplifier stage. The output signals of each stage in the amplification process are introduced into a low-pass filter to reduce and suppress high-frequency noise.
[0105] Step 2: After Hampel outlier removal and digital filtering and smoothing, the digital signal is processed using cross-correlation-based time alignment and background reference differential processing. The current measured magnetic signal is... The background reference signal is Then the cross-correlation function between the two is:
[0106] ;
[0107] It is a cross-correlation function, representing the time delay. The degree of correlation between two signals under the condition of one sampling point Represents a discrete-time index. It is a discrete delay index, representing the offset of the signal in the sampling sequence;
[0108] when When, it indicates that the background reference signal is generated relative to the magnetic signal. Time offset of each sampling point; by finding the cross-correlation function The location of the maximum value allows us to obtain the optimal offset between the two signals.
[0109] ;
[0110] Time alignment is performed, followed by background differencing.
[0111] ;
[0112] in These are the amplitude matching coefficients, used to compensate for amplitude deviations under different sampling conditions; This indicates that the magnetic anomaly signal after background subtraction processing retains the local magnetic anomaly components caused by underground targets, thereby enhancing the salience of target features;
[0113] Step 3: After completing cross-correlation alignment and background subtraction, perform uniform time alignment and slicing. After obtaining magnetic signal segments of uniform length, perform data augmentation on each segment.
[0114] Step 4: After enhancement processing, the current magnetic signal segment is normalized and time-domain statistical features are extracted to describe the distribution characteristics and trends of the magnetic signal as a whole. Magnetic signal features are extracted from the frequency domain. After completing the extraction of time-domain and frequency-domain information, the system splices and fuses the features of each part to form a unified magnetic feature expression.
[0115] The sound signal processing flow is as follows:
[0116] Step 1: Resonance Spectrum Feature Extraction;
[0117] Directional sound source transmits frequency sweep excitation signal :
[0118] ;
[0119] in, For time variables, The signal amplitude, This is the starting frequency of the resonant sweep. The resonant sweep slope is used to control the linear change of signal frequency over time.
[0120] Step 2: FMCW distance positioning and distance gating;
[0121] Introducing a range positioning mechanism based on FMCW (linear frequency modulated continuous wave), assuming the transmitted signal is:
[0122] ;
[0123] in For time variables, The signal amplitude, This is the starting frequency of the FMCW. The FMCW frequency modulation slope is used to obtain the beat frequency after the echo signal is demodulated. Then the target distance satisfies:
[0124] ;
[0125] in Given the speed of sound, the echo time interval corresponding to the target can be determined based on the mapping relationship between distance and echo time.
[0126] Let the distance gate interval be The corresponding time window is:
[0127] ;
[0128] The collected echo signal Convert to complete audio data sorted by time;
[0129] Step 3: Time-slicing processing of the acoustic signal;
[0130] The gated acoustic signal is divided into multiple time segments according to a fixed time window length, so that each time segment is consistent with the corresponding magnetic signal segment in the time dimension, thereby constructing a unified data sample.
[0131] S4. Extract magnetic and acoustic signal features and standardize them;
[0132] For magnetic signals: After obtaining magnetic signal segments of uniform length, data augmentation processing is performed on each segment; after augmentation processing, the current magnetic signal segment is normalized, and time-domain statistical features are extracted to describe the distribution characteristics and trends of the magnetic signal as a whole. Magnetic signal features are extracted from the frequency domain. After completing the extraction of time-domain and frequency-domain information, the system splices and fuses the features of each part to form a unified magnetic feature expression.
[0133] Signal enhancement is achieved by using Gaussian noise superposition, random amplitude scaling, and signal polarity reversal. Gaussian noise superposition is used to simulate electronic equipment noise and slight environmental disturbances. Random amplitude scaling is used to simulate signal strength changes caused by different detection distances, burial depths, or attitude changes. Signal polarity reversal is used to simulate polarity reversal caused by changes in sensor orientation or magnetic field orientation.
[0134] After enhancement processing, the current magnetic signal segment is normalized, resulting in the normalized magnetic signal time series. Treated as time-domain feature vectors It is represented as:
[0135] ;
[0136] in, and These represent the mean and standard deviation of the current magnetic signal segment, respectively. A small constant introduced to prevent the denominator from being zero; the normalized time-domain sequence. It can itself serve as one of the time-domain representations of magnetic signals, used to preserve fine-grained information about how the original waveform changes over time;
[0137] Time-domain statistical feature extraction is used to describe the overall distribution characteristics and trends of magnetic signals. Time-domain statistical feature signals include mean, standard deviation, maximum value, minimum value, median, skewness, and kurtosis. The statistical feature vector composed of these statistical quantities is denoted as... .
[0138] Performing a Fast Fourier Transform on the normalized magnetic signal segment yields its spectral representation:
[0139] ;
[0140] in, Representing the frequency index, the first 128 dimensions of the spectral amplitude are taken as the main frequency domain features, and the corresponding frequency domain features can be expressed as:
[0141] ;
[0142] After extracting information from the time and frequency domains, the data is concatenated and fused to form a unified magnetic feature representation, resulting in the final magnetic feature vector. Represented as:
[0143] ;
[0144] in, This represents the normalized time-domain sequence features. Represents the time-domain statistical eigenvector. This represents the frequency domain eigenvector obtained by the Fast Fourier Transform and then logarithmically compressed.
[0145] For acoustic signals: After time-slicing the acoustic signal, a short-time Fourier transform is performed on each audio signal sample to obtain the time-frequency representation of the acoustic signal.
[0146] ;
[0147] in, For time frame indexing, Angular frequency, Indicates the time of the sound signal With angular frequency The following time-frequency representation, This is a sliding window function used to extract signals within a specific time period. This represents the complex exponential basis function, used to implement frequency domain transformation;
[0148] After obtaining the time-frequency representation, to highlight the resonant frequency characteristics in the acoustic signal, a Mel filter bank is introduced to perform nonlinear frequency mapping on the spectrum and extract the Mel spectral features. ;
[0149] in Represents the Mel frequency dimension. Represents a time frame. ,in Represents the real number field. For frequency dimension, Represent a set of real numbers 3D matrix This refers to the number of time frames.
[0150] To enhance dynamic range adaptability, Perform logarithmic amplitude compression:
[0151] ;
[0152] After obtaining the Mel spectral features, data augmentation is performed on the spectral features to improve the model's robustness in complex environments. This includes random noise perturbation, local frequency perturbation, and temporal occlusion to simulate acoustic variations under different detection distances and environmental conditions. The augmented spectral features are then standardized to obtain the final acoustic feature vector. .
[0153] S5. Construct a multimodal feature fusion model;
[0154] The magnetic signal features are used as query vectors and the acoustic spectrum features are used as key vectors. These are input into a multimodal attention structure. The intermodal correlation weights are calculated through a multi-head attention mechanism, and the acoustic features are reconstructed by weighting. This achieves the fusion of magnetic and acoustic signals in a unified representation space, resulting in a fused feature representation.
[0155] Using magnetic features as the query vector and acoustic features as the key and value vectors, dynamic attention to acoustic modal information from magnetic modes is achieved. Definition:
[0156] ;
[0157] in Represents the query vector. Represents the key vector. Represents a value vector. , , The weight matrix is a learnable matrix;
[0158] The correlation weights between multimodal features are calculated based on the attention mechanism, and the expression is as follows:
[0159] ;
[0160] in This represents the similarity matrix between the query vector and the key vector. For feature dimension, This is a normalization function used to convert relevance into attention weights;
[0161] Through the aforementioned cross-attention mechanism, the correlation between magnetic and acoustic modes can be adaptively learned, thereby obtaining a fused multimodal feature representation. .
[0162] S6. Perform target identification and output the results;
[0163] The fused features are input into a lightweight classification structure, and the probability of target presence is output through multi-layer linear transformation and nonlinear mapping. Based on a set threshold, it is determined whether there is a shallowly buried explosive target in the current detection area, and the discrimination result is output.
[0164] In the target discrimination stage, a multilayer perceptron structure is used to implement the classifier. This classifier consists of two linear transformations and a nonlinear activation function, and the classification output is expressed as follows:
[0165] ;
[0166] in This represents the predicted probability of the target category. , This is the weight matrix. , For bias terms, It is a non-linear activation function. Used to convert the output into a probability distribution.
[0167] During the model training phase, the loss function is optimized and updated to obtain stable discrimination performance;
[0168] The loss function is expressed in the form of cross-entropy:
[0169] ;
[0170] Indicates the true label, This means that the model predicts the probability by minimizing the loss function, and finally outputs the probability prediction result of the presence of shallowly buried explosives.
[0171] A shallow-buried explosive detection system based on multimodal fusion includes a magnetic signal acquisition unit, an acoustic signal acquisition unit, a processing unit, and an output unit;
[0172] The magnetic signal acquisition unit includes a TMR magnetic sensor, which is used to collect magnetic anomaly signals caused by underground targets. The collected analog signals are successively amplified and low-pass filtered, and then converted into narrowband measurement signals and analog-to-digital conversion to obtain digital magnetic signals, which are then transmitted to the processing unit for processing.
[0173] The acoustic signal acquisition unit includes a directional sound source and a microphone array. The directional sound source is used to emit frequency-sweeping sound waves to the ground. The sound waves propagate in the underground medium and couple with the underground target structure to generate a vibration response, forming an echo signal. The echo signal is received by the microphone array and transmitted to the processing unit for processing.
[0174] The processing unit is used to receive signals collected by the magnetic sensor and microphone, and to perform multimodal fusion and target recognition processing internally;
[0175] The output unit outputs the detection results.
[0176] like Figure 1 As shown, the magnetoacoustic multimodal shallow-buried explosive detection method of the present invention is constructed according to a hierarchical and progressive processing logic. The overall process consists of a signal acquisition stage, a signal processing stage, a feature construction stage, a multimodal fusion stage, and a target discrimination stage in sequence. Each stage is connected sequentially through a data interface to form a complete processing link.
[0177] During system operation, magnetic and acoustic signals are first synchronously acquired using a magnetic sensor and a microphone array. The magnetic sensor detects magnetic anomaly changes caused by underground targets, while the microphone array acquires underground echo signals generated after sound source excitation, thereby obtaining raw data reflecting different physical characteristics of underground targets.
[0178] After data acquisition, the magnetic and acoustic signals enter their respective preprocessing paths. The magnetic signal first enters the magnetic signal processing module, where filtering, noise reduction, and signal smoothing improve signal quality. Subsequently, the processing unit performs cross-correlation alignment and background difference processing on the magnetic signal to eliminate the effects of slow drift of the environmental magnetic field and system noise, thereby enhancing the response characteristics of magnetic anomalies caused by underground targets.
[0179] Meanwhile, the acoustic signal enters the acoustic signal processing module. First, through distance positioning and spatial gating processing, the effective echo region corresponding to the target spatial location is selected from the continuous acoustic echo signal, thereby suppressing interference caused by multipath reflections and echoes from non-target areas.
[0180] After completing the above processing, the processing unit enters the unified time alignment and slicing stage for magnetic and acoustic signals. This stage synchronously divides the data in the two paths through a unified time window, so that the magnetic and acoustic signals form corresponding data segments within the same time period, thereby ensuring the consistency of the two modal data in the time dimension.
[0181] After time alignment, the two signal paths proceed to the feature construction stage. For the magnetic signal path, data augmentation is first performed to simulate magnetic field variations under different environmental conditions, improving the system's adaptability to complex environments. Subsequently, time-domain and frequency-domain features are extracted from the augmented magnetic signal, and the obtained magnetic features are standardized to eliminate dimensional differences between different features.
[0182] For the acoustic signal path, resonance spectrum features are first extracted from the acoustic signal. By analyzing the energy distribution of the echo signal at different frequencies, acoustic spectrum features that reflect the resonance characteristics of the underground target structure are obtained. Subsequently, the extracted acoustic spectrum features are enhanced to improve the robustness of the system under different detection conditions, and the acoustic features are standardized to ensure that the feature scales are consistent across different samples.
[0183] After standardizing the two modal features, the processing unit projects the magnetic and acoustic features into a unified feature space through the feature dimension mapping module, so that different modal features can be correlated and modeled in the same representation space.
[0184] The process then proceeds to the multimodal feature fusion stage. In this stage, the correlation between magnetic and acoustic features is modeled through a multimodal feature association mechanism, achieving deep fusion of the two modal information to obtain a fused feature representation that simultaneously contains magnetic anomaly information and acoustic resonance information.
[0185] Finally, in the target discrimination stage, the fused features are input into the classification model for target identification, outputting the probability or category of the underground target's presence. Through this multimodal fusion discrimination method, the system can comprehensively utilize both magnetic anomaly response and acoustic resonance response information, improving the detection accuracy and anti-interference capability of shallowly buried explosive targets under complex environmental conditions.
[0186] In this specification, the present invention has been described with reference to specific embodiments. These embodiments are preferred embodiments of the present patent and are not intended to limit the scope of the invention. It should be noted that the present invention is not limited to the specific embodiments described above. Improvements, variations, combinations, substitutions, etc., made by those skilled in the art without departing from the principles of the present invention are all within the scope of protection claimed in the claims of the present invention.
[0187] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.
Claims
1. A method for detecting shallowly buried explosives based on multimodal fusion, characterized in that, Includes the following steps: S1. Synchronous acquisition of magnetic and acoustic signals; S2. Preprocess the magnetic and acoustic signals; S3. Time alignment and data slicing of magnetic and acoustic signals; the magnetic and acoustic signals are segmented according to a unified time window, the length of each time window is a pre-set fixed time period, and the magnetic signal segments and acoustic signal segments within the same time window are matched accordingly. S4. Extract magnetic and acoustic signal features and standardize them; The magnetic signal processing flow is as follows: Step 1: The TMR magnetic sensor acquires the magnetic field signal. The acquired analog signal is then amplified and low-pass filtered in multiple stages to convert it into a narrowband measurement signal, which is then converted from analog to digital to obtain a digital magnetic signal. ; Step 2: After Hampel outlier removal and digital filtering and smoothing, the digital signal is processed using cross-correlation-based time alignment and background reference differential processing. The current measured magnetic signal is... The background reference signal is Then the cross-correlation function between the two is: ; It is a cross-correlation function, representing the time delay. The degree of correlation between two signals under the condition of one sampling point Represents a discrete-time index. It is a discrete delay index, representing the offset of the signal in the sampling sequence; when When, it indicates that the background reference signal is generated relative to the magnetic signal. Time offset of each sampling point; by finding the cross-correlation function The location of the maximum value allows us to obtain the optimal offset between the two signals. ; Time alignment is performed, followed by background differencing. ; in These are the amplitude matching coefficients, used to compensate for amplitude deviations under different sampling conditions; This indicates that the magnetic anomaly signal after background subtraction processing retains the local magnetic anomaly components caused by underground targets, thereby enhancing the salience of target features; Step 3: After completing cross-correlation alignment and background subtraction, perform uniform time alignment and slicing. After obtaining magnetic signal segments of uniform length, perform data augmentation on each segment. Step 4: After enhancement processing, the current magnetic signal segment is normalized and time-domain statistical features are extracted to describe the distribution characteristics and trends of the magnetic signal as a whole. Magnetic signal features are extracted from the frequency domain. After the extraction of time-domain and frequency-domain information, the features of each part are spliced and fused to form a unified magnetic feature expression. The sound signal processing flow is as follows: Step 1: Resonance Spectrum Feature Extraction; Directional sound source transmits frequency sweep excitation signal : ; in, For time variables, The signal amplitude, This is the starting frequency of the resonant sweep. The resonant sweep slope is used to control the linear change of signal frequency over time. Step 2: FMCW distance positioning and distance gating; Introducing a range positioning mechanism based on FMCW (linear frequency modulated continuous wave), assuming the transmitted signal is: ; in For time variables, The signal amplitude, This is the starting frequency of the FMCW. The FMCW frequency modulation slope is used to obtain the beat frequency after the echo signal is demodulated. Then the target distance satisfies: ; in Given the speed of sound, the echo time interval corresponding to the target can be determined based on the mapping relationship between distance and echo time. Let the distance gate interval be The corresponding time window is: ; The collected echo signal Convert to complete audio data sorted by time; Step 3: Time-slicing processing of the acoustic signal; The gated acoustic signal is divided into multiple time segments according to a fixed time window length, so that each time segment is consistent with the corresponding magnetic signal segment in the time dimension, thereby constructing a unified data sample; S5. Construct a multimodal feature fusion model; The magnetic signal features are used as query vectors and the acoustic spectrum features are used as key vectors. These are input into a multimodal attention structure. The intermodal correlation weights are calculated through a multi-head attention mechanism, and the acoustic features are reconstructed by weighting. This achieves the fusion of magnetic and acoustic signals in a unified representation space, resulting in a fused feature representation. S6. Perform target identification and output the results; The fused features are input into a lightweight classification structure, and the probability of target presence is output through multi-layer linear transformation and nonlinear mapping. Based on a set threshold, it is determined whether there is a shallowly buried explosive target in the current detection area, and the discrimination result is output.
2. The method for detecting shallowly buried explosives based on multimodal fusion according to claim 1, characterized in that, In step S2, after the magnetic signal is preprocessed, it undergoes time alignment and background difference processing based on cross-correlation to suppress environmental magnetic field interference. After preprocessing, the acoustic signal is used to obtain the resonant response characteristics of the target structure through frequency sweep excitation, and then combined with the FMCW signal for spatial positioning and distance gating.
3. The method for detecting shallowly buried explosives based on multimodal fusion according to claim 1, characterized in that, Signal enhancement is achieved by using Gaussian noise superposition, random amplitude scaling, and signal polarity reversal. Gaussian noise superposition is used to simulate electronic equipment noise and slight environmental disturbances. Random amplitude scaling is used to simulate signal strength changes caused by different detection distances, burial depths, or attitude changes. Signal polarity reversal is used to simulate polarity reversal caused by changes in sensor orientation or magnetic field orientation. After enhancement processing, the current magnetic signal segment is normalized, resulting in the normalized magnetic signal time series. Treated as time-domain feature vectors It is represented as: ; in, and These represent the mean and standard deviation of the current magnetic signal segment, respectively. A small constant introduced to prevent the denominator from being zero; the normalized time-domain sequence. It can be used as one of the time-domain representations of magnetic signals to preserve fine-grained information about how the original waveform changes over time; Time-domain statistical feature extraction is used to describe the overall distribution characteristics and trends of magnetic signals. Time-domain statistical feature signals include mean, standard deviation, maximum value, minimum value, median, skewness, and kurtosis. The statistical feature vector composed of these statistical quantities is denoted as... .
4. The method for detecting shallowly buried explosives based on multimodal fusion according to claim 3, characterized in that, Performing a Fast Fourier Transform on the normalized magnetic signal segment yields its spectral representation: ; in, Representing the frequency index, the first 128 dimensions of the spectral amplitude are taken as the main frequency domain features, and the corresponding frequency domain features can be expressed as: ; After extracting information from the time and frequency domains, the data is concatenated and fused to form a unified magnetic feature representation, resulting in the final magnetic feature vector. Represented as: ; in, This represents the normalized time-domain sequence features. Represents the time-domain statistical eigenvector. This represents the frequency domain eigenvector obtained by the Fast Fourier Transform and then logarithmically compressed.
5. The method for detecting shallowly buried explosives based on multimodal fusion according to claim 1, characterized in that, Perform a short-time Fourier transform on each audio signal sample to obtain the time-frequency representation of the audio signal: ; in, For time frame indexing, Angular frequency, Indicates the time of the sound signal With angular frequency The following time-frequency representation, This is a sliding window function used to extract signals within a specific time period. This represents the complex exponential basis function, used to implement frequency domain transformation; After obtaining the time-frequency representation, to highlight the resonant frequency characteristics in the acoustic signal, a Mel filter bank is introduced to perform nonlinear frequency mapping on the spectrum and extract the Mel spectral features. ; in Represents the Mel frequency dimension. Represents a time frame. ,in Represents the real number field. For frequency dimension, Represent a set of real numbers 3D matrix Number of time frames; right Perform logarithmic amplitude compression: ; After obtaining the Mel spectral features, data augmentation processing is performed on the spectral features to simulate the acoustic variation characteristics under different detection distances and environmental conditions. The augmented spectral features are then standardized to obtain the final acoustic feature vector. .
6. The method for detecting shallowly buried explosives based on multimodal fusion according to claim 1, characterized in that, Step S5 projects the magnetic and acoustic features onto a representation space of the same dimension using a linear mapping function, thereby obtaining a unified dimension feature representation. This leads to the multimodal feature fusion stage, where a cross-attention mechanism is used to model the correlation between the magnetic and acoustic modal features. The steps are as follows: Using magnetic features as the query vector and acoustic features as the key and value vectors, dynamic attention to acoustic modal information from magnetic modes is achieved. Definition: ; in Represents the query vector. Represents the key vector. Represents a value vector. , , The weight matrix is a learnable matrix; The correlation weights between multimodal features are calculated based on the attention mechanism, and the expression is as follows: ; in This represents the similarity matrix between the query vector and the key vector. For feature dimension, This is a normalization function used to convert relevance into attention weights; By employing a cross-attention mechanism, the correlation between magnetic and acoustic modes is adaptively learned, thereby obtaining a fused multimodal feature representation. .
7. The method for detecting shallowly buried explosives based on multimodal fusion according to claim 6, characterized in that, Step S6: In the target discrimination stage, a multilayer perceptron structure is used to implement the classifier. The classifier consists of two layers of linear transformation and a nonlinear activation function. The classification output is expressed as: ; in This represents the predicted probability of the target category. , This is the weight matrix. , For bias terms, It is a non-linear activation function. Used to convert the output into a probability distribution; During the model training phase, the loss function is optimized and updated to obtain stable discrimination performance; The loss function is expressed in the form of cross-entropy: ; Indicates the true label, This means that the model predicts the probability by minimizing the loss function, and finally outputs the probability prediction result of the presence of shallowly buried explosives.
8. A shallow-buried explosive detection system based on multimodal fusion, employing the shallow-buried explosive detection method based on multimodal fusion as described in any one of claims 1-7, characterized in that, It includes a magnetic signal acquisition unit, an acoustic signal acquisition unit, a processing unit, and an output unit; The magnetic signal acquisition unit includes a TMR magnetic sensor, which is used to collect magnetic anomaly signals caused by underground targets. The collected analog signals are successively amplified and low-pass filtered, and then converted into narrowband measurement signals and analog-to-digital conversion to obtain digital magnetic signals, which are then transmitted to the processing unit for processing. The sound signal acquisition unit includes a directional sound source and a microphone array; A directional sound source is used to emit frequency-sweeping sound waves into the ground. The sound waves propagate in the underground medium and couple with the underground target structure to generate a vibration response, forming an echo signal. The echo signal is received by a microphone array and transmitted to a processing unit for processing. The processing unit is used to receive signals collected by the magnetic sensor and microphone, and to perform multimodal fusion and target recognition processing internally; The output unit outputs the detection results.