Box-type space hiding person detection method and system based on multi-modal feature fusion and signal purification

By combining a laser vibrometer with a vibration sensor, and employing wavelet decomposition, empirical mode decomposition, and cross-modal feature decoupling modules, high-precision detection of personnel hiding in box-shaped spaces was achieved. This solved the problems of low detection accuracy and signal contamination, and improved the accuracy of detection and feature extraction.

CN121901797APending Publication Date: 2026-04-21HUNAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN UNIV
Filing Date
2026-01-12
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing methods for detecting people hiding in box-shaped spaces are greatly affected by external conditions in scenarios with high accuracy requirements, leading to reduced detection accuracy, prominent false detections and false negatives, difficulties in signal extraction, insufficient feature extraction, and severe signal contamination in laser vibration measurement methods.

Method used

By combining a laser vibrometer with a vibration sensor, multimodal feature fusion and signal purification are performed through wavelet decomposition, binary ensemble empirical mode decomposition, MLP-Integrator module and cascaded cross-modal feature decoupling module to extract and separate human physiological signals and environmental vibration signals.

Benefits of technology

It improves detection accuracy, reduces false detections and missed detections, accurately extracts human physiological information, solves the signal contamination problem, and achieves high-precision detection of people hiding in box-like spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901797A_ABST
    Figure CN121901797A_ABST
Patent Text Reader

Abstract

The invention discloses a box-type space hiding person detection method based on multi-modal feature fusion and signal purification, and the method comprises the steps: collecting a surface vibration signal of a to-be-detected target through a laser vibration meter, and collecting an overall vibration signal of the to-be-detected target through a vibration sensor, and inputting the collected surface vibration signal and the overall vibration signal into a pre-trained human physiological feature vibration detection model to obtain a detection result. The technical problems that when a traditional detection method is applied to a personnel hiding detection scene with a high precision requirement, due to the fact that the detection target is high in concealment and is greatly influenced by external conditions, the detection precision is remarkably reduced, and the problems of false detection and missing detection are prominent can be solved; and the technical problems of difficult signal extraction and insufficient feature extraction in the existing contact-based physical vibration detection method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning and security detection technology, and more specifically, relates to a method and system for detecting people hiding in box-shaped spaces based on multimodal feature fusion and signal cleanup. Background Technology

[0002] Assisting customs and border inspection personnel in effectively detecting individuals concealed within container spaces has become a key challenge in the security field.

[0003] Existing methods for detecting people hiding in box-like spaces can be categorized into three main types: The first type is traditional detection technology, which relies on direct perception or indirect inference. Examples include thermal imaging technology capturing infrared radiation from the human body, acoustic detection depending on the signal reflection of sound waves, and gas concentration analysis inferring the presence of people by detecting abnormal concentrations of metabolic gases such as CO2. The second type is vibration analysis-based detection technology, which has the advantages of being non-invasive and universally applicable. For example, contact-based physical vibration detection involves directly fixing vibration sensors to the surface of the box to collect weak vibration signals transmitted to the box by human respiration and heartbeat. The third type is non-contact laser vibration measurement technology, which utilizes the laser Doppler effect. A laser vibrometer emits a laser beam and receives its reflected light to accurately measure the vibration parameters of the box surface and analyzes these parameters to obtain information about the interior of the box.

[0004] However, the above-mentioned methods for detecting people hiding in box-shaped spaces all have some drawbacks: (1) When the above traditional detection methods are applied to the scenario of personnel concealment detection with high accuracy requirements, the disadvantages of the detection target being highly concealed and greatly affected by external conditions (such as ambient temperature fluctuations, traffic noise, ventilation conditions, etc.) become prominent, resulting in a significant reduction in detection accuracy and prominent problems of false detection and missed detection. (2) The above-mentioned physical vibration detection method based on contact is too idealistic in its implementation and too demanding in its installation conditions. The coupling stability between the sensor and the box with different materials and curvatures is difficult to guarantee, which causes a large amount of environmental vibration information to be mixed into the target vibration information, resulting in difficulty in signal extraction and insufficient feature extraction. (3) In the above-mentioned non-contact laser vibration measurement method, the signal transmission process from the signal source to the surface of the box is not a simple signal copying, but a complex process affected by the characteristics of the transmission path, resonance phenomenon, multi-path superposition, nonlinear factors and environmental noise. This results in the obtained vibration signal on the surface of the box not only including human physiological signals, but also signal components that are contaminated by environmental vibration signals but still contain physiological information components, thus affecting the accuracy of the detection results. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a method and system for detecting personnel concealment in enclosures by combining a laser vibrometer and a vibration sensor. Its purpose is to solve the technical problems of traditional detection methods, which, when applied to high-precision personnel concealment detection scenarios, suffer from significantly reduced detection accuracy due to the high concealment of the target and the significant influence of external conditions, resulting in prominent false detections and missed detections. It also addresses the technical problems of existing contact-based physical vibration detection methods, such as difficulties in signal extraction and insufficient feature extraction, and the technical problem that existing non-contact laser vibrometer methods obtain vibration signals from the enclosure surface that include not only human physiological signals but also signal components contaminated by environmental vibration signals but still containing physiological information, thus affecting the accuracy of the detection results.

[0006] To achieve the above objectives, according to one aspect of the present invention, a method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing is provided, comprising the following steps: Step 1: Use a laser vibrometer to collect the surface vibration signal of the target to be tested, and use a vibration sensor to collect the overall vibration signal of the target to be tested; Step 2: Input the surface vibration signal and overall vibration signal collected in Step 1 into the pre-trained human physiological characteristic vibration detection model to obtain the detection results.

[0007] Preferably, the surface vibration signal of the target to be detected includes the natural vibration of the environment and the target to be detected, as well as the weak vibration signal generated by the transmission of physiological signals from the human body to the surface of the target to be detected; The overall vibration signal of the target under test is the mechanical vibration signal generated by the natural vibration of the environment and the target under test; Step one is as follows: First, calibrate the laser vibrometer and set up a test scene within 1 meter of the target to be tested and sample for more than 3 seconds to collect the surface vibration signal of the target to be tested; then, install the vibration sensor at multiple key positions of the target to be tested according to the five-point sampling method, and ensure that it covers the main vibration area of ​​the target to be tested, in order to collect the overall vibration signal of the target to be tested.

[0008] Preferably, the human physiological characteristic vibration detection model comprises four parts connected in sequence: a wavelet decomposition module, a binary set empirical mode decomposition module, an MLP-Integrator module, and a cascaded cross-modal feature decoupling module; The specific structure of the wavelet decomposition module is as follows: The first layer is the initial decomposition layer, whose input is the original signal a of surface vibration. This layer first uses the db4 wavelet basis to perform low-pass filtering and 2x downsampling on the input original signal a. Then, the downsampling result is subjected to linear interpolation to obtain the first-layer approximate component A1. Subsequently, the first-layer approximate component A1 is subjected to high-pass filtering and 2x downsampling. Finally, the downsampling result is subjected to linear interpolation to obtain the high-frequency subband signal component D1. The second layer is the depth decomposition layer. Its input is the first-layer approximation component A1 output by the first layer. This layer first performs low-pass filtering and 2x downsampling on the first-layer approximation component A1 through the db4 wavelet basis, and then performs cubic spline interpolation on the downsampling result to obtain the low-frequency subband signal component A2. Then, the low-frequency subband signal component A2 is subjected to high-pass filtering and 2x downsampling, and then the downsampling result is subjected to cubic spline interpolation to obtain the mid-high frequency subband signal component D2. The input to the binary set empirical mode decomposition module is the low-frequency subband signal component A2 obtained from the wavelet decomposition module, and the overall vibration signal x(t) of the target at time t collected by the laser vibrometer. This module first preprocesses the overall vibration signal x(t) of the target at time t, then fuses the preprocessed result with the low-frequency subband signal component A2. Subsequently, it performs multiple empirical mode decomposition processes on the fused result to obtain multiple intrinsic mode components (IMFs). Finally, it selects the lowest-frequency IMFs from all IMFs and fuses them into the environmental vibration signal component F. env And output it.

[0009] Preferably, the processing procedure of the binary set empirical mode decomposition module includes the following sub-steps: (2-1) The overall vibration signal x(t) of the target under test at time t is subjected to low-frequency filtering using a Butterworth low-pass filter, and the overall vibration signal after low-frequency filtering is processed by Min-Max normalization to obtain the main signal z of the first cycle of the first intrinsic mode component at time t. 1,1 (t); (2-2) Set counter k=1 and counter i=1; (2-3) Inject the time sequence A2(t) corresponding to the low-frequency subband signal component A2 obtained by the wavelet decomposition module at time t into the main signal z of the i-th cycle of the k-th eigenmode component at time t. k,i In (t), the mixed signal x of the i-th cycle of the k-th intrinsic mode component at time t is obtained. k,i (t): ; Where λ is the amplitude coefficient, and λ∈[0.1, 0.3]; (2-4) Obtain the mixed signal x of the k-th intrinsic mode component obtained in step (2-3) during the i-th cycle. k,i All time sampling points in (t) are obtained, and the local maxima and local minima corresponding to each time sampling point are obtained; (2-5) Based on the local maxima and local minima corresponding to all time sampling points obtained in step (2-4), and using cubic spline interpolation, fit the upper and lower envelopes at time t, and calculate the mean amplitude of the upper and lower envelopes at time t to obtain the mean envelope m of the k-th intrinsic mode component in the i-th cycle at time t. k,i (t); (2-6) For the main signal z of the i-th cycle of the k-th intrinsic mode component at time t k,i (t) and the mean envelope m of the k-th intrinsic mode component in the i-th cycle obtained in step (2-5) at time t k,i (t) Perform subtraction operation on each time sampling point to obtain the difference signal h of the k-th intrinsic mode component in the i-th cycle at time t. k,i (t); (2-7) Determine whether the difference signal h of the k-th intrinsic mode component in the i-th cycle at time t obtained in step (2-6) exists. k,i The difference between the number of local extrema and the number of zero crossings of (t) is less than or equal to 1, and the mean envelope m k,i The amplitude of (t) is less than or equal to the threshold ε; if so, the difference signal h is taken. k,i (t) is the kth intrinsic mode component IMF k If the result is positive, proceed to step (2-8); otherwise, take the difference signal h. k,i (t) is the main signal z of the (i+1)th cycle of the kth intrinsic mode component at time t. k,i+1 (t), set counter i = i + 1, and return to step (2-3); (2-8) Determine if counter k is equal to 5. If yes, proceed to step (2-9). Otherwise, take the difference signal h of the (i+1)th cycle of the kth intrinsic mode component at time t. k,i (t) is the main signal z of the (k+1)th intrinsic mode component in the first cycle. k+1,1 (t), set counter k=k+1, counter i=1, and return to step (2-3); (2-9) Select the three lowest frequency intrinsic mode components (IMF3, IMF4, and IMF5) from all the obtained intrinsic mode components, and perform weighted fusion of the three with weights of 0.2, 0.3, and 0.5. Input the fusion result into a low-pass filter and perform downsampling to obtain the environmental vibration signal component F.env ; Preferably, the MLP-Integrator module includes a modal preprocessing layer, a time-frequency feature matrix construction layer, and a multilayer perceptron interaction layer; The first layer of the MLP-Integrator module is the modal preprocessing layer, whose input is the environmental vibration signal component F output from the binary ensemble empirical mode decomposition module. env The layer also outputs the mid-to-high frequency sub-band signal component D2 and the high frequency sub-band signal component D1 from the wavelet decomposition module. This layer then analyzes the environmental vibration signal component F... env The time sequences corresponding to the mid-to-high frequency sub-band signal component D2 and the high frequency sub-band signal component D1 are all subjected to temporal convolution and phase calibration processes to obtain the corresponding time sequences z. env z(t), z1(t), and z2(t); The above processing steps of the modal preprocessing layer include the following sub-steps: (3-1-1) A one-dimensional convolution kernel of size 3 is used to perform convolution operation on the time sequence corresponding to the high-frequency subband signal component D1 to obtain the convolution sequence c1(t), and Hilbert transform is performed on the convolution sequence c1(t) to obtain the feature sequence z1(t). Specifically, the feature sequence z1(t) is equal to: ; in, Let z1(t) represent the Hilbert transform, where the real part of the feature sequence z1(t) is the convolution sequence c1(t), and the imaginary part is the phase shift of the convolution sequence c1(t) after the Hilbert transform. (3-1-2) Obtain the time sequence length N corresponding to the high-frequency subband signal component D1. D And determine the length N. D Is it greater than the environmental vibration signal component F? env The corresponding time sequence length, if so, is in the environmental vibration signal component F. env The corresponding time series sequence is cyclically padded at the end, with the padded length being the difference between the two lengths. The padded content is from F... env The sequence begins with a loop that repeats to obtain the new environmental vibration signal component F. env If the length is N, proceed to step (3-1-3); otherwise, take the length as N. D Environmental vibration signal component F env As a component of the vibration signal in the new environment, F env ', and proceed to step (3-1-3); (3-1-3) Determine whether the length of the time sequence corresponding to the high-frequency subband signal component D1 is greater than the length of the time sequence corresponding to the mid-high frequency subband signal component D2. If so, perform cyclic padding at the end of the time sequence corresponding to the mid-high frequency subband signal component D2. The padding length is the difference between the two lengths. The padding content is a cyclic repetition starting from the beginning of the F2 sequence to obtain a new mid-high frequency subband signal component D2', and proceed to step (3-1-4). Otherwise, take a length of N. D The mid-to-high frequency subband signal component D2 is used as the new mid-to-high frequency subband signal component D2' and proceeds to step (3-1-4). (3-1-4) A one-dimensional convolution kernel of size 3 is used to analyze the vibration signal component F of the new environment. env The corresponding time sequence and the time sequence corresponding to the new mid-to-high frequency sub-band signal component D2 are convolved to obtain convolution sequences c respectively. env c(t) and c2(t), and respectively for the convolution sequence c env Perform Hilbert transforms on c1(t) and c2(t) respectively to obtain the characteristic sequence z. env z(t) and z2(t); ; ;

[0010] Preferably, the second layer of the MLP-Integrator module is a time-frequency feature matrix construction layer, whose input is the feature sequence z1(t) and z2(t) obtained from the first layer. env z1(t) and z2(t), this layer extracts the feature sequences z1(t) and z2(t) respectively. env Extract the corresponding energy value, phase, and frequency from z(t) and z2(t) to construct the feature matrix corresponding to the feature sequence; The above processing steps for the time-frequency feature matrix construction layer include the following sub-steps: (3-2-1) Calculate the characteristic sequences z1(t) and z2(t) respectively. env The instantaneous energy values ​​a1(t) and a2(t) at time t are z1(t) and z2(t). env (t) and a2(t); ; in ; (3-2-2) Calculate the characteristic sequences z1(t) and z2(t) respectively. env The instantaneous phases ϕ1(t) and ϕ2(t) at time t. env (t) and ϕ2(t); ; (3-2-3) For the instantaneous phases ϕ1(t) and ϕ1(t) obtained in step (3-2-1) respectively... env Taking the first derivatives of ϕ1(t) and ϕ2(t) to obtain their instantaneous frequencies f1(t) and f2(t) at time t, respectively. env f(t) and f2(t); ; (3-2-4) Construct the time-feature channel matrix E1∈R corresponding to the high-frequency subband signal component D1. N*d Environmental vibration signal component F env The corresponding time-feature channel matrix E env ∈R N*d And the time-feature channel matrix E2∈R corresponding to the mid-to-high frequency subband signal component D2. N*d Where N is the number of sampling points of the time-feature channel matrix in the time dimension, and d is the number of feature channels of the time-feature channel matrix in the feature channel dimension; (3-2-5) Analyze the time-feature channel matrices E1 and E2 obtained in step (3-2-4) respectively. env Z-score normalization is performed on each feature channel in E2 to obtain the normalized feature matrices E1' and E2 respectively. env 'and E2'.

[0011] ; in μ 1,j and σ 1,j It is the mean and standard deviation of the j-th feature channel of time-feature channel E1, μ env,j and σ env,j It is the time-feature channel E env The mean and standard deviation of the j-th feature channel, μ 2,j and σ 2,j It is the mean and standard deviation of the j-th feature channel of the time-feature channel matrix E2.

[0012] Preferably, the third layer of the MLP-Integrator module is a multilayer perceptron interaction layer, whose input is the normalized feature matrix E1', E2', E3' output from the time-frequency feature matrix construction layer. env 'and E2', this layer on the normalized feature matrix E env E1' and E2' are subjected to matrix concatenation, feature channel mixing, dimension transpose, mode mixing, dimension reversal and feature fusion processing in sequence to obtain the impurity tensor U and the normalized feature matrix E1'. The above processing procedure of the multilayer perceptron interaction layer includes the following sub-steps: (3-3-1) Use the Stack function to concatenate the standardized feature matrix Eenv 'and E2' perform bimodal fusion to obtain the input tensor X∈R N*M*d ; Where M is the number of modes of the input tensor X in the modal dimension; (3-3-2) Extract all sub-tensors from the v-th discrete feature channel of the input tensor X obtained in step (3-3-1) to form a single-feature-channel feature sequence set X. u,v ∈R N*M*1 The feature sequence set is input into the feature channel mixing block for nonlinear transformation, and the nonlinearly transformed feature sequence is residually connected with the input tensor X to obtain the cross-feature channel intermediate tensor K, where u∈[1,M], v∈[1,d]; (3-3-3) Swap the modal dimension of the cross-feature channel intermediate feature K obtained in step (3-3-2) with the feature channel dimension to obtain the intermediate tensor X'∈R. N*d*M ; (3-3-4) Extract all sub-tensors from the u-th mode of the intermediate tensor X' obtained in step (3-3-3) to form the feature sequence set X of the single mode. v,u ∈R N*d*1 The feature sequence set is input into the modal mixing block for nonlinear transformation to obtain the nonlinear transformation result. The nonlinear transformation result is then residually connected with the intermediate tensor X' to obtain the cross-modal intermediate tensor T. (3-3-5) Swap the modal dimension and feature channel dimension of the cross-modal intermediate tensor T obtained in step (3-3-4) again to obtain the intermediate tensor T; (3-3-6) Perform element-wise multiplication on the intermediate tensor T obtained in step (3-3-5) to obtain the impurity tensor U∈R. N *M*d .

[0013] Preferably, the specific structure of the cascaded cross-modal feature decoupling module is as follows: The first layer is the impurity tensor preprocessing layer. Its input is the impurity tensor U output by the MLP-Integrator module. This module performs a global average pooling operation on the impurity tensor U in the modality dimension to obtain the feature tensor. This feature tensor is then input into a linear projection layer for linear transformation to obtain the reference matrix C∈R. N*d And output; The second layer is the feature orthogonal projection decomposition layer. Its inputs are the normalized feature matrix E1' output from the multilayer perceptron interaction layer of the MLP-Integrator module and the reference matrix C output from the impurity tensor preprocessing layer. This layer performs L2 norm normalization on the normalized feature matrix E1' and the reference matrix C respectively to obtain the normalized feature references H1 and H2. c Obtain the normalized feature benchmarks H1 and H c The correlation coefficient matrix between them is then multiplied element-wise with the normalized characteristic benchmark H1 to obtain the correlation coefficient matrix between the normalized characteristic benchmark H1 and the normalized characteristic benchmark H1. c Orthogonal projections along the direction, i.e., common characteristic components H 1,c ; The third layer is the cleaned feature reconstruction and output layer, whose input is the normalized feature reference H1 and common feature components H obtained from the feature orthogonal projection decomposition layer. 1,c This layer removes common feature components H from the normalized feature benchmark H1 based on the learnable weight matrix α. 1,c To obtain the purified target modal features H clean And the target modal feature H clean Input a lightweight multilayer perceptron network to obtain the final prediction result.

[0014] Preferably, the human physiological characteristic vibration detection model is trained through the following steps: (A1) Obtain a multimodal vibration signal dataset consisting of a surface vibration signal dataset p and a global vibration signal dataset q, and divide the multimodal vibration signal dataset into a training set and a test set in a ratio of 7:3; Specifically, vibration datasets involving various typical box-type spaces under multiple working conditions are collected. These include a surface vibration signal dataset p, composed of multiple surface vibration signals from multiple typical box-type spaces collected by a laser vibrometer, and a total vibration signal dataset q, composed of multiple total vibration signals from multiple typical box-type spaces collected by a vibration sensor. Each sample from both datasets corresponds to a binary classification label. Then, samples from dataset p and dataset q are sequentially combined to form a sample of a multimodal vibration signal dataset, thus constituting the multimodal vibration signal dataset. (A2) Initialize the human physiological characteristic vibration detection model to obtain the initialized human physiological characteristic vibration detection model; (A3) For each sample in the training set obtained in step (A1), the portion of the sample corresponding to the surface vibration signal dataset p is input into the wavelet decomposition module of the human physiological characteristic vibration detection model initialized in step (A2) to obtain the low-frequency subband signal component A2 and the mid-to-high frequency subband signal component D2 corresponding to the sample. (A4) For each sample in the training set obtained in step (A1), extract the part of the sample corresponding to the overall vibration signal dataset q to obtain the corresponding overall vibration signal x(t), and input it along with the low-frequency subband signal component A2 obtained in step (A3) into the binary set empirical mode decomposition module in the human physiological characteristic vibration detection model initialized in step (A2) to obtain the environmental vibration signal component corresponding to the sample. (A5) For each sample in the training set obtained in step (A1), the environmental vibration signal component corresponding to the sample obtained in step (5-4), and the low-frequency sub-band signal component A2 and the mid-to-high frequency sub-band signal component D2 corresponding to the sample obtained in step (A3) are respectively input into the MLP-Integrator module in the human physiological characteristic vibration detection model after initialization in step (A2) to obtain the impurity tensor U and the normalized feature matrix E1' corresponding to the sample respectively. (A6) For each sample in the training set obtained in step (A1), input the impurity tensor U and the normalized feature matrix E1' obtained in step (A5) into the cascaded cross-modal feature decoupling module in the human physiological feature vibration detection model initialized in step (A2) to obtain the final prediction probability corresponding to that sample. , where i∈[1, the total number of samples in the training set N]; (A7) For each sample in the training set obtained in step (A1), the final predicted probability corresponding to that sample is obtained in step (A6). Obtain the final predicted probability of the sample obtained in step (A6). With the true label of the sample Binary cross-entropy loss between: ; (A8) For each sample in the training set of step (A1), the binary cross-entropy loss corresponding to the sample obtained in step (5-7) is used to iteratively optimize the human physiological feature vibration detection model that integrates multi-scale features using gradient descent. During the training process, the loss function between the forward propagation output and the real label is used as a guide to update the model weights through backpropagation until the human physiological feature vibration detection model that integrates multi-scale features reaches the preset number of iterations. The optimal parameters of the human physiological feature vibration detection model that integrates multi-scale features are obtained at this time, thus obtaining the initially trained human physiological feature vibration detection model. (A9) Use the test set obtained in step (A1) to test the human physiological characteristic vibration detection model that was initially trained in step (A8) and integrates multi-scale features, so as to obtain the final trained human physiological characteristic vibration detection model.

[0015] According to another aspect of the present invention, a box-type space concealed personnel detection system based on multimodal feature fusion and signal cleansing is provided, comprising the following modules: The first module is used to collect surface vibration signals of the target under test using a laser vibrometer and to collect overall vibration signals of the target under test using a vibration sensor. The second module is used to input the surface vibration signals and overall vibration signals collected by the first module into a pre-trained human physiological characteristic vibration detection model to obtain detection results.

[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: 1. This invention employs a wavelet decomposition module to perform wavelet decomposition on the overall vibration signal of the box, and a binary ensemble empirical mode decomposition module to perform ensemble empirical mode decomposition on the surface vibration signal of the box. Therefore, it can solve the problem of the existing traditional detection technology being greatly affected by external conditions. 2. This invention employs a binary set empirical mode decomposition module, which fuses and integrates the low-frequency signal components of the surface vibration signal obtained by the wavelet decomposition module with the overall vibration signal obtained at the same time in step (2-1) in steps (2-2) to (2-9) through empirical mode decomposition (EMD). This significantly increases the proportion of low-frequency environmental vibration signals in the obtained intrinsic mode component (IMF) results, thereby extracting the environmental vibration signal components mixed in the surface vibration signal. Therefore, by transforming the difficulty into targeted focusing and preliminary feature extraction of environmental vibration, this invention solves the technical problems of difficulty in extracting human physiological vibration information and insufficient feature extraction in existing physical vibration detection methods. 3. This invention employs an MLP-Integrator module, which preprocesses the mid-to-high frequency sub-band signal components obtained by the wavelet decomposition module and the environmental vibration signal components obtained by the binary set empirical mode decomposition module, constructs a standardized feature matrix, and then performs a dual-branch interactive architecture and residual connection to achieve deep information fusion and feature purification across feature channels and modes. This accurately captures and separates the impurity tensor that expresses the physiological information features of the human body contaminated by environmental vibration signals, thus solving the technical problem of signal contamination caused by the signal transmission path in existing laser vibration measurement methods. 4. This invention obtains mid-to-high frequency sub-band signal components and high-frequency sub-band signal components from surface vibration signals through a wavelet decomposition module, and obtains environmental vibration signal components from the overall vibration signal through a binary set empirical mode decomposition module. Because it employs an MLP-Integrator module, it first standardizes the mid-to-high frequency sub-band signal components, high-frequency sub-band signal components, and environmental vibration signal components through a time-frequency feature matrix construction layer to obtain a standardized feature matrix. Then, through a multilayer perceptron interaction layer, it extracts similar features from the standardized feature matrices corresponding to the mid-to-high frequency sub-band signal components and environmental vibration signal components to obtain an impurity tensor. Because it employs a cascaded cross-modal feature decoupling module, it first performs global average pooling of the impurity tensor along the modal dimension through a tensor preprocessing layer to obtain a reference matrix. Then, through a feature orthogonal projection decomposition layer, it performs L2 norm normalization and orthogonal projection on the standardized feature matrix corresponding to the high-frequency sub-band signal components and the reference matrix. Using a learnable weight matrix α, it adaptively extracts redundant components shared with environmental noise from the fused features, ultimately outputting highly pure core features that characterize the target's physiological information. 5. This invention promotes the technical concept of multimodal fusion and provides a new method for human concealment security detection based on human physiological signal detection. Attached Figure Description

[0017] Figure 1 This is a flowchart of the method for detecting people hiding in a box-shaped space based on multimodal feature fusion and signal purification according to the present invention; Figure 2 This is a schematic diagram of the training process of the vibration detection model for human physiological characteristics of the present invention; Figure 3 This is a schematic diagram of the operation of the vibration detection model for human physiological characteristics of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0019] The basic idea of ​​this invention is to address the problem of overlapping physiological information and environmental interference in vibration signals during the detection of people hiding in box-shaped spaces. It collects surface vibration signals using a laser vibrometer and overall vibration signals using a vibration sensor, and constructs a progressive processing framework of "signal decomposition - interference extraction - feature fusion - decoupling and purification": First, using a wavelet decomposition module with a db4 wavelet basis, the surface vibration signal is decomposed in two layers according to the logic of "coarse decomposition followed by fine decomposition," dividing it into a high-frequency subband D1 focusing on key physiological characteristics, a mid-to-high-frequency subband D2 containing pollutant physiological information, and a low-frequency subband A2 focusing on environmental interference, thus solving the frequency overlap problem and providing targeted input; then, using a binary ensemble empirical mode decomposition module, the low-frequency subband A2 is integrated into the overall vibration signal for integrated EMD, enhancing the identification of environmental vibration signals and obtaining the environmental vibration signal components; subsequently, ML... The P-Integrator module first performs temporal convolution, Hilbert transform, and length alignment on the environmental vibration signal components, mid-to-high frequency sub-band signal components, and high frequency sub-band signal components to generate a time-frequency feature matrix containing instantaneous energy, phase, and frequency, and completes Z-score normalization. Then, the normalized time-frequency feature matrix of the environmental vibration signal components and mid-to-high frequency sub-band signal components is input into the dual-branch architecture and residual connection is performed to achieve deep fusion across modalities and feature channels, resulting in an impurity tensor that expresses the physiological information characteristics of the human body contaminated by environmental vibration signals. Finally, the cascaded cross-modal feature decoupling module performs global average pooling and orthogonal projection on the modal dimension of the impurity tensor, and adaptively removes impurity components from the high frequency signal through a learnable weight matrix α, accurately separating and purifying the human physiological information, and then inputting it into a multilayer perceptron to complete the accurate detection of people hiding in the box.

[0020] like Figure 1 As shown, this invention provides a method for detecting people hiding in a box-shaped space based on multimodal feature fusion and signal cleansing, including the following steps: (i) Use a laser vibrometer to collect surface vibration signals of the target to be tested (mainly the natural vibration of the environment and the target to be tested, as well as the weak vibration signals generated by the transmission of physiological signals of the human body to the surface of the target to be tested), and use a vibration sensor to collect the overall vibration signals of the target to be tested (mainly the mechanical vibration signals generated by the natural vibration of the environment and the target to be tested). Specifically, this step involves first calibrating the laser vibrometer and setting up a test scenario within 1 meter of the target to be tested, sampling for more than 3 seconds to reduce data errors in surface vibration detection; then, installing vibration sensors at multiple key locations on the target to be tested using a five-point sampling method (for trucks, this includes the inner wall of the cargo compartment, the chassis, and the engine), ensuring that they cover the main vibration areas of the target to be tested.

[0021] (ii) Input the surface vibration signal and overall vibration signal collected in step (i) into the pre-trained human physiological characteristic vibration detection model to obtain the detection results.

[0022] like Figure 3 As shown, the human physiological characteristic vibration detection model of the present invention comprises four parts connected in sequence: a wavelet decomposition module, a binary set empirical mode decomposition module, an MLP-Integrator module, and a cascaded cross-modal feature decoupling module.

[0023] The specific structure of the wavelet decomposition module is as follows: The first layer is the initial decomposition layer, whose input is the original signal 'a' of surface vibration. This layer first uses the db4 wavelet basis to perform low-pass filtering (coefficients: 0.1629, 0.5449, 0.5449, 0.1629) and 2x downsampling on the input original signal 'a'. Then, the downsampling result is subjected to a linear interpolation to obtain the first-layer approximate component A1. Subsequently, the first-layer approximate component A1 is subjected to high-pass filtering (coefficients: 0.1629, 0.5449, -0.5449, 0.1629) and 2x downsampling on the first-layer approximate component A1. Finally, the downsampling result is subjected to a linear interpolation to obtain the high-frequency subband signal component D1. The second layer is the depth decomposition layer. Its input is the first-layer approximation component A1 output by the first layer. This layer first performs low-pass filtering and 2x downsampling on the first-layer approximation component A1 through the db4 wavelet basis, and then performs cubic spline interpolation on the downsampling result to obtain the low-frequency subband signal component A2. Then, the low-frequency subband signal component A2 is subjected to high-pass filtering and 2x downsampling, and then the downsampling result is subjected to cubic spline interpolation to obtain the mid-high frequency subband signal component D2. The advantage of the wavelet decomposition module lies in its utilization of the higher vibration frequency of physiologically relevant signals compared to environmental vibrations. Through a two-level wavelet decomposition logic—first coarse, then fine—the surface vibration signal is split into three sub-bands with different frequencies but a unified format: a high-frequency sub-band D1 focusing on key physiological features; a mid-to-high-frequency sub-band D2 focusing on components contaminated by environmental vibrations but still containing physiological information; and a low-frequency sub-band A2 focusing on components of environmental vibrations. This approach solves the problem of overlapping frequencies between physiological information and interference signals in the original signal and provides a highly targeted input foundation for subsequent signal processing. The signal length is then recovered using cubic spline interpolation, which better preserves mid-to-low-frequency features than linear interpolation, avoiding the loss of key information in A2 and D2 and ensuring the effectiveness of subsequent integrated processing.

[0024] The input to the binary set empirical mode decomposition module is the low-frequency subband signal component A2 obtained from the wavelet decomposition module, and the overall vibration signal x(t) of the target at time t collected by the laser vibrometer. This module first preprocesses the overall vibration signal x(t) of the target at time t, then fuses the preprocessed result with the low-frequency subband signal component A2. Subsequently, it performs multiple empirical mode decomposition processes on the fused result to obtain multiple intrinsic mode functions (IMFs). Finally, it selects the lowest-frequency IMFs from all IMFs and fuses them into the environmental vibration signal component F. env And output it.

[0025] The above processing procedure of the binary set empirical mode decomposition module includes the following sub-steps: (2-1) The overall vibration signal x(t) of the target under test at time t is subjected to low-frequency filtering using a Butterworth low-pass filter (the cutoff frequency f ranges from 10Hz to 50Hz, and can be adaptively adjusted according to the vibration frequency characteristics of the target under test). The overall vibration signal after low-frequency filtering is then processed using Min-Max normalization to obtain the main signal z of the first cycle of the first intrinsic mode component at time t. 1,1 (t); (2-2) Set counter k=1 and counter i=1; (2-3) Inject the time sequence A2(t) corresponding to the low-frequency subband signal component A2 obtained by the wavelet decomposition module at time t into the main signal z of the i-th cycle of the k-th eigenmode component at time t. k,i In (t), the mixed signal x of the i-th cycle of the k-th intrinsic mode component at time t is obtained. k,i (t); Specifically, the mixed signal x of the k-th intrinsic mode component in the i-th cycle at time t k,i (t) equals: ; Where λ is the amplitude coefficient, and λ∈[0.1, 0.3]; (2-4) Obtain the mixed signal x of the k-th intrinsic mode component obtained in step (2-3) during the i-th cycle. k,i All time sampling points in (t) are obtained, and the local maxima and local minima corresponding to each time sampling point are obtained; (2-5) Based on the local maxima and local minima corresponding to all time sampling points obtained in step (2-4), and using cubic spline interpolation, fit the upper and lower envelopes at time t, and calculate the mean amplitude of the upper and lower envelopes at time t to obtain the mean envelope m of the k-th intrinsic mode component in the i-th cycle at time t. k,i (t); (2-6) For the main signal z of the i-th cycle of the k-th intrinsic mode component at time t k,i (t) and the mean envelope m of the k-th intrinsic mode component in the i-th cycle obtained in step (2-5) at time t k,i (t) Perform subtraction operation on each time sampling point to obtain the difference signal h of the k-th intrinsic mode component in the i-th cycle at time t. k,i (t); (2-7) Determine whether the difference signal h of the k-th intrinsic mode component in the i-th cycle at time t obtained in step (2-6) exists. k,i The difference between the number of local extrema and the number of zero crossings of (t) is less than or equal to 1, and the mean envelope m k,i The amplitude of (t) is less than or equal to the threshold ε. If so, the difference signal h is taken. k,i (t) is the kth intrinsic mode component IMF k If the result is positive, proceed to step (2-8); otherwise, take the difference signal h. k,i (t) is the main signal z of the (i+1)th cycle of the kth intrinsic mode component at time t. k,i+1 (t), set counter i = i + 1, and return to step (2-3); Specifically, the threshold ε ranges from [0.2, 0.3], with a preferred value of 0.25, based on the following: Basis 1: The balance between decomposition accuracy and efficiency. When ε is less than 0.2, although it can further improve the pure oscillation characteristics of the IMF components, it will increase the number of EMD screening cycles, leading to a longer overall mode decomposition time and reduced module real-time performance. When ε is greater than 0.3, although it can reduce the number of cycles, it is easy for the IMF components to retain local trend terms, which will affect the subsequent environmental vibration signal components F. env The extraction error increased; Basis 2: Scene adaptability of vibration signals. The overall vibration signal of the target to be detected is characterized by low / mid-low frequencies dominating. Taking ε to around 0.25 can accurately filter out local envelope shifts caused by interference, ensuring that the decomposed IMF components can effectively characterize the inherent frequency features of environmental vibrations. (2-8) Determine if counter k is equal to 5. If yes, proceed to step (2-9). Otherwise, take the difference signal h of the (i+1)th cycle of the kth intrinsic mode component at time t. k,i(t) is the main signal z of the (k+1)th intrinsic mode component in the first cycle. k+1,1 (t), set counter k=k+1, counter i=1, and return to step (2-3); (2-9) Select the three lowest frequency intrinsic mode components (IMF3, IMF4, and IMF5) from all the obtained intrinsic mode components, and perform weighted fusion of the three with weights of 0.2, 0.3, and 0.5. Input the fusion result into a low-pass filter and perform downsampling (sample one value every other point according to a downsampling factor of 2) to obtain the environmental vibration signal component F. env ; The advantages of the binary ensemble empirical mode decomposition module are twofold. First, it abandons the pursuit of detection difficulty and accuracy in the traditional optimization path, instead focusing on the characteristic that the installation conditions of vibration sensors are too harsh, leading to a large amount of environmental vibration information being mixed into the target vibration information, and extracts the environmental vibration signal. Second, compared with the traditional ensemble empirical mode decomposition using white noise, inputting the low-frequency sub-band signal obtained by wavelet decomposition into the overall vibration signal for empirical mode decomposition can avoid the random interference of traditional white noise, specifically enhance the frequency band focusing ability of the low-frequency part, and significantly improve the identification of low-frequency environmental noise mixed in the surface vibration signal.

[0026] The specific structure of the MLP-Integrator module is as follows: The first layer is the modal preprocessing layer, whose input is the environmental vibration signal component F output by the binary ensemble empirical mode decomposition module. env The layer also outputs the mid-to-high frequency sub-band signal component D2 and the high frequency sub-band signal component D1 from the wavelet decomposition module. This layer then analyzes the environmental vibration signal component F... env The time sequences corresponding to the mid-to-high frequency sub-band signal component D2 and the high frequency sub-band signal component D1 are all subjected to temporal convolution and phase calibration processes to obtain the corresponding time sequences z. env z(t), z1(t), and z2(t); The above processing steps of the modal preprocessing layer include the following sub-steps: (3-1-1) A one-dimensional convolution kernel of size 3 is used to perform convolution operation on the time sequence corresponding to the high-frequency subband signal component D1 to obtain the convolution sequence c1(t), and Hilbert transform is performed on the convolution sequence c1(t) to obtain the feature sequence z1(t). Specifically, the feature sequence z1(t) is equal to: ; in, Let z1(t) represent the Hilbert transform, where the real part of the feature sequence z1(t) is the convolution sequence c1(t), and the imaginary part is the phase shift of the convolution sequence c1(t) after the Hilbert transform. (3-1-2) Obtain the time sequence length N corresponding to the high-frequency subband signal component D1. D And determine the length N. D Is it greater than the environmental vibration signal component F? env The corresponding time sequence length, if so, is in the environmental vibration signal component F. env The corresponding time series sequence is cyclically padded at the end, with the padded length being the difference between the two lengths. The padded content is from F... env The sequence begins with a loop that repeats to obtain the new environmental vibration signal component F. env If the length is N, proceed to step (3-1-3); otherwise, take the length as N. D Environmental vibration signal component F env As a component of the vibration signal in the new environment, F env ', and proceed to step (3-1-3); (3-1-3) Determine whether the length of the time sequence corresponding to the high-frequency subband signal component D1 is greater than the length of the time sequence corresponding to the mid-high frequency subband signal component D2. If so, perform cyclic padding at the end of the time sequence corresponding to the mid-high frequency subband signal component D2. The padding length is the difference between the two lengths. The padding content is a cyclic repetition starting from the beginning of the F2 sequence to obtain a new mid-high frequency subband signal component D2', and proceed to step (3-1-4). Otherwise, take a length of N. D The mid-to-high frequency subband signal component D2 is used as the new mid-to-high frequency subband signal component D2' and proceeds to step (3-1-4). (3-1-4) A one-dimensional convolution kernel of size 3 is used to analyze the vibration signal component F of the new environment. env The corresponding time sequence and the time sequence corresponding to the new mid-to-high frequency sub-band signal component D2 are convolved to obtain convolution sequences c respectively. env c(t) and c2(t), and respectively for the convolution sequence c env Perform Hilbert transforms on c1(t) and c2(t) respectively to obtain the characteristic sequence z. env z(t) and z2(t); ; ; The advantage of the modal preprocessing layer of the MLP-Integrator module is that, based on the length of the high-frequency subband signal component D1, the lengths of the three types of signals are aligned by completion or truncation, avoiding subsequent feature misalignment due to length differences, laying a unified foundation for cross-signal feature processing, realizing the unified feature morphology of the three types of signals, namely environmental vibration signal component, mid-to-high frequency subband signal component D2, and high-frequency subband signal component D1, and initially extracting effective information to obtain a time sequence of uniform length; The second layer is the time-frequency feature matrix construction layer, whose input is the feature sequence z1(t) and z2(t) obtained from the first layer. env z1(t) and z2(t), this layer extracts the feature sequences z1(t) and z2(t) respectively. env Extract the corresponding energy value, phase, and frequency from z1(t) and z2(t) to construct the feature matrix corresponding to the feature sequence.

[0027] The above processing steps for the time-frequency feature matrix construction layer include the following sub-steps: (3-2-1) Calculate the characteristic sequences z1(t) and z2(t) respectively. env The instantaneous energy values ​​a1(t) and a2(t) at time t are z1(t) and z2(t). env (t) and a2(t); ; in ; (3-2-2) Calculate the characteristic sequences z1(t) and z2(t) respectively. env The instantaneous phases ϕ1(t) and ϕ2(t) at time t. env (t) and ϕ2(t); ; (3-2-3) For the instantaneous phases ϕ1(t) and ϕ1(t) obtained in step (3-2-1) respectively... env Taking the first derivatives of ϕ1(t) and ϕ2(t) to obtain their instantaneous frequencies f1(t) and f2(t) at time t, respectively. env f(t) and f2(t); ; (3-2-4) Construct the time-feature channel matrix E1∈R corresponding to the high-frequency subband signal component D1. N*d Environmental vibration signal component F env The corresponding time-feature channel matrix E env ∈R N*d And the time-feature channel matrix E2∈R corresponding to the mid-to-high frequency subband signal component D2. N*d Where N is the number of sampling points of the time-feature channel matrix in the time dimension, and d is the number of feature channels of the time-feature channel matrix in the feature channel dimension; Specifically, the time-feature channel matrices E1 and E env The horizontal dimension of E2 is the time sampling point t, and the vertical dimension consists of three feature channels: energy channel, phase channel, and frequency channel. The time-feature channel matrices E1 and E2... env The energy values ​​of E2 at time t are the instantaneous energy values ​​a1(t) and a2(t) obtained in step (3-2-1) at time t, respectively. env (t) and a2(t), time-frequency energy matrices E1, E env And the phase of E2 at time t are respectively the instantaneous phases ϕ1(t) and ϕ2(t) obtained in step (3-2-2) at time t. env (t) and ϕ2(t), time-frequency energy matrices E1, E env And the frequencies of E2 at time t are the instantaneous frequencies f1(t) and f2(t) obtained in step (3-2-3) at time t, respectively. env f(t) and f2(t).

[0028] (3-2-5) Analyze the time-feature channel matrices E1 and E2 obtained in step (3-2-4) respectively. env Z-score normalization is performed on each feature channel in E2 to obtain the normalized feature matrices E1' and E2 respectively. env 'and E2'.

[0029] ; in μ 1,j and σ 1,j It is the mean and standard deviation of the j-th feature channel of time-feature channel E1, μ env,j and σ env,j It is the time-feature channel E env The mean and standard deviation of the j-th feature channel, μ 2,j and σ 2,j It is the mean and standard deviation of the j-th feature channel of the time-feature channel matrix E2.

[0030] The advantage of the multilayer perceptron interaction layer in the MLP-Integrator module lies in its ability to extract the amplitude of the oscillation represented by the signal at any given time (i.e., the instantaneous energy value representing the strength of the energy) from the time-series sequences corresponding to the three signals, extract the instantaneous phase representing the phase position of the oscillation at any given time from the argument of the time-series sequence, and extract the instantaneous frequency describing the dominant signal component and time-varying characteristics from the time-series sequence. These three components form the time-feature channel matrices of the three components, and Z-score normalization is performed on each of them. This not only visualizes the time-frequency correlation information of the time-domain signal but also eliminates the scale differences between different feature channels. This allows the model to clearly observe the vibration amplitude at different times and frequencies, making the time-frequency features more comparable. This amplifies the environmental vibration features and obtains a normalized feature matrix, providing high-quality structured input for cross-modal deep interaction.

[0031] The third layer is the multilayer perceptron interaction layer, whose input is the normalized feature matrix E1' and E2' output from the time-frequency feature matrix construction layer. env 'and E2', this layer on the normalized feature matrix E env E1' and E2' are subjected to matrix concatenation, feature channel mixing, dimension transpose, mode mixing, dimension reversal and feature fusion processing in sequence to obtain the impurity tensor U and the normalized feature matrix E1'. The above processing procedure of the multilayer perceptron interaction layer includes the following sub-steps: (3-3-1) Use the Stack function to concatenate the standardized feature matrix E env 'and E2' perform bimodal fusion to obtain the input tensor X∈R N*M*d ; Where M is the number of modes of the input tensor X in the modal dimension; (3-3-2) Extract all sub-tensors from the v-th discrete feature channel of the input tensor X obtained in step (3-3-1) to form a single-feature-channel feature sequence set X. u,v ∈R N*M*1 The feature sequence set is input into the feature channel mixing block for nonlinear transformation, and the nonlinearly transformed feature sequence is residually connected with the input tensor X to obtain the cross-feature channel intermediate tensor K, where u∈[1,M], v∈[1,d]; Specifically, the feature channel mixing block includes two cascaded multilayer perceptron (MLP) units. Each MLP unit processes the feature sequence by first using a fully connected layer (with the parameter being a learnable weight matrix γ) to perform a linear transformation on the feature sequence, then activating the linearly transformed feature sequence using the ReLU activation function, and finally standardizing the activation result using a BatchNorm layer to obtain the standardized result.

[0032] (3-3-3) Swap the modal dimension of the cross-feature channel intermediate feature K obtained in step (3-3-2) with the feature channel dimension to obtain the intermediate tensor X'∈R. N*d*M ; (3-3-4) Extract all sub-tensors from the u-th mode of the intermediate tensor X' obtained in step (3-3-3) to form the feature sequence set X of the single mode. v,u ∈R N*d*1 The feature sequence set is input into the modal mixing block for nonlinear transformation to obtain the nonlinear transformation result. The nonlinear transformation result is then residually connected with the intermediate tensor X' to obtain the cross-modal intermediate tensor T. Specifically, the modality mixing block includes two cascaded MLP units. Each MLP unit processes the feature sequence by first using a fully connected layer (whose parameter is a learnable weight matrix θ) to perform a linear transformation on the feature sequence, then using the ReLU activation function to activate the linearly transformed feature sequence, and finally using a BatchNorm layer to standardize the activation result to obtain a standardized result.

[0033] (3-3-5) Swap the modal dimension and feature channel dimension of the cross-modal intermediate tensor T obtained in step (3-3-4) again to obtain the intermediate tensor T; (3-3-6) Perform element-wise multiplication on the intermediate tensor T obtained in step (3-3-5) to obtain the impurity tensor U∈R. N *M*d .

[0034] The advantage of the multilayer perceptron interaction layer in the MLP-Integrator module lies in its ability to concatenate the standardized feature matrices corresponding to the mid-to-high frequency subband signal components and the environmental vibration signal components. It also utilizes a dual-branch architecture to enhance the interaction between feature channel dimensions and modal dimensions, achieving feature channel and modal mixing. Furthermore, residual connections preserve the original features, and dimensional transposition adapts to interaction requirements, enabling deep information fusion and feature purification across feature channels and modalities. This allows the model to effectively transmit and share information from different feature channels and modalities, thereby accurately extracting the impurity tensor that expresses the physiological information features of the human body contaminated by environmental vibration signals, providing core support for subsequent prediction tasks.

[0035] The specific structure of the cascaded cross-modal feature decoupling module is as follows: The first layer is the impurity tensor preprocessing layer. Its input is the impurity tensor U output by the MLP-Integrator module. This module performs a global average pooling operation on the impurity tensor U in the modality dimension to obtain the feature tensor. This feature tensor is then input into a linear projection layer for linear transformation to obtain the reference matrix C∈R. N*d And output; The second layer is the feature orthogonal projection decomposition layer. Its inputs are the normalized feature matrix E1' output from the multilayer perceptron interaction layer of the MLP-Integrator module and the reference matrix C output from the impurity tensor preprocessing layer. This layer performs L2 norm normalization on the normalized feature matrix E1' and the reference matrix C respectively to obtain the normalized feature references H1 and H2. c Obtain the normalized feature benchmarks H1 and H c The correlation coefficient matrix between them is then multiplied element-wise with the normalized characteristic benchmark H1 to obtain the correlation coefficient matrix between the normalized characteristic benchmark H1 and the normalized characteristic benchmark H1. c Orthogonal projections along the direction, i.e., common characteristic components H 1,c ; The third layer is the cleaned feature reconstruction and output layer, whose input is the normalized feature reference H1 and common feature components H obtained from the feature orthogonal projection decomposition layer. 1,c This layer removes common feature components H from the normalized feature reference H1 based on a learnable weight matrix α (where each column of the weight matrix α corresponds to the elimination coefficient of a feature channel of the normalized feature reference H1, and the initial value of each elimination coefficient is 0.2). 1,c To obtain the purified target modal features H clean And the target modal feature H clean Input a lightweight multilayer perceptron network to obtain the final prediction result.

[0036] Specifically, a lightweight multilayer perceptron network is used to describe the target modal features H.clean The following operations are performed: First, the target modality feature is linearly transformed through a fully connected layer. Then, the target modality feature after linear transformation is activated by the Sigmoid activation function to generate a binary classification probability as the final prediction result and output it.

[0037] The advantage of the cascaded cross-modal feature decoupling module lies in its ability to achieve precise and controllable feature stripping by constructing a common feature benchmark and introducing adaptive, learnable weights. This avoids the drawback of traditional methods that tend to over-remove effective information when removing noise. While maximally suppressing environmental interference, it ensures the complete preservation of key physiological features, thoroughly separating redundant components shared with environmental noise from the extracted mixed features. This results in highly pure core features that characterize the target's physiological information, ultimately directly improving the accuracy and robustness of subsequent prediction models.

[0038] like Figure 2 As shown, the human physiological characteristic vibration detection model of the present invention is obtained through the following steps: (A1) Obtain a multimodal vibration signal dataset consisting of a surface vibration signal dataset p and a global vibration signal dataset q, and divide the multimodal vibration signal dataset into a training set and a test set in a ratio of 7:3; Specifically, vibration datasets are collected from various typical box-type spaces (such as standard containers, vans, and cargo containers) under multiple operating conditions. These include a surface vibration signal dataset p, composed of multiple surface vibration signals from multiple typical box-type spaces collected by a laser vibrometer, and a comprehensive vibration signal dataset q, composed of multiple overall vibration signals from multiple typical box-type spaces collected by vibration sensors. Each sample from both datasets corresponds to a binary classification label (1 represents a person hiding inside the container, 0 represents an empty container or one containing only cargo). Samples from dataset p and dataset q are then sequentially combined to form a multimodal vibration signal dataset, ultimately constituting the multimodal vibration signal dataset.

[0039] The advantage of this step is that by integrating multi-source data, a large-scale, high-quality training foundation covering a wide range of real-world scenarios (different vehicle types, road conditions, speeds, and environmental noise) is constructed. The strict synchronization and alignment of the dual-modal signals ensures temporal coherence, making it possible for the model to learn the deep physical correlations and differences between the "surface" and "whole" of vibration signals, which is a fundamental prerequisite for the subsequent multimodal fusion algorithm to take effect.

[0040] (A2) Initialize the human physiological characteristic vibration detection model to obtain the initialized human physiological characteristic vibration detection model; Specifically, this step involves using Xavier to uniformly initialize the learnable weight matrix γ in the feature channel mixing block of the MLP-Integrator module, the learnable weight matrix θ in the modal mixing block, and the learnable weight matrix α in the third layer of the cascaded cross-modal feature decoupling module. The training hyperparameters are configured as follows: batch size is set to 64, initial learning rate is set to 1×10⁻³, the optimizer is Adam (β1=0.9, β2=0.999), and the weight decay coefficient is set to 1×10⁻³. 4 L2 regularization was applied. The total number of training epochs was set to 200, and the learning rate was set to decay to 0.1 times its original value at 60% and 85% of the total epochs.

[0041] The advantage of this step is that Xavier initialization ensures stable input-output variance in each network layer during the early training phase, which facilitates effective gradient propagation in deeper modules (such as wavelet decomposition, EEMD, MLP-Integrator, and cascaded structures with feature decoupling), accelerating convergence. The Adam optimizer, combining warm-up and step-descent learning rate strategies, enables stable exploration in the early training phase and fine-grained convergence in the later phase. The introduction of weight decay explicitly penalizes the norms of the weight matrices γ, θ, etc., effectively constraining model complexity, improving its generalization ability, and preventing overfitting in complex multimodal signal fitting.

[0042] (A3) For each sample in the training set obtained in step (A1), the portion of the sample corresponding to the surface vibration signal dataset p is input into the wavelet decomposition module of the human physiological characteristic vibration detection model initialized in step (A2) to obtain the low-frequency subband signal component A2 and the mid-to-high frequency subband signal component D2 corresponding to the sample. (A4) For each sample in the training set obtained in step (A1), extract the part of the sample corresponding to the overall vibration signal dataset q to obtain the corresponding overall vibration signal x(t), and input it along with the low-frequency subband signal component A2 obtained in step (A3) into the binary set empirical mode decomposition module in the human physiological characteristic vibration detection model initialized in step (A2) to obtain the environmental vibration signal component corresponding to the sample. (A5) For each sample in the training set obtained in step (A1), the environmental vibration signal component corresponding to the sample obtained in step (5-4), and the low-frequency sub-band signal component A2 and the mid-to-high frequency sub-band signal component D2 corresponding to the sample obtained in step (A3) are respectively input into the MLP-Integrator module in the human physiological characteristic vibration detection model after initialization in step (A2) to obtain the impurity tensor U and the normalized feature matrix E1' corresponding to the sample respectively. (A6) For each sample in the training set obtained in step (A1), input the impurity tensor U and the normalized feature matrix E1' obtained in step (A5) into the cascaded cross-modal feature decoupling module in the human physiological feature vibration detection model initialized in step (A2) to obtain the final prediction probability corresponding to that sample. , where i∈[1, the total number of samples in the training set N]; (A7) For each sample in the training set obtained in step (A1), the final predicted probability corresponding to that sample is obtained in step (A6). Obtain the final predicted probability of the sample obtained in step (A6). With the true label of the sample Binary cross-entropy loss between: ; The advantage of this step is that the forward propagation fully realizes the end-to-end computation graph from the original vibration signal to the determination of the presence of personnel, verifying the coherence and correctness of each module of the system (including the core MLP-Integrator). The loss calculation provides a clear optimization objective for the entire model, driving all parameters, including γ, θ, and α, to adjust collaboratively in the direction of minimizing the classification error.

[0043] (A8) For each sample in the training set of step (A1), the binary cross-entropy loss corresponding to the sample obtained in step (5-7) is used to iteratively optimize the human physiological feature vibration detection model that integrates multi-scale features using gradient descent. During the training process, the loss function between the forward propagation output and the real label is used as a guide to update the model weights through backpropagation until the human physiological feature vibration detection model that integrates multi-scale features reaches the preset number of iterations (100 times in this invention). The optimal parameters of the human physiological feature vibration detection model that integrates multi-scale features are obtained at this time, thereby obtaining the initially trained human physiological feature vibration detection model. (A9) Use the test set obtained in step (A1) to evaluate and screen the performance of the human physiological feature vibration detection model that integrates multi-scale features initially trained in step (A8) (monitor the generalization ability of the model through multiple iterative tests on unseen data) until the detection accuracy reaches the optimal level, thereby obtaining the final trained human physiological feature vibration detection model.

[0044] This invention introduces multimodal collaboration, wavelet multiscale decomposition, intelligent feature fusion and purification mechanisms, and combines the advantages of laser vibration measurement in acquiring relatively good human physiological information with the characteristic that vibration sensing, although not effective in detecting human physiological information, can simultaneously acquire rich environmental vibration information and human physiological information affected by environmental pollution, thus providing a new method for detecting personnel concealment.

[0045] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing, characterized in that, Includes the following steps: Step 1: Use a laser vibrometer to collect the surface vibration signal of the target to be tested, and use a vibration sensor to collect the overall vibration signal of the target to be tested; Step 2: Input the surface vibration signal and overall vibration signal collected in Step 1 into the pre-trained human physiological characteristic vibration detection model to obtain the detection results.

2. The method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing according to claim 1, characterized in that, The surface vibration signal of the target to be detected includes the natural vibration of the environment and the target itself, as well as the weak vibration signal generated by the transmission of physiological signals from the human body to the surface of the target. The overall vibration signal of the target under test is the mechanical vibration signal generated by the natural vibration of the environment and the target under test; Step one is as follows: First, calibrate the laser vibrometer and set up a test scene within 1 meter of the target to be tested and sample for more than 3 seconds to collect the surface vibration signal of the target to be tested; then, install the vibration sensor at multiple key positions of the target to be tested according to the five-point sampling method, and ensure that it covers the main vibration area of ​​the target to be tested, in order to collect the overall vibration signal of the target to be tested.

3. The method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing according to claim 1 or 2, characterized in that, The human physiological characteristic vibration detection model consists of four sequentially connected parts: a wavelet decomposition module, a binary set empirical mode decomposition module, an MLP-Integrator module, and a cascaded cross-modal feature decoupling module. The specific structure of the wavelet decomposition module is as follows: The first layer is the initial decomposition layer, whose input is the original signal a of surface vibration. This layer first uses the db4 wavelet basis to perform low-pass filtering and 2x downsampling on the input original signal a. Then, the downsampling result is subjected to linear interpolation to obtain the first-layer approximate component A1. Subsequently, the first-layer approximate component A1 is subjected to high-pass filtering and 2x downsampling. Finally, the downsampling result is subjected to linear interpolation to obtain the high-frequency subband signal component D1. The second layer is the depth decomposition layer. Its input is the first-layer approximation component A1 output by the first layer. This layer first performs low-pass filtering and 2x downsampling on the first-layer approximation component A1 through the db4 wavelet basis, and then performs cubic spline interpolation on the downsampling result to obtain the low-frequency subband signal component A2. Then, the low-frequency subband signal component A2 is subjected to high-pass filtering and 2x downsampling, and then the downsampling result is subjected to cubic spline interpolation to obtain the mid-high frequency subband signal component D2. The input to the binary set empirical mode decomposition module is the low-frequency subband signal component A2 obtained from the wavelet decomposition module, and the overall vibration signal x(t) of the target at time t collected by the laser vibrometer. This module first preprocesses the overall vibration signal x(t) of the target at time t, then fuses the preprocessed result with the low-frequency subband signal component A2. Subsequently, it performs multiple empirical mode decomposition processes on the fused result to obtain multiple intrinsic mode components (IMFs). Finally, it selects the lowest-frequency IMFs from all IMFs and fuses them into the environmental vibration signal component F. env And output it.

4. The method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing according to any one of claims 1 to 3, characterized in that, The processing steps of the binary set empirical mode decomposition module include the following sub-steps: (2-1) The overall vibration signal x(t) of the target under test at time t is subjected to low-frequency filtering using a Butterworth low-pass filter, and the overall vibration signal after low-frequency filtering is processed by Min-Max normalization to obtain the main signal z of the first cycle of the first intrinsic mode component at time t. 1,1 (t); (2-2) Set counter k=1 and counter i=1; (2-3) Inject the time sequence A2(t) corresponding to the low-frequency subband signal component A2 obtained by the wavelet decomposition module at time t into the main signal z of the i-th cycle of the k-th eigenmode component at time t. k,i In (t), the mixed signal x of the k-th intrinsic mode component in the i-th cycle at time t is obtained. k,i (t): ; Where λ is the amplitude coefficient, and λ∈[0.1, 0.3]; (2-4) Obtain the mixed signal x of the k-th intrinsic mode component obtained in step (2-3) during the i-th cycle. k,i All time sampling points in (t) are obtained, and the local maxima and local minima corresponding to each time sampling point are obtained; (2-5) Based on the local maxima and local minima corresponding to all time sampling points obtained in step (2-4), cubic spline interpolation is used for fitting to obtain the upper and lower envelopes at time t. The mean amplitude of the upper and lower envelopes at time t is calculated to obtain the mean envelope m of the i-th cycle of the k-th intrinsic mode component at time t. k,i (t); (2-6) For the main signal z of the i-th cycle of the k-th intrinsic mode component at time t k,i (t) and the mean envelope m of the k-th intrinsic mode component in the i-th cycle obtained in step (2-5) at time t k,i (t) Perform subtraction operation on each time sampling point to obtain the difference signal h of the k-th intrinsic mode component in the i-th cycle at time t. k,i (t); (2-7) Determine whether the difference signal h of the k-th intrinsic mode component in the i-th cycle at time t obtained in step (2-6) exists. k,i The difference between the number of local extrema and the number of zero crossings of (t) is less than or equal to 1, and the mean envelope m k,i The amplitude of (t) is less than or equal to the threshold ε; if so, the difference signal h is taken. k,i (t) is the kth intrinsic mode component IMF k If the result is positive, proceed to step (2-8); otherwise, take the difference signal h. k,i (t) is the main signal z of the (i+1)th cycle of the kth intrinsic mode component at time t. k,i+1 (t), set counter i = i + 1, and return to step (2-3); (2-8) Determine if counter k is equal to 5. If yes, proceed to step (2-9). Otherwise, take the difference signal h of the (i+1)th cycle of the kth intrinsic mode component at time t. k,i (t) is the main signal z of the (k+1)th intrinsic mode component in the first cycle. k+1,1 (t), set counter k=k+1, counter i=1, and return to step (2-3); (2-9) Select the three lowest frequency intrinsic mode components (IMF3, IMF4, and IMF5) from all the obtained intrinsic mode components, and perform weighted fusion of the three with weights of 0.2, 0.3, and 0.

5. Input the fusion result into a low-pass filter and perform downsampling to obtain the environmental vibration signal component F. env。 5. The method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing according to claim 4, characterized in that, The MLP-Integrator module includes a modal preprocessing layer, a time-frequency feature matrix construction layer, and a multilayer perceptron interaction layer; The first layer of the MLP-Integrator module is the modal preprocessing layer, whose input is the environmental vibration signal component F output from the binary ensemble empirical mode decomposition module. env The layer also outputs the mid-to-high frequency sub-band signal component D2 and the high frequency sub-band signal component D1 from the wavelet decomposition module. This layer then analyzes the environmental vibration signal component F... env The time sequences corresponding to the mid-to-high frequency subband signal component D2 and the high frequency subband signal component D1 are all subjected to temporal convolution and phase calibration processes to obtain the corresponding time sequences z. env z(t), z1(t), and z2(t); The above processing steps of the modal preprocessing layer include the following sub-steps: (3-1-1) A one-dimensional convolution kernel of size 3 is used to perform convolution operation on the time sequence corresponding to the high-frequency subband signal component D1 to obtain the convolution sequence c1(t), and Hilbert transform is performed on the convolution sequence c1(t) to obtain the feature sequence z1(t). Specifically, the feature sequence z1(t) is equal to: ; in, Let z1(t) represent the Hilbert transform, where the real part of the feature sequence z1(t) is the convolution sequence c1(t), and the imaginary part is the phase shift of the convolution sequence c1(t) after the Hilbert transform. (3-1-2) Obtain the time sequence length N corresponding to the high-frequency subband signal component D1. D And determine the length N. D Is it greater than the environmental vibration signal component F? env The corresponding time sequence length, if so, is in the environmental vibration signal component F. env The corresponding time series sequence is cyclically padded at the end, with the padded length being the difference between the two lengths. The padded content is from F... env The sequence begins with a loop that repeats to obtain the new environmental vibration signal component F. env If the length is N, proceed to step (3-1-3); otherwise, take the length as N. D Environmental vibration signal component F env As a component of the vibration signal in the new environment, F env ', and proceed to step (3-1-3); (3-1-3) Determine whether the length of the time sequence corresponding to the high-frequency subband signal component D1 is greater than the length of the time sequence corresponding to the mid-high frequency subband signal component D2. If so, perform cyclic padding at the end of the time sequence corresponding to the mid-high frequency subband signal component D2. The padding length is the difference between the two lengths. The padding content is a cyclic repetition starting from the beginning of the F2 sequence to obtain a new mid-high frequency subband signal component D2', and proceed to step (3-1-4). Otherwise, take a length of N. D The mid-to-high frequency subband signal component D2 is used as the new mid-to-high frequency subband signal component D2' and proceeds to step (3-1-4). (3-1-4) A one-dimensional convolution kernel of size 3 is used to analyze the vibration signal component F of the new environment. env The corresponding time sequence and the time sequence corresponding to the new mid-to-high frequency sub-band signal component D2 are convolved to obtain convolution sequences c respectively. env c(t) and c2(t), and respectively for the convolution sequence c env Perform Hilbert transforms on c1(t) and c2(t) respectively to obtain the characteristic sequence z. env z(t) and z2(t); ; 。 6. The method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing according to claim 5, characterized in that, The second layer of the MLP-Integrator module is the time-frequency feature matrix construction layer, whose input is the feature sequence z1(t) and z2(t) obtained from the first layer. env z1(t) and z2(t), this layer extracts the feature sequences z1(t) and z2(t) respectively. env Extract the corresponding energy value, phase, and frequency from z(t) and z2(t) to construct the feature matrix corresponding to the feature sequence; The above processing steps for the time-frequency feature matrix construction layer include the following sub-steps: (3-2-1) Calculate the characteristic sequences z1(t) and z2(t) respectively. env The instantaneous energy values ​​a1(t) and a2(t) at time t are z1(t) and z2(t). env (t) and a2(t); ; in ; (3-2-2) Calculate the characteristic sequences z1(t) and z2(t) respectively. env The instantaneous phases ϕ1(t) and ϕ2(t) at time t. env (t) and ϕ2(t); ; (3-2-3) For the instantaneous phases ϕ1(t) and ϕ1(t) obtained in step (3-2-1) respectively... env Taking the first derivatives of ϕ(t) and ϕ2(t) to obtain their instantaneous frequencies f1(t) and f2(t) at time t, respectively. env f(t) and f2(t); ; (3-2-4) Construct the time-feature channel matrix E1∈R corresponding to the high-frequency subband signal component D1. N*d Environmental vibration signal component F env The corresponding time-feature channel matrix E env ∈R N*d And the time-feature channel matrix E2∈R corresponding to the mid-to-high frequency subband signal component D2. N*d Where N is the number of sampling points of the time-feature channel matrix in the time dimension, and d is the number of feature channels of the time-feature channel matrix in the feature channel dimension; (3-2-5) Analyze the time-feature channel matrices E1 and E2 obtained in step (3-2-4) respectively. env Z-score normalization is performed on each feature channel in E2 to obtain the normalized feature matrices E1' and E2 respectively. env 'and E2'; ; in μ 1,j and σ 1,j It is the mean and standard deviation of the j-th feature channel of time-feature channel E1, μ env,j and σ env,j It is the time-feature channel E env The mean and standard deviation of the j-th feature channel, μ 2,j and σ 2,j It is the mean and standard deviation of the j-th feature channel of the time-feature channel matrix E2.

7. The method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing according to claim 6, characterized in that, The third layer of the MLP-Integrator module is a multilayer perceptron interaction layer, whose input is the normalized feature matrix E1', E2', E3' output from the time-frequency feature matrix construction layer. env 'and E2', this layer on the normalized feature matrix E env E1' and E2' are subjected to matrix concatenation, feature channel mixing, dimension transpose, mode mixing, dimension reversal and feature fusion processing in sequence to obtain the impurity tensor U and the normalized feature matrix E1'. The above processing procedure of the multilayer perceptron interaction layer includes the following sub-steps: (3-3-1) Use the Stack function to concatenate the standardized feature matrix E env 'and E2' perform bimodal fusion to obtain the input tensor X∈R N*M*d ; ; Where M is the number of modes of the input tensor X in the modal dimension; (3-3-2) Extract all sub-tensors from the v-th discrete feature channel of the input tensor X obtained in step (3-3-1) to form a single-feature-channel feature sequence set X. u,v ∈R N*M*1 The feature sequence set is input into the feature channel mixing block for nonlinear transformation, and the nonlinearly transformed feature sequence is residually connected with the input tensor X to obtain the cross-feature channel intermediate tensor K, where u∈[1,M], v∈[1,d]; (3-3-3) Swap the modal dimension of the cross-feature channel intermediate feature K obtained in step (3-3-2) with the feature channel dimension to obtain the intermediate tensor X'∈R. N*d*M ; (3-3-4) Extract all sub-tensors from the u-th mode of the intermediate tensor X' obtained in step (3-3-3) to form the feature sequence set X of the single mode. v,u ∈R N*d*1 The feature sequence set is input into the modal mixing block for nonlinear transformation to obtain the nonlinear transformation result. The nonlinear transformation result is then residually connected with the intermediate tensor X' to obtain the cross-modal intermediate tensor T. (3-3-5) Swap the modal dimension and feature channel dimension of the cross-modal intermediate tensor T obtained in step (3-3-4) again to obtain the intermediate tensor T; (3-3-6) Perform element-wise multiplication on the intermediate tensor T obtained in step (3-3-5) to obtain the impurity tensor U∈R. N*M*d .

8. The method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing according to claim 7, characterized in that, The specific structure of the cascaded cross-modal feature decoupling module is as follows: The first layer is the impurity tensor preprocessing layer. Its input is the impurity tensor U output by the MLP-Integrator module. This module performs a global average pooling operation on the impurity tensor U in the modality dimension to obtain the feature tensor. This feature tensor is then input into a linear projection layer for linear transformation to obtain the reference matrix C∈R. N*d And output; The second layer is the feature orthogonal projection decomposition layer. Its inputs are the normalized feature matrix E1' output from the multilayer perceptron interaction layer of the MLP-Integrator module and the reference matrix C output from the impurity tensor preprocessing layer. This layer performs L2 norm normalization on the normalized feature matrix E1' and the reference matrix C respectively to obtain the normalized feature references H1 and H2. c Obtain the normalized feature benchmarks H1 and H c The correlation coefficient matrix between them is then multiplied element-wise with the normalized characteristic benchmark H1 to obtain the correlation coefficient matrix between the normalized characteristic benchmark H1 and the normalized characteristic benchmark H1. c Orthogonal projections along the direction, i.e., common characteristic components H 1,c ; The third layer is the cleaned feature reconstruction and output layer, whose input is the normalized feature reference H1 and common feature components H obtained from the feature orthogonal projection decomposition layer. 1,c This layer removes common feature components H from the normalized feature benchmark H1 based on the learnable weight matrix α. 1,c To obtain the purified target modal features H clean And the target modal feature H clean Input a lightweight multilayer perceptron network to obtain the final prediction result.

9. The method for detecting concealed persons in a box-shaped space based on multimodal feature fusion and signal cleansing according to claim 8, characterized in that, The human physiological characteristic vibration detection model is trained through the following steps: (A1) Obtain a multimodal vibration signal dataset consisting of a surface vibration signal dataset p and a global vibration signal dataset q, and divide the multimodal vibration signal dataset into a training set and a test set in a ratio of 7:3; Specifically, vibration datasets involving various typical box-type spaces under multiple working conditions are collected. These include a surface vibration signal dataset p, composed of multiple surface vibration signals from multiple typical box-type spaces collected by a laser vibrometer, and a total vibration signal dataset q, composed of multiple total vibration signals from multiple typical box-type spaces collected by a vibration sensor. Each sample from both datasets corresponds to a binary classification label. Then, samples from dataset p and dataset q are sequentially combined to form a sample of a multimodal vibration signal dataset, thus constituting the multimodal vibration signal dataset. (A2) Initialize the human physiological characteristic vibration detection model to obtain the initialized human physiological characteristic vibration detection model; (A3) For each sample in the training set obtained in step (A1), the portion of the sample corresponding to the surface vibration signal dataset p is input into the wavelet decomposition module of the human physiological characteristic vibration detection model initialized in step (A2) to obtain the low-frequency subband signal component A2 and the mid-to-high frequency subband signal component D2 corresponding to the sample. (A4) For each sample in the training set obtained in step (A1), extract the part of the sample corresponding to the overall vibration signal dataset q to obtain the corresponding overall vibration signal x(t), and input it along with the low-frequency subband signal component A2 obtained in step (A3) into the binary set empirical mode decomposition module in the human physiological characteristic vibration detection model initialized in step (A2) to obtain the environmental vibration signal component corresponding to the sample. (A5) For each sample in the training set obtained in step (A1), the environmental vibration signal component corresponding to the sample obtained in step (5-4), and the low-frequency sub-band signal component A2 and the mid-to-high frequency sub-band signal component D2 corresponding to the sample obtained in step (A3) are respectively input into the MLP-Integrator module in the human physiological characteristic vibration detection model after initialization in step (A2) to obtain the impurity tensor U and the normalized feature matrix E1' corresponding to the sample respectively. (A6) For each sample in the training set obtained in step (A1), input the impurity tensor U and the normalized feature matrix E1' obtained in step (A5) into the cascaded cross-modal feature decoupling module in the human physiological characteristic vibration detection model initialized in step (A2) to obtain the final prediction probability corresponding to that sample. , where i∈[1, the total number of samples in the training set N]; (A7) For each sample in the training set obtained in step (A1), the final predicted probability corresponding to that sample is obtained in step (A6). Obtain the final predicted probability of the sample obtained in step (A6). With the true label of the sample Binary cross-entropy loss between: ; (A8) For each sample in the training set of step (A1), the binary cross-entropy loss corresponding to the sample obtained in step (5-7) is used to iteratively optimize the human physiological feature vibration detection model that integrates multi-scale features using gradient descent. During the training process, the loss function between the forward propagation output and the real label is used as a guide to update the model weights through backpropagation until the human physiological feature vibration detection model that integrates multi-scale features reaches the preset number of iterations. The optimal parameters of the human physiological feature vibration detection model that integrates multi-scale features are obtained at this time, thus obtaining the initially trained human physiological feature vibration detection model. (A9) Use the test set obtained in step (A1) to test the human physiological characteristic vibration detection model that was initially trained in step (A8) and integrates multi-scale features, so as to obtain the final trained human physiological characteristic vibration detection model.

10. A system for detecting concealed personnel in a box-type space based on multimodal feature fusion and signal cleansing, characterized in that, Includes the following modules: The first module is used to collect surface vibration signals of the target under test using a laser vibrometer and to collect overall vibration signals of the target under test using a vibration sensor. The second module is used to input the surface vibration signals and overall vibration signals collected by the first module into a pre-trained human physiological characteristic vibration detection model to obtain detection results.