A micro-motion detection method for ground moving targets in perimeter security

By integrating the architecture of stacked convolutional neural networks and long short-term memory modules, combined with preprocessing and attention mechanisms, the computational load and robustness issues of deep learning models in perimeter security ground moving target detection are solved, achieving efficient anomaly recognition and classification.

CN120178310BActive Publication Date: 2025-10-03JILIN UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510662800.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-10-03
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

Existing deep learning models overly rely on the stacking of network depth in perimeter security ground motion target detection, which leads to increased computational load and decreased robustness, making it difficult to effectively identify seismic anomalies.

Method used

It adopts a fusion architecture of stacked convolutional neural networks and long short-term memory modules, combined with the mask mechanism and spatial attention mechanism of the preprocessing module to achieve dynamic capture of multi-dimensional features and anomaly identification.

Benefits of technology

It achieves efficient recognition of ground moving targets with an accuracy rate of 97%, an improvement of 36% compared to traditional models, and optimizes computing speed and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178310B_ABST
    Figure CN120178310B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of intelligent sensing technology and relates to a method for detecting micro-motion on ground moving targets for perimeter security. The method includes collecting seismic data of moving targets; preprocessing the data to construct a dataset; constructing a joint spatiotemporal feature extraction model, using a CNN to capture the spatiotemporal features of short-term models, and identifying different features in the data, thereby achieving intelligent identification of suspicious targets; and continuously capturing physical quantities related to time series changes through an LSTM. A temporal pattern attention mechanism is introduced into the joint spatiotemporal feature extraction model to dynamically adjust feature weights, allowing the model to focus more on key times and signals and suppress irrelevant noise. The acquired model is then trained and tested to identify abnormal shallow surface vibration data and automatically extract and classify it. The detection method of the present invention can achieve a recognition accuracy of 97% in five ground moving target classification tasks, which is a 36% improvement in processing speed compared to traditional models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent sensing technology, and specifically relates to a ground vibration signal sensing method, and more particularly to a micro-motion detection method for ground moving targets for perimeter security. Background Art

[0002] The security of certain sensitive areas, such as perimeters, within cities, and around them, is paramount. Efficient identification of moving targets is a key task in a variety of applications, from regional security to environmental monitoring. Seismic sensing technology is a highly sensitive and robust passive sensing technology that can effectively detect moving targets. Node seismometers, as terminal sensing devices for intelligent environments, are also powerful tools for maintaining the security of sensitive areas. They can be completely buried and tightly coupled to the soil, making them weatherproof. Regarding the recognition of ground vibration signals, machine learning remains the core classification technology framework in the field of target identification. Its operating mechanism is as follows: when the classifier output is judged to be environmental noise or interference signals, the sensing system automatically filters out the invalid information. Upon detecting an abnormal signal signature that matches the warning target, the system immediately triggers a multi-level response mechanism, simultaneously outputting classification decisions and warning instructions.

[0003] With the evolution of intelligent algorithm systems, this field has achieved a paradigm shift from traditional shallow models to deep neural networks. Although traditional methods have the advantage of low computational complexity, they have inherent defects in the analysis of nonlinear features of time-varying vibration signals, which directly restricts the improvement of their recognition accuracy. In contrast, deep learning methods break through the limitations of manual feature engineering through end-to-end feature learning mechanisms and can autonomously explore the multi-level representation of vibration signals in the time and frequency domains. However, existing deep models generally have architectural design flaws that over-rely on the stacking of network depth, which not only significantly increases the computational load but also weakens the model's robustness in noisy environments. Therefore, constructing a deep network architecture that integrates signal mechanism cognition and lightweight design has become a key breakthrough in improving the performance of target classification systems. Summary of the Invention

[0004] The purpose of the present invention is to provide a micro-motion detection method for ground moving targets for perimeter security. It adopts a fusion architecture of stacked convolutional neural networks and long short-term memory modules, integrates the mask mechanism and spatial attention mechanism of the preprocessing module, and realizes dynamic capture of multi-dimensional features to solve the problems of automatic extraction of earthquake characteristics and intelligent classification of moving targets, and efficiently identify seismic anomalies.

[0005] The purpose of the present invention is achieved through the following technical solutions:

[0006] A method for detecting micro-motion of ground moving targets for perimeter security comprises the following steps:

[0007] A. Collection of moving target seismic monitoring data:

[0008] Collect shallow surface vibration signals induced by ground moving targets, capture seismic waves generated by moving targets, and output original time domain vibration signals;

[0009] B. Preprocess the acquired moving target earthquake monitoring data and construct a data set:

[0010] The acquired raw time-domain vibration signals are preprocessed, with continuous signals segmented, standardized, and initially extracted, then divided into sliding windows of fixed length to ensure that the final learning results cover the entire dataset to meet real-time detection requirements. 80% of the dataset is used as a training set, and the rest as a test set.

[0011] C. Build a joint spatiotemporal feature extraction model, using a fusion architecture of stacked CNN and LSTM to achieve spatiotemporal feature fusion and feature enhancement. CNN is used to capture the spatiotemporal features of short-term models and identify different features in the data, thereby completing the intelligent identification of suspicious targets. LSTM is also used to continuously capture the physical quantities related to time series changes.

[0012] D. Introducing the temporal pattern attention mechanism into the joint spatiotemporal feature extraction model, it selects relevant time series, remembers important features across the entire time domain, and dynamically adjusts feature weights, allowing the model to focus more on key times and key signals while suppressing irrelevant noise.

[0013] E. Train and test the model obtained in step D to identify shallow surface abnormal vibration data and automatically extract and classify it.

[0014] Furthermore, in step A, shallow surface vibration signals induced by moving targets are collected. Specifically, shallow surface vibration signals triggered by five common types of moving targets, namely human activities, wheeled vehicle driving, tracked vehicle driving, aircraft flight, and natural noise are collected.

[0015] Furthermore, step B is specifically as follows: after obtaining the time series data, first construct unbalanced category samples of five categories of moving targets; first, fill the data to unify the unbalanced categories into a standard size matrix to obtain a standard sample, and use a Boolean mask matrix to cover the top of the unbalanced data set to indicate the filling position; then apply a sliding window on the time series data, and generate a subsequence for each window; for each window, perform feature extraction, label extraction, normalization and sliding window in sequence; finally, convert the time series data into a set of input features and corresponding labels; thereby, completing the construction of the data set.

[0016] Furthermore, step C specifically involves first using four convolutional layers to locate and extract spatial features, and connecting a maximum pooling layer after each convolutional layer to perform downsampling operations; then, adding an LSTM-based temporal attention module after the last CNN layer to accurately identify and retain important dynamic behaviors of time series of arbitrary lengths and hide redundant information; finally, after the expansion layer, a fully connected layer and a softmax classifier are accessed in sequence to classify events according to the learned features.

[0017] Furthermore, CNN consists of alternating stacks of convolutional layers and pooling layers. There is an activation function between each convolutional layer and pooling layer to accelerate the convergence of the model. The function σ is the ReLU activation function. The pooling layer is used for downsampling operations and can be expressed as:

[0018] ;

[0019] ;

[0020] in, It is The i-th feature in the layer, It is The j-th feature in the layer, represents the kernel connected to the function, is the deviation of this function;

[0021] After passing the data output from the convolution module to the long short-term memory module, the decision on what to keep and what to discard is done with the help of the activation function, which can choose to forget and save important data; when the data is multiplied by 1, it means that it is retained, and the data obtained from the input gate means that the state has been updated; finally, with the help of the output gate, the information carried is determined, and the new state and hidden state are transferred to the next time step; its input gate, update gate, forget gate, and output gate can be defined as:

[0022] ;

[0023] ;

[0024] ;

[0025] ;

[0026] in, represents the forget gate, represents the input gate, represents the output gate, represents the update gate, represents the input vector, Indicates the hidden state, Represents a memory unit, represents the weight, Indicates bias, Represents the input vector To the Gate of Oblivion The weight matrix, Represents the input vector To the input gate The weight matrix, Represents the input vector To the output gate The weight matrix, Represents the input vector To the update gate The weight matrix, Indicates the hidden state at the last moment To the Gate of Oblivion The weight matrix, Indicates the hidden state at the last moment To the input gate The weight matrix, Indicates the hidden state at the last moment To the output gate The weight matrix, Indicates the hidden state at the last moment To the update gate The weight matrix, Represents the forget gate The bias term, Represents the update gate The bias term, Represents the input gate The bias term, Represents the output gate The bias term.

[0027] Furthermore, in step D, given the current input state, the context vector is extracted from the state vector. The weighted sum of each sub-state in the current state represents the information associated with the current time step. The query vector is defined as a key and value for learning different weights. The scoring function is used to calculate the correlation between its input vectors and can be defined as:

[0028] ;

[0029] The query vector is defined as , Used to learn keys and values ​​based on different weights, Represents a weight matrix, which is used to perform linear transformation on the vector and adjust the weight of each dimension of information to capture the association between different features.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The present invention proposes a hierarchical standardization processing method for data sets, which enhances data balance through adaptive sample partitioning; the present method constructs a joint spatiotemporal feature extraction model, and strengthens the target signal representation through deep nonlinear mapping; the present invention introduces an embedded attention weight allocation strategy to optimize the focusing ability of important feature channels; the present method can achieve a recognition accuracy of 97% in the five-category ground motion target classification task, which is 36% faster than the processing speed of traditional models; the present method effectively solves the problem of over-reliance on network depth, realizes the automatic extraction of earthquake features and intelligent classification of motion targets, and provides an efficient motion target identification solution for seismic anomalies. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 Schematic diagram for visualizing preprocessed data in a time sliding window;

[0034] Figure 2 Schematic diagram of the micro-motion detection method for ground moving targets in perimeter security according to the present invention;

[0035] Figure 3 Schematic diagram of the convolutional neural network of the present invention;

[0036] Figure 4 Schematic diagram of the attention mechanism embedding module of the present invention. DETAILED DESCRIPTION

[0037] The present invention will be further described below in conjunction with embodiment:

[0038] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0039] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.

[0040] The present invention provides a method for detecting micro-motion of a ground moving target for perimeter security, comprising the following steps:

[0041] A. Collection of earthquake monitoring data of moving targets.

[0042] Collect shallow surface vibration signals induced by ground moving targets, capture seismic waves generated by moving targets, and output original time domain vibration signals.

[0043] In the present invention, shallow surface vibration signals induced by moving targets are collected, specifically, shallow surface vibration signals triggered by five common types of moving targets, namely human activities, wheeled vehicle driving, tracked vehicle driving, aircraft flying and natural noise are collected.

[0044] B. Preprocess the acquired moving target earthquake monitoring data and construct a data set.

[0045] The acquired raw time-domain vibration signals are preprocessed, segmenting the continuous signals, normalizing them, and performing initial feature extraction. The signals are then divided into sliding windows of fixed lengths to ensure that the final learning results cover the entire dataset, meeting real-time detection requirements. 80% of the dataset is used as a training set, and the remainder as a test set.

[0046] Specifically, after acquiring time series data, we first construct unbalanced class samples for the five categories of moving objects. First, we pad the data to unify the unbalanced categories into a standard-sized matrix to obtain a standard sample. A Boolean mask matrix is ​​then overlaid on top of the unbalanced dataset to indicate the padding location. A sliding window is then applied to the time series data, generating a subsequence for each window. For each window, feature extraction, label extraction, normalization, and sliding windowing are performed sequentially. Finally, the time series data is converted into a set of input features and corresponding labels. This completes the dataset construction.

[0047] C. Build a joint spatiotemporal feature extraction model, using a fusion architecture of stacked convolutional neural networks (CNNs) and long short-term memory neural networks (LSTMs) to achieve spatiotemporal feature fusion and feature enhancement, so as to more comprehensively describe the shallow surface vibration signals of moving targets.

[0048] In this invention, CNN is used to capture the spatiotemporal characteristics of short-term models and identify different features in the data, thereby completing the intelligent identification of suspicious targets; and LSTM is used to process the long-term dependencies in the earthquake monitoring system and continuously capture the physical quantities related to time series changes.

[0049] Specifically, four convolutional layers are used to locate and extract spatial features, and each convolutional layer is followed by a max pooling layer for downsampling. An LSTM-based temporal attention module is then added after the final CNN layer to accurately identify and retain important dynamic behaviors in time series of arbitrary lengths while hiding redundant information. Finally, after the expansion layer, a fully connected layer and a softmax classifier are used to classify events based on the learned features.

[0050] Specifically, CNN consists of alternating stacks of convolutional layers and pooling layers. There is an activation function between each convolutional layer and pooling layer to accelerate the convergence of the model. The function σ is the ReLU activation function. The pooling layer is used for downsampling operations and can be expressed as:

[0051] ;

[0052] ;

[0053] in, It is The i-th feature in the layer, It is The j-th feature in the layer, represents the kernel connected to the function, is the deviation of this function.

[0054] After passing the data output from the convolutional module to the long short-term memory module, the decision on what to keep and what to discard is made with the help of the activation function, which allows the choice of forgetting and saving important data. When the data is multiplied by 1, it means it is retained, and the data obtained from the input gate means that the state has been updated. Finally, with the help of the output gate, the information carried is determined, and the new state and hidden state are transferred to the next time step. Its input gate, update gate, forget gate, and output gate can be defined as:

[0055] ;

[0056] ;

[0057] ;

[0058] ;

[0059] in, represents the forget gate, represents the input gate, represents the output gate, represents the update gate, represents the input vector, Indicates the hidden state, Represents a memory unit, represents the weight, Indicates bias, Represents the input vector To the Gate of Oblivion The weight matrix, Represents the input vector To the input gate The weight matrix, Represents the input vector To the output gate The weight matrix, Represents the input vector To the update gate The weight matrix, Indicates the hidden state at the last moment To the Gate of Oblivion The weight matrix, Indicates the hidden state at the last moment To the input gate The weight matrix, Indicates the hidden state at the last moment To the output gate The weight matrix, Indicates the hidden state at the last moment To the update gate The weight matrix, Represents the forget gate The bias term, Represents the update gate The bias term, Represents the input gate The bias term, Represents the output gate The bias term.

[0060] D. Introduce the temporal pattern attention mechanism into the joint spatiotemporal feature extraction model, select relevant time series, remember important features of the entire time domain, and dynamically adjust feature weights to make the model pay more attention to key times and key signals and suppress irrelevant noise.

[0061] Given the current input state, the context vector is extracted from the state vector. The weighted sum of each sub-state in the current state represents the information associated with the current time step. The query vector is defined as a key and value for learning based on different weights. The scoring function is used to calculate the correlation between its input vectors and can be defined as:

[0062] ;

[0063] The query vector is defined as , Used to learn keys and values ​​based on different weights, Represents a weight matrix, which is used to perform linear transformation on the vector and adjust the weight of each dimension of information to capture the association between different features.

[0064] E. Train and test the model obtained in step D to identify shallow surface abnormal vibration data and automatically extract and classify it.

[0065] The detection device used in the aforementioned method for detecting micro-motion on moving ground targets for perimeter security includes a moving target seismic motion acquisition module, a time sliding window data preprocessing module, a spatiotemporal feature integration module, and an attention mechanism embedding module. The detection device utilizes a fusion architecture of a stacked convolutional neural network and a long-short-term memory module, integrating the preprocessing module's masking mechanism and spatial attention mechanism to dynamically capture multi-dimensional features.

[0066] The moving target seismic motion acquisition module completes the data perception part, collects shallow surface vibration signals through node seismometers, captures seismic waves generated by moving targets, and outputs original time domain vibration signals. The time sliding window data preprocessing module processes continuous signals in segments, standardizes and completes initial feature extraction, and divides them into sliding windows of fixed length to meet real-time detection requirements. A Boolean mask matrix is ​​added to the time sliding window data preprocessing module to ignore the padding position of adaptive samples and avoid processing meaningless padding values, thereby improving the effectiveness of the model. The spatiotemporal feature joint module is used for spatiotemporal feature fusion and feature enhancement to more comprehensively describe the seismic motion signals of moving targets. The core task of the attention mechanism embedding module is to dynamically adjust the feature weights so that the model pays more attention to key times and key signals and suppresses irrelevant noise.

[0067] Example 1:

[0068] A method for detecting micro-motion of ground moving targets for perimeter security comprises the following steps:

[0069] 1. Collection of earthquake monitoring data of moving targets.

[0070] A detection device was installed on a field with flat terrain and relatively uniform geological conditions to conduct on-site inspections. The detection device is based on a wireless ad hoc network and is capable of real-time monitoring. The detection device includes a node seismometer that can collect shallow surface vibration signals induced by five common ground moving targets, capture the seismic waves generated by the moving targets, and output the original time-domain vibration signal. During the data collection process, only one moving target moves near the node seismometer at a time. The trajectory of the controlled target is a round-trip straight line parallel to the road. Specific collection methods are used for the trajectories and data of different types of targets to avoid unacceptable energy differences between different types of targets. Finally, the collected signals are transmitted to the control end for data processing using a computer.

[0071] 2. Preprocess the acquired moving target seismic monitoring data and construct a data set.

[0072] The acquired original time-domain vibration signal is preprocessed, the continuous signal is segmented, standardized and initially extracted, and divided into sliding windows of fixed length to obtain the required data set to meet real-time detection needs.

[0073] Obtain the time series and first construct imbalanced category samples of the five categories of moving targets. Since the information and natural noise contained in the aircraft are complex, a large amount of data is required to learn, and all samples cannot be fully guaranteed in actual engineering. Therefore, the constructed imbalanced category samples include 1-minute seismic signals of human footsteps, 4-minute seismic signals of wheeled vehicles, 2-minute seismic signals of tracked vehicles, 22-minute seismic signals caused by aircraft, and 30-minute seismic signals caused by natural noise. First, a padding operation is applied to unify the imbalanced data into a standard size matrix. The Boolean mask matrix is ​​first overlaid on the top of the imbalanced dataset to indicate which positions are filled. During the calculation of the deep learning model, these filled positions will be ignored to avoid processing meaningless filled values, thereby improving the effectiveness of the model. Then, a sliding window is applied on top of the time series data, and a subsequence is generated for each window. For each window, feature extraction, label extraction, normalization, and sliding window operations are performed sequentially. Finally, the time series data is converted into a set of input features and corresponding labels, such as Figure 1 As shown. Thus, the construction of the data set is completed. In this embodiment, the preprocessing of the signal ensures that the final learning result covers the entire data set, wherein 80% of the data set is used as a training set and the rest is used as a test set.

[0074] 3. A joint spatiotemporal feature extraction model is constructed, employing a fusion architecture of stacked convolutional neural networks (CNNs) and long short-term memory neural networks (LSTMs) to achieve spatiotemporal feature fusion and enhancement, thereby more comprehensively describing the shallow surface seismic signals of moving targets. By utilizing CNNs to capture the spatiotemporal features of short-term models, the neural network can learn valuable information from seismic data without relying on explicit physical equations, identifying distinct features within the data and thus intelligently identifying suspicious targets. Furthermore, by processing long-term dependencies within the detection device through LSTMs, the system can continuously capture physical quantities related to time series changes.

[0075] Four convolutional layers are used to localize and extract spatial features, and each convolutional layer is followed by a max pooling layer for downsampling. An LSTM-based temporal attention module is then added after the final CNN layer to accurately identify and retain important dynamic behaviors in time series of arbitrary lengths while hiding redundant information. Finally, after the expansion layer, a fully connected layer and a softmax classifier are used to classify events based on the learned features.

[0076] Specifically, CNN consists of alternating stacks of convolutional and pooling layers, e.g. Figure 3 As shown, there is an activation function between each convolution layer and pooling layer to accelerate the convergence of the model. The function σ is the ReLU activation function. The pooling layer is used for downsampling operation and can be expressed as:

[0077] ;

[0078] ;

[0079] in, It is The i-th feature in the layer, It is The j-th feature in the layer, represents the kernel connected to the function, is the deviation of this function.

[0080] After passing the data output from the convolutional module to the long short-term memory module, the decision on what to keep and what to discard is made with the help of the activation function, which allows the choice of forgetting and saving important data. When the data is multiplied by 1, it means it is retained, and the data obtained from the input gate means that the state has been updated. Finally, with the help of the output gate, the information carried is determined, and the new state and hidden state are transferred to the next time step. Its input gate, update gate, forget gate, and output gate can be defined as:

[0081] ;

[0082] ;

[0083] ;

[0084] ;

[0085] in, represents the forget gate, represents the input gate, represents the output gate, represents the update gate, represents the input vector, Indicates the hidden state, Represents a memory unit, represents the weight, Indicates bias.

[0086] in, Represents the input vector To the Gate of Oblivion The weight matrix is ​​used to measure the influence of input information on the forget gate decision. Represents the input vector To the input gate The weight matrix is ​​used to measure the decision of the input information on the input gate. Represents the input vector To the output gate The weight matrix is ​​used to measure the impact of input information on the output gate decision. Represents the input vector To the update gate The weight matrix is ​​used to measure the impact of input information on the update gate. Indicates the hidden state at the last moment To the Gate of Oblivion The weight matrix reflects the effect of the hidden state on the forget gate at the previous moment. Indicates the hidden state at the last moment To the input gate The weight matrix represents the influence of the hidden state on the input gate at the previous moment. Indicates the hidden state at the last moment To the output gate The weight matrix reflects the effect of the hidden state on the output gate at the previous moment. Indicates the hidden state at the last moment To the update gate The weight matrix represents the effect of the previous hidden state on the update gate.

[0087] in, Represents the forget gate The bias term is used to adjust the weighted sum result when calculating the forget gate value, so that the model has stronger expressive power. Represents the update gate The bias term is used to adjust the weighted sum result when calculating the update gate value. Represents the input gate The bias term is used to adjust the weighted sum result when calculating the input gate value. Represents the output gate The bias term is used to adjust the weighted sum result when calculating the output gate value.

[0088] The overall framework of the spatiotemporal feature joint extraction model constructed in this embodiment is as follows: Figure 2 As shown in the figure, this example uses four convolutional layers to locate and extract spatial features, with a max pooling layer following each convolutional layer for downsampling. An LSTM-based temporal attention module is then added after the last CNN layer to accurately identify and retain important dynamic behaviors in time series of any length while hiding redundant information. Finally, after the expansion layer, a fully connected layer and a softmax classifier are sequentially applied to classify events based on the learned features. The detailed structure of this model includes four convolutional layers (CLSs), four max pooling (MP) layers, one dropout layer, one LSTM layer, one TAM layer, and one fully connected (FC) layer. The number of convolutional kernels in the four CLSs is 18, 36, 72, and 32, respectively, all with a size of 2×1. The stride of all kernels is 1. Batch normalization is used after each convolution and before activation to improve convergence. The LSTM layer has 64 hidden units, and the number of layers in the TAM network is 125. The number of neurons in the output layer is determined by the number of categories of the moving objects. Furthermore, regularization is added to the model to improve the model's generalization ability. In summary, CNN-LSTAM automatically extracts earthquake features and intelligently classifies earthquakes through multi-layer nonlinear learning and multi-layer learning and inference in the time domain.

[0089] 4. Introduce the temporal pattern attention mechanism into the joint spatiotemporal feature extraction model, select relevant time series, remember important features in the entire time domain, and dynamically adjust feature weights, so that the model pays more attention to key times and key signals and suppresses irrelevant noise.

[0090] Given the current input state, the context vector is extracted from the state vector. The weighted sum of each sub-state in the current state represents the information associated with the current time step. The query vector is defined as a key and value for learning based on different weights. The scoring function is used to calculate the correlation between its input vectors and can be defined as:

[0091] ;

[0092] The query vector is defined as , Used to learn keys and values ​​based on different weights, Represents a weight matrix, which is used to perform linear transformation on the vector and adjust the weight of each dimension of information to capture the association between different features.

[0093] Regarding the problems of training instability and gradient disappearance, the long short-term memory module cannot remember early interdependencies, and the addition of a dynamic bidirectional mechanism will lead to a sharp increase in the number of parameters and prolonged calculation time. To solve this problem, a temporal pattern attention mechanism can be introduced. The traditional attention mechanism only extracts the previous time step, selects relevant hidden states, and ignores information from other time steps. In contrast, the attention mechanism proposed in this embodiment selects relevant time series, can remember important features of the entire time domain, is applicable to various data sets including non-periodic data sets, saves calculation time, and helps alleviate the scale insensitivity and complexity of neural networks. In this embodiment, the attention mechanism embedding module is shown in Figure 4.

[0094] 5. Train and test the model obtained in step 4 to identify shallow surface abnormal vibration data and automatically extract and classify it.

[0095] The detection device used in the aforementioned method for detecting micro-motion on the ground for perimeter security includes a moving target seismic motion acquisition module, a time sliding window data preprocessing module, a spatiotemporal feature integration module, and an attention mechanism embedding module. The detection device utilizes a fusion architecture of a stacked convolutional neural network and a long-short-term memory module, integrating the preprocessing module's masking mechanism and spatial attention mechanism to dynamically capture multi-dimensional features.

[0096] The moving target seismic motion acquisition module completes the data perception part, collects shallow surface vibration signals through node seismometers, captures seismic waves generated by moving targets, and outputs original time domain vibration signals. The time sliding window data preprocessing module processes continuous signals in segments, standardizes and completes initial feature extraction, and divides them into sliding windows of fixed length to meet real-time detection requirements. A Boolean mask matrix is ​​added to the time sliding window data preprocessing module to ignore the padding position of adaptive samples and avoid processing meaningless padding values, thereby improving the effectiveness of the model. The spatiotemporal feature joint module is used for spatiotemporal feature fusion and feature enhancement to more comprehensively describe the seismic motion signals of moving targets. The core task of the attention mechanism embedding module is to dynamically adjust the feature weights so that the model pays more attention to key times and key signals and suppresses irrelevant noise.

[0097] The detection method used in this embodiment can achieve a recognition accuracy of 97% in the classification task of five types of ground moving targets, which is 36% faster than the processing speed of traditional models.

[0098] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the concept of the present invention. The scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for detecting micro-motion of ground moving targets for perimeter security, characterized in that: The following steps are involved: A. Collection of moving target seismic monitoring data: Collect shallow surface vibration signals induced by ground moving targets, capture seismic waves generated by moving targets, and output original time domain vibration signals; B. Preprocess the acquired moving target earthquake monitoring data and construct a data set: The acquired raw time-domain vibration signals are preprocessed, with continuous signals segmented, standardized, and initially extracted, then divided into sliding windows of fixed length to ensure that the final learning results cover the entire dataset to meet real-time detection requirements. 80% of the dataset is used as a training set, and the remainder as a test set. Specifically, after acquiring the time series data, we first construct imbalanced category samples for the five categories of moving objects. First, we fill the data to unify the imbalanced categories into a standard size matrix to obtain a standard sample, which is then overlaid on top of the imbalanced dataset using a Boolean mask matrix to indicate the filling position. Then, we apply a sliding window on the time series data and generate a subsequence for each window. For each window, we perform feature extraction, label extraction, normalization, and sliding windowing in sequence. Finally, we convert the time series data into a set of input features and corresponding labels. This completes the construction of the dataset. C. Build a joint spatiotemporal feature extraction model, using a fusion architecture of stacked CNN and LSTM to achieve spatiotemporal feature fusion and feature enhancement. CNN is used to capture the spatiotemporal features of short-term models and identify different features in the data, thereby completing the intelligent identification of suspicious targets. LSTM is also used to continuously capture the physical quantities related to time series changes. Specifically, four convolutional layers are used to locate and extract spatial features, and a maximum pooling layer is connected after each convolutional layer to perform downsampling. Then, an LSTM-based temporal attention module is added after the last CNN layer to accurately identify and retain important dynamic behaviors of time series of arbitrary lengths and hide redundant information. Finally, after the expansion layer, a fully connected layer and a softmax classifier are sequentially accessed to classify events based on the learned features. CNN consists of alternating stacks of convolutional layers and pooling layers. There is an activation function between each convolutional layer and pooling layer to accelerate the convergence of the model. The function σ is the ReLU activation function. The pooling layer is used for downsampling operations and can be expressed as: σ(y)=max(0,y) in, is the i-th feature in the l-th layer, is the jth feature in the l-1th layer, represents the kernel connected to the function, is the deviation of this function; After passing the data output from the convolution module to the long short-term memory module, the decision on what to keep and what to discard is done with the help of the activation function, which can choose to forget and save important data; when the data is multiplied by 1, it means that it is retained, and the data obtained from the input gate means that the state has been updated; finally, with the help of the output gate, the information carried is determined, and the new state and hidden state are transferred to the next time step; its input gate, update gate, forget gate, and output gate can be defined as: f t =σ(w xf x t +w hf h t-1 +b f ) i t =σ(w xi x t +w hi h t-1 +b i ) o t =σ(w xo x t +w ho h t-1 +b o ) g t =tanh(w xc x t +w hc h t-1 +b c ) Among them, f t represents the forget gate, i t represents the input gate, o t represents the output gate, g t represents the update gate, x represents the input vector, h represents the hidden state, c represents the memory unit, w represents the weight, and b represents the bias; D. Introducing the temporal pattern attention mechanism into the joint spatiotemporal feature extraction model, it selects relevant time series, remembers important features across the entire time domain, and dynamically adjusts feature weights, allowing the model to focus more on key times and key signals while suppressing irrelevant noise. Given the current input state, the context vector is extracted from the state vector. The weighted sum of each sub-state in the current state represents the information associated with the current time step. The query vector is defined as a key and value for learning based on different weights. The scoring function is used to calculate the correlation between its input vectors and can be defined as: Here, the query vector is defined as h i , h t Used to learn keys and values ​​based on different weights; E. Train and test the model obtained in step D to identify shallow surface abnormal vibration data and automatically extract and classify it.

2. The method for detecting micro-motion of ground moving targets for perimeter security according to claim 1, characterized in that: In step A, shallow surface vibration signals induced by moving targets are collected. Specifically, shallow surface vibration signals triggered by five common types of moving targets are collected, namely human activities, wheeled vehicle driving, tracked vehicle driving, aircraft flying, and natural noise.

3. The detection device adopted by the micro-motion detection method for ground moving targets for perimeter security according to claim 1 includes a moving target seismic motion acquisition module, a time sliding window data preprocessing module, a spatiotemporal feature combination module and an attention mechanism embedding module; the moving target seismic motion acquisition module completes the data perception part, collects shallow surface vibration signals through node seismometers, captures seismic waves generated by moving targets, and outputs original time domain vibration signals; the time sliding window data preprocessing module segments the continuous signal, standardizes it and completes the initial feature extraction, and divides it into sliding windows of fixed length to meet the real-time detection requirements; a Boolean mask matrix is ​​added to the time sliding window data preprocessing module to ignore the filling position of the adaptive sample; the spatiotemporal feature combination module is used for spatiotemporal feature fusion and feature enhancement; the attention mechanism embedding module can dynamically adjust the feature weights so that the model pays more attention to key times and key signals and suppresses irrelevant noise.

Citation Information

Patent Citations

  • Deep fusion network production line fault prediction method based on deep learning

    CN119357769A

  • TTPA-LSTM soft measurement method based on multiple sampling rates

    CN119673323A

  • Intelligent gold bonding wire length identification method and system

    CN119785050A

  • Automatic side slope rolling stone micro-seismic signal detection method

    CN119846692A