Perimeter sound signal identification method and system based on distributed optical fiber sensing
Through distributed fiber sensing technology, combined with MobileNetV3 and LSTM networks, a perimeter sound signal recognition system is built, which solves the problem of high deployment and maintenance costs of perimeter security systems in complex environments, and realizes efficient and accurate perimeter sound signal monitoring and identification.
Patent Information
- Application Number
- CN202510464806.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
AI Technical Summary
The existing perimeter security system is costly to deploy and maintain in complex environments, and it is difficult to effectively process and analyze perimeter security data, and it is impossible to effectively identify and monitor sound signals, resulting in failure to prevent security threats in a timely manner.
Distributed fiber sensing technology is adopted to collect and preprocess fiber signal image data, build and train fiber sensing network models, combine MobileNetV3 and LSTM networks to perform data noise reduction and feature extraction, monitor and identify perimeter sound signals in real time, and adapt to different environments.
It realizes efficient and accurate identification and monitoring of perimeter sound signals in complex environments, reduces the impact of noise, improves the computing efficiency and adaptability of the model, is suitable for a variety of equipment, and reduces deployment and maintenance costs.
Smart Images

Figure CN120388223A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optical fiber sensing, and particularly to a method and system for identifying perimeter sound signals based on distributed optical fiber sensing. Background Art
[0002] Distributed optical fiber sensing technology uses optical fiber as a sensing element and a signal transmission medium. By detecting vibration signals at different positions along the optical fiber, real-time and continuous monitoring of the perimeter area can be achieved. Optical fiber can shield electromagnetic interference and radio frequency interference, and has characteristics such as explosion-proof and high stability. It can adapt to harsh working environments and has the advantages of low cost and high efficiency. Perimeter security systems are widely used in high-security requirement places such as airports, military bases, and nuclear power plants, and the market demand continues to grow. At the same time, with the promotion of smart city construction, the application of perimeter security systems in civilian fields such as business parks and residential communities is also becoming more and more popular. Perimeter security systems need to work in various complex environments, such as different terrains and climate conditions. These environmental factors may affect the performance of the sensor, and the system needs to have good adaptability, and the deployment and maintenance costs are relatively high. And how to effectively process and analyze the data generated by perimeter security and extract valuable information is also a challenge. In view of this situation, the present invention proposes a method for identifying perimeter sound signals based on distributed optical fiber sensing. By monitoring perimeter sound and vibration signals, the perimeter situation can be monitored in real time, perimeter damage can be prevented from occurring, and perimeter security can be ensured. Summary of the Invention
[0003] In order to solve the above-mentioned problems, the present invention provides a method and system for identifying perimeter sound signals based on distributed optical fiber sensing. By collecting and processing perimeter sound signals, for certain intrusion events, real-time monitoring, rapid alarm, accurate positioning can be carried out to prevent the occurrence of perimeter security damage events.
[0004] In the first aspect, a method for identifying perimeter sound signals based on distributed optical fiber sensing provided by the present invention adopts the following technical solution:
[0005] A method for identifying perimeter sound signals based on distributed optical fiber sensing includes:
[0006] Collecting distributed optical fiber signal image data;
[0007] Performing data preprocessing on the collected distributed optical fiber signal image data;
[0008] Constructing a distributed optical fiber sensing network model;
[0009] Training the distributed optical fiber sensing network model by using the distributed optical fiber signal image data;
[0010] Deploy the trained model to the on-site environment for use and fine-tune it according to the on-site environment.
[0011] Furthermore, the acquisition of distributed fiber optic signal image data includes building a distributed fiber optic sound signal recognition system. The distributed fiber optic sound signal recognition system includes a fiber optic perimeter security monitoring and alarm host DAS, which is respectively connected to the overhead optical cable and the buried optical cable. Among them, first, arrange a several-shaped overhead section of about 130m; and conduct 2 buried sections, bury the buried optical cable under 50 cm of soil, each with a length of 20m; arrange a straight-shaped overhead section, and leave 30m of coiled optical cable between different scenarios.
[0012] Furthermore, the data preprocessing of the acquired distributed fiber optic signal image data includes preprocessing the DAS data, including data denoising and downsampling to reduce noise and unnecessary data fluctuations; then, the processed data is segmented by means of a sliding window, converted into images, and the images are screened. Select those with event signals as event sample data and those without event signals as negative samples.
[0013] Furthermore, the data denoising includes using the non-local means method for denoising. Classify and weighted average the regions with the same nature in the same image to obtain the denoised image, and use redundant information to remove noise, which is used to filter out Gaussian noise in the image. Among them, given a discrete noisy image v = {v(i)|i ∈ I}, the estimated value NL[v](i) of pixel i is calculated as the weighted average of all pixels in the image. The filtering process of NLM is expressed as:
[0014]
[0015] where w represents the weight.
[0016] Furthermore, the data preprocessing of the acquired distributed fiber optic signal image data also includes, before the image data is input into the MobileNetV3 network, first adjusting the image to the expected input size of MobileNetV3, that is, 224x224 pixels, and then normalizing the image using specific mean and standard deviation values to improve the stability and convergence speed of model training.
[0017] Furthermore, the construction of the distributed fiber optic sensing network model includes combining the MobileNetV3 network and the LSTM network to build a neural network framework. Among them, based on the lightweight neural network architecture of the MobileNetV3 network, depthwise separable convolution, inverted residual structure, and squeeze-and-excitation structure are introduced, and the hard-swish activation function is used to replace the traditional activation function to improve the operation efficiency of the model and be applicable to low-computing-power hardware devices.
[0018] Further, the model training of the distributed optical fiber sensing network using the distributed optical fiber signal image data includes inputting the preprocessed image into the MobileNetV3 network to extract event signal features. In the inverted residual structure, the number of input channels is first expanded through 1x1 convolution, then depthwise separable convolution is applied to reduce the amount of computation and the number of parameters, and finally the number of channels is compressed through 1x1 convolution, so as to maintain the computational efficiency while increasing the network depth.
[0019] Further, the model training of the distributed optical fiber sensing network using the distributed optical fiber signal image data also includes inputting the extracted event signal features into the LSTM for recognition. During the training process, the training of specific data is enhanced by optimizing the forget gate. Among them, the weights and biases of the forget gate are dynamically adjusted according to the current input and the previous state, so that the network adaptively adjusts the degree of forgetting according to the loss value of the previous batch of data.
[0020] Further, the step of enabling the network to adaptively adjust the degree of forgetting according to the loss value of the previous batch of data includes adding the loss value of the previous batch of data to the input of the forget gate, and dynamically adjusting the forget gate through an algorithm strategy with the loss value of the current data. When information is discarded by the Sigmoid function, the smaller the error between the current loss value and the loss value of the previous batch, the more stable the model is and the less information is discarded; the larger the error, the greater the model fluctuation and the more information is discarded. Among them, the input of the current batch is expressed as:
[0021]
[0022] where, x i represents the input at time t, represents the input weight, E represents the forgetting factor, h t―1 represents the output at time t - 1, b f represents the bias term of the forget gate, and the calculation method of E is:
[0023]
[0024] where, L t represents the loss value at time t, and L t―1 represents the loss value at time t - 1.
[0025] Further, the network adaptively adjusts the degree of forgetting according to the loss value of the previous batch of data, including dynamically adjusting the weight of the forgetting gate through an improved algorithm strategy, so that the model flexibly controls the degree of information forgetting according to the difference between the loss values of the current batch and the previous batch. When the error between the loss values of the current batch and the previous batch is less than the set value, the model tends to be stable, and the discarded information is at a low level at this time, thus ensuring the high-fidelity recognition of the model for stable signals; when the error is greater than the set value, the model fluctuates greatly and the discarded information is at a high level at this time, which is conducive to quickly adjusting the model state and avoiding the decline in recognition accuracy caused by overfitting.
[0026] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the method for perimeter sound signal recognition based on distributed optical fiber sensing.
[0027] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the method for perimeter sound signal recognition based on distributed optical fiber sensing.
[0028] In summary, the present invention has the following beneficial technical effects:
[0029] The present invention designs a method for perimeter sound signal recognition based on distributed optical fiber sensing, deeply combines the spatial feature extraction ability of MobileNetV3 and the timing processing ability of LSTM, constructs a powerful end-to-end model, extracts event information, predicts event types, and achieves accurate positioning and real-time monitoring.
[0030] 1) A method of non-local means (NLM) is used to denoise the data. After filtering, the image has high clarity, does not lose details, and reduces the impact of data background noise on the training results.
[0031] 2) Through the analysis of the application scenarios of distributed optical fiber sensing, the MobileNetV3 network and the LSTM network are combined to build a neural network framework. On the one hand, rich spatial features can be extracted from the data to provide strong feature support for visual tasks; on the other hand, the timing features of optical fiber signals can be better utilized to sense the dynamic changes of event signals, and the computational efficiency and model size are optimized, the number of parameters is reduced, making the model more lightweight and suitable for running on various devices.
[0032] 3) During the training process, some optimizations and improvements were made to the LSTM network. The forget gate was improved, and the gating strategy was adjusted to better control the information flow, improve the network's training ability for difficult samples, and accelerate the model convergence speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is the model structure diagram of the method of the present invention in Embodiment 1 of the present invention;
[0034] Figure 2 is the flowchart of the MobileNetV3 network in Embodiment 1 of the present invention;
[0035] Figure 3 is the flowchart of the distributed optical fiber system in Embodiment 1 of the present invention;
[0036] Figure 4 is the waterfall diagram of various events in Embodiment 1 of the present invention;
[0037] Figure 5 is the LSTM network structure diagram in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0038] The present invention will be further described in detail below with reference to the accompanying drawings.
[0039] Embodiment 1
[0040] Refer to Figure 1 , a perimeter sound signal recognition method based on distributed optical fiber sensing in this embodiment includes:
[0041] Collect distributed optical fiber signal image data;
[0042] Perform data preprocessing on the collected distributed optical fiber signal image data;
[0043] Construct a distributed optical fiber sensing network model;
[0044] Use the distributed optical fiber signal image data to train the distributed optical fiber sensing network model;
[0045] Deploy the trained model to the on-site environment for use and fine-tune it according to the on-site environment.
[0046] Specifically:
[0047] S1. Build a distributed optical fiber sensing monitoring system, connect the optical cable to the host device through a jumper, and adjust parameters such as the pulse width and sampling points according to the on-site situation to make it reach the best state.
[0048] Among them, a distributed fiber optic sound signal recognition system is built, which is divided into overhead optical cables and buried optical cables. At the beginning, there is a "J"-shaped overhead section about 130 m long; then there are 2 buried sections. The buried optical cables are buried 50 cm underground and are about 20 m long respectively; then there is another "I"-shaped overhead section about 60 m long. There is about 30 m of coiled optical cable left between different scenarios. Install a fiber optic perimeter security monitoring and alarm host (DAS) and necessary signal processing equipment in the control center, then connect the optical cable to the DAS device, and monitor the current status of the optical cable through the distributed fiber optic sound signal recognition system management software. The equipment includes functions such as broken fiber alarm setting, zoning setting, data collection, device alarm, and alarm record query.
[0049] S2. Collect data. According to the built scenarios, collect various types of event data, including five types of events: personnel climbing, drones, personnel walking, manual excavation, and background. For each type of event, 3 - 5 people take turns to carry out, so as to enrich the diversity of data. And the time for a single type of event is about one hour. When the event is in progress, save the global data.
[0050] Among them, after the host device is debugged, collect data in the order of background, personnel climbing, drones, personnel walking, and manual excavation. When collecting data, 3 - 5 people take turns to carry out to prevent the signal of a single type of event from being too fixed. Collect data for about one hour for each type of event and save the data.
[0051] S3. Data processing. First, preprocess the DAS data, including data denoising and downsampling, to reduce noise and unnecessary data fluctuations. Then, use the sliding window method to segment the processed data and convert it into images. Screen the images, select those with event signals as event sample data, and those without event signals as negative samples.
[0052] Among them, use the non - local means method for denoising. Classify and weighted average the regions with the same nature in the same image to obtain the denoised picture. The more images are weighted, the better the denoising effect will be. This algorithm uses the redundant information commonly existing in natural images to remove noise and can better filter out Gaussian noise in the image.
[0053] Given a discrete noisy image v = {v(i)|i ∈ I}, the estimated value NL[v](i) of pixel i is calculated as the weighted average of all pixels in the image. The filtering process of NLM can be represented by the following formula:
[0054]
[0055] Among them, w represents the weight.
[0056] In the non - local means denoising algorithm, the reasonable allocation of weights is the core of achieving effective denoising. This algorithm breaks through the spatial limitation of traditional local filtering by establishing the cross - domain similarity association between pixels. Specifically, when calculating the estimated value of pixel \(i\), instead of simply averaging its adjacent pixels, it searches the entire image to find regions with similar texture patterns to it. The quantification of this similarity requires precise mathematical expressions.
[0057] There are many ways to measure similarity. The most commonly used method is to estimate based on the square of the difference in brightness between two pixels. Due to the presence of noise, a single pixel is not reliable, so its neighborhood is used. Only when the neighborhood similarity is high can it be said that the similarity between the two pixels is high. The most commonly used method to measure the similarity between two image patches is to calculate the Euclidean distance between them:
[0058]
[0059] where \(Z(i)\) is the normalization constant.
[0060]
[0061] where \(a\) is the standard deviation of the Gaussian kernel.
[0062] When calculating the Euclidean distance, the weights of pixels at different positions are different. The closer to the center of the block, the greater the weight, and the farther from the center, the smaller the weight. The weights follow a Gaussian distribution.
[0063] Generally speaking, considering the algorithm complexity, the search area is approximately \(21\times21\), and the blocks for similarity comparison can be \(7\times7\). In practice, it is often necessary to select appropriate parameters according to the noise. When the standard deviation \(\sigma\) of Gaussian noise is larger, in order to improve the algorithm robustness, it is necessary to increase the block area and also increase the search area. At the same time, the filtering coefficient \(h\) and \(\sigma\) are positively correlated: \(h = k\sigma\). When the block becomes larger, \(k\) needs to be appropriately reduced.
[0064] S4. After data denoising, the data is sliced into images, and different signals formed by different events can be clearly viewed. The specific effect is as shown in the appendix Figure 4 After screening out the images with obvious event characteristics, use the annotation tool to annotate various events. The image data cannot be too few, about 2500 for each type of image.
[0065] All the data is divided into a training set, a validation set, and a test set according to the ratio of \(7:2:1\). The batchsize is 64, the optimizer uses Adam, the initial learning rate is 0.001, the training device uses the 3080Ti graphics card, and train for 500 rounds.
[0066] S5. Model construction
[0067] Considering making full use of the efficient feature extraction ability of MobileNetV3 and the temporal data perception ability of LSTM, MobileNetV3 and LSTM are combined. This architecture can give play to the advantages of lightweight convolutional neural networks and recurrent neural networks and is suitable for processing data with spatio-temporal characteristics.
[0068] The specific architecture of the distributed optical fiber sensing network model includes a multi-level feature fusion and spatio-temporal correlation analysis module. In the framework where the MobileNetV3 network and the LSTM network work together, first, the preprocessed 224×224 pixel image is input into the depthwise separable convolutional network of MobileNetV3 for spatial feature extraction. Its improved SE (Squeeze-and-Excitation) attention module is used to enhance the features of the abnormal vibration area in the optical fiber signal, and the hard-swish activation function is used to enhance the expression ability of high-frequency features. The feature tensor processed by the inverted residual structure is dimensionally reduced through the global average pooling layer, and a feature vector sequence with a dimension of 1280 is output. The LSTM network uses a bidirectional double-layer structure to process temporal features, and each time step receives the feature vector output by MobileNetV3. The network sets 256 hidden units, and the gating mechanism is used to capture the propagation law of the optical fiber signal in the time dimension, especially for modeling the common mechanical vibration waveforms in distributed optical fiber sensing. A dropout rate of 0.3 is introduced between the LSTM layers to prevent overfitting. At the same time, an adaptive temporal attention mechanism is connected to the output end, and the features of key time nodes are dynamically weighted through a learnable weight matrix to improve the detection sensitivity to sudden events.
[0069] The specific construction process is as follows:
[0070] (1) Spatial feature extraction layer
[0071] MobileNetV3 is used as the backbone network, and its core lies in the collaborative design of depthwise separable convolution and the attention mechanism. After the input image is normalized to a size of 224×224, it enters 16 convolutional modules for processing:
[0072] The first 3 layers use 3×3 standard convolution (stride = 2) to quickly compress the spatial dimension to 112×112;
[0073] The subsequent 13 layers are stacked by inverted residual structures, and each module includes:
[0074] Expansion convolution (1×1, expansion coefficient 6) → Depth convolution (3×3 or 5×5) → SE attention module (compression ratio 0.25) → 1×1 projection convolution
[0075] Finally, a feature tensor of 7×7×1280 is output, which is reduced to a 1280-dimensional feature vector through global average pooling.
[0076] (2) Temporal processing layer
[0077] The spatial feature sequence output by MobileNetV3 is input into a bidirectional LSTM network:
[0078] A two-layer structure is designed, with each layer containing 512 hidden units
[0079] The temporal window is set to 10 frames, and each frame corresponds to 1280-dimensional input features
[0080] A variant of the gated recurrent unit (GRU) is introduced and implemented using optimized LSTM:
[0081]
[0082] Dropout with a rate of 0.4 and Layer Normalization are added to the output layer to suppress overfitting
[0083] (3) Classification decision layer
[0084] The spatio-temporal features are mapped to the target category space through a fully connected network:
[0085] The first-layer FC: 2048 neurons, activated by ReLU, and the final hidden state of the input LSTM (1024-dimensional) is input
[0086] The second-layer FC: 512 neurons, activated by Swish, and L2 regularization is added (λ = 0.001)
[0087] The output-layer FC: the number of neurons is equal to the number of categories, and it is connected to the SoftMax function to generate a probability distribution:
[0088] (C is the total number of categories).
[0089] S6. Before the image data is input into the MobileNetV3 network, the image is first adjusted to the expected input size of MobileNetV3, which is 224x224 pixels. Then the image is normalized using specific mean and standard deviation values to improve the stability and convergence speed of model training. The image is processed by scaling and cropping to adjust it to the expected input size of 224x224 pixels of the MobileNetV3 model. MobileNetV3 is a lightweight deep learning model that uses a variety of advanced techniques during feature extraction to optimize computational efficiency and model performance. The structure is as shown in the appendix Figure 2 as follows.
[0090] Model training. Innovatively combine the MobileNetV3 network and the LSTM network in the field of distributed optical fiber sensing to build a new training network framework, the structure of which is as shown in the appendix Figure 1 as follows. The specific training process is as follows:
[0091] Input the preprocessed images into the MobileNetV3 network to extract event signal features. The inverted residual structure first expands the number of input channels through 1x1 convolution, then applies depthwise separable convolution to greatly reduce the computational complexity and the number of parameters, and finally compresses the number of channels through 1x1 convolution, thereby maintaining the computational efficiency while increasing the network depth.
[0092] When constructing a model based on MobileNetV3, the preprocessed images are first input into the MobileNetV3 network to extract event signal features. One of the core structures of MobileNetV3 is the inverted residual structure (Inverted Residual Block), which extracts features through a series of efficient convolution operations while significantly reducing the computational complexity and the number of parameters.
[0093] The feature maps extracted by MobileNetV3 and the time series features processed by LSTM will be fused, adopting a multi-scale spatio-temporal feature interaction mechanism. The hierarchical spatial features (output dimension 1280×7×7) extracted by the MobileNetV3 backbone network and the temporal feature vectors (dimension 256×T) output by LSTM are aligned and fused through a cross-modal attention module. Specifically, a bidirectional feature pyramid network (BiFPN) is designed as a connection bridge to upsample the temporal features of LSTM to the resolution matching different hierarchical feature maps of MobileNetV3, and weighted summation is performed through learnable weight coefficients to form a composite feature tensor with both spatial details and motion trajectory information. These features are then sent to the network detection head part to obtain detection results, and the detection results will be organized into an array containing bounding box coordinates, confidence levels, and class labels. These results can be used for subsequent visualization or further processing.
[0094] The specific algorithm steps include:
[0095] 1. Input feature map: Assume the shape of the input feature map is (N, C in, H, W), where N is the batch size, C in is the number of input channels, and H and W are the height and width of the feature map.
[0096] 2. Expand the number of channels: Expand the number of input channels from C in to C exp through 1x1 convolution to obtain a feature map with the shape of (N, Cexp, H, W).
[0097] 3. Depthwise separable convolution:
[0098] Depth convolution: Perform convolution operations on each channel separately, keeping the spatial dimension of the feature map unchanged, and the output shape remains (N, C_exp, H, W).
[0099] Pointwise convolution: Use 1x1 convolution to perform cross-channel fusion on the output of depth convolution, and the output shape is (N, C_exp, H, W).
[0100] 4. Compress the number of channels: Compress the number of channels from C_exp to C_out through 1x1 convolution, and the final output feature map shape is (N, C_out, H, W).
[0101] Then, input the extracted event signal features into the LSTM for recognition. LSTM has great advantages in dealing with temporal problems. It can utilize historical data through long-term and short-term memories. Positive historical data accelerates the convergence speed of the training process and has a good guiding effect on the training of the model. However, negative historical data will hinder the training of the network model. In this paper, by optimizing the forget gate and proposing the operation of the hard sample pool, the training of specific data is enhanced, and the model performance is improved. The LSTM network structure is as shown in the appendix Figure 5 as follows. Among them,
[0102] The original LSTM unit contains three gating mechanisms (input gate i t , forget gate f t , output gate o t ) and cell state c t . Its forward propagation formula is:
[0103] f t = σ(W f · [h t-1 , x t + b f )
[0104] i t = σ(W i · [h t-1 , x t + b i )
[0105]
[0106] o t = σ(W o · [h t-1 , x t + b o )
[0107] h t = o t ☉ tanh(C t )
[0108] where ⊙ is the Hadamard product and σ is the Sigmoid function.
[0109] A loss-guided forgetting gate adaptive mechanism is proposed to address the problem that the traditional forgetting gate processes historical information without discrimination:
[0110] The input of the current batch is represented as:
[0111]
[0112] where x i represents the input at time t, represents the input weight, E represents the forgetting factor, and h t―1 represents the output at time t - 1, and b f represents the bias term of the forgetting gate. The calculation method of E is:
[0113]
[0114] where L t represents the loss value at time t, and L t―1 represents the loss value at time t - 1.
[0115] This design enables the network to focus on the current input features preferentially during the unstable training stage (loss mutation), avoiding the propagation of false memories.
[0116] During the training process, the model gradually converges, but there are always fluctuations and mutations, indicating that this part of the data is not well learned and the degree of fitting is low. These are difficult samples that need to be emphasized for learning. The loss value of the previous batch of data is added to the input of the forgetting gate, and the forgetting gate is dynamically adjusted through an algorithm strategy in combination with the loss value of the current data. When information is discarded through the Sigmoid function, the smaller the error between the current loss value and the loss value of the previous batch, the more stable the model and the less information is discarded; the larger the error, the greater the model fluctuation and the more information is discarded. The input of the current batch is as follows:
[0117]
[0118] where x i represents the input at time t, represents the input weight, E represents the forgetting factor, and h t―1 represents the output at time t - 1, and b f represents the bias term of the forgetting gate. The calculation method of E is as follows:
[0119]
[0120] where L tRepresents the loss value at time t, L t―1 Represents the loss value at time t-1.
[0121] Dynamically adjust the weights and biases of the forget gate according to the current input and previous states, enabling the network to adaptively adjust the degree of forgetting based on the loss value of the previous batch of data.
[0122] Algorithm steps for dynamically adjusting the weights and biases of the forget gate:
[0123] 1. Calculate the initial weights and biases of the forget gate:
[0124] The output of the forget gate is jointly determined by the current input and the hidden state at the previous moment, and its calculation formula is:
[0125]
[0126] where f t is the output of the forget gate, σ is the Sigmoid activation function, is the weight matrix of the forget gate, b f is the bias vector of the forget gate.
[0127] 2. Introduce the loss value to adjust the weights and biases of the forget gate:
[0128] To enable the network to adaptively adjust the degree of forgetting based on the loss value of the previous batch of data, we incorporate the loss value loss t-1 of the previous batch into the weight and bias update process of the forget gate. Specifically, we define a regulation factor E for dynamically adjusting the weights and biases of the forget gate:
[0129]
[0130] 3. Update the weights and biases of the forget gate:
[0131] In each training batch, dynamically update the weights and biases of the forget gate according to the loss value loss t of the current batch and the loss value loss t-1 of the previous batch:
[0132] W′ f =W f +E·ΔW f
[0133] b′ f =b f +E·Δb f
[0134] where ΔWf and Δbf are the gradient update amounts of the weights and biases calculated based on the current training batch.
[0135] 4. Apply the updated forget gate parameters:
[0136] Recalculate the output of the forgetting gate using the updated weight Wf′ and bias bf′ to better control the degree of information forgetting during subsequent training processes.
[0137] Ensure the stable training of the network and improve the model training speed. And establish a hard sample pool to strengthen the training of hard samples. If the training loss value of the current batch of data increases, put this batch of data into the hard sample pool and strengthen the training of hard samples subsequently to improve the adaptability of the model.
[0138] S8. After the training is completed, deploy the trained model to the on-site environment for use and fine-tune it according to the on-site environment.
[0139] Embodiment 2
[0140] This embodiment provides a perimeter sound signal recognition system based on distributed optical fiber sensing, including:
[0141] A data acquisition module, configured to
[0142] A computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the perimeter sound signal recognition method based on distributed optical fiber sensing.
[0143] A terminal device, including a processor and a computer-readable storage medium, where the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the perimeter sound signal recognition method based on distributed optical fiber sensing.
[0144] The above are all preferred embodiments of the present invention. Without restricting the protection scope of the present invention accordingly, therefore: All equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.
Claims
1. A perimeter sound signal recognition method based on distributed optical fiber sensing, characterized in that, Including: Collecting distributed fiber optic signal image data; Performing data preprocessing on the collected distributed fiber optic signal image data; Constructing a distributed fiber optic sensing network model; Using the distributed fiber optic signal image data to train the distributed fiber optic sensing network model; Deploying the trained model to the on-site environment for use and fine-tuning according to the on-site environment.
2. The perimeter sound signal recognition method based on distributed optical fiber sensing according to claim 1, characterized in that The collecting of the distributed fiber optic signal image data includes building a distributed fiber optic sound signal recognition system, and the distributed fiber optic sound signal recognition system includes a fiber optic perimeter security monitoring alarm host DAS. The fiber optic perimeter security monitoring alarm host DAS is respectively connected to a networked optical cable and a buried optical cable. Among them, first, arrange a several-shaped networked segment of about 130m; and perform 2 buried segments, bury the buried segment optical cable under 50 cm of soil, each with a length of 20m; arrange a linear networked segment, and leave 30m of coiled optical cable between different scenarios.
3. The perimeter sound signal recognition method based on distributed optical fiber sensing according to claim 2, characterized in that, The performing of data preprocessing on the collected distributed fiber optic signal image data includes preprocessing the DAS data, including data denoising and downsampling to reduce noise and unnecessary data fluctuations; then segmenting the processed data in a sliding window manner, converting it into an image and screening the image, selecting those with event signals as event sample data and those without event signals as negative samples.
4. The perimeter sound signal recognition method based on distributed optical fiber sensing according to claim 3, characterized in that, The data denoising includes using the non-local means method for denoising, classifying and weighted-averaging regions with the same nature in the same image to obtain a denoised picture, using redundant information to remove noise, which is used to filter Gaussian noise in the image. Among them, given a discrete noisy image v = {v(i)|i ∈ I}, the estimated value NL[v](i) of pixel i is calculated as the weighted average of all pixels in the image, and the filtering process of NLM is expressed as: where w represents the weight.
5. The perimeter sound signal recognition method based on distributed optical fiber sensing according to claim 4, characterized in that The performing of data preprocessing on the collected distributed fiber optic signal image data also includes, before the image data is input into the MobileNetV3 network, first adjusting the image to the expected input size of MobileNetV3, that is, 224x224 pixels, and then normalizing the image using specific mean and standard deviation values to improve the stability and convergence speed of model training.
6. The perimeter sound signal recognition method based on distributed optical fiber sensing according to claim 5, characterized in that The constructing of the distributed fiber optic sensing network model includes combining the MobileNetV3 network and the LSTM network to build a neural network framework. Among them, based on the lightweight neural network architecture of the MobileNetV3 network, introducing depthwise separable convolution, inverted residual structure and squeeze-and-excitation structure, and using the hard-swish activation function to replace the traditional activation function to improve the operation efficiency of the model and be applicable to low-computing-power hardware devices.
7. A perimeter sound signal recognition method based on distributed optical fiber sensing according to claim 6, characterized in that The model training of the distributed optical fiber sensing network using the distributed optical fiber signal image data includes inputting the preprocessed image into the MobileNetV3 network to extract event signal features. In the inverted residual structure, the number of input channels is first expanded through a 1x1 convolution, then a depthwise separable convolution is applied to reduce the computational amount and the number of parameters, and finally the number of channels is compressed through a 1x1 convolution, so as to maintain the computational efficiency while increasing the network depth.
8. A perimeter sound signal recognition method based on distributed optical fiber sensing according to claim 7, characterized in that, The model training of the distributed optical fiber sensing network using the distributed optical fiber signal image data also includes inputting the extracted event signal features into the LSTM for recognition. During the training process, the training of specific data is enhanced by optimizing the forget gate. Among them, the weights and biases of the forget gate are dynamically adjusted according to the current input and the previous state, so that the network adaptively adjusts the forgetting degree according to the loss value of the previous batch of data.
9. A perimeter sound signal recognition method based on distributed optical fiber sensing according to claim 8, characterized in that The network adaptively adjusting the forgetting degree according to the loss value of the previous batch of data includes adding the loss value of the previous batch of data to the input of the forget gate, and dynamically adjusting the forget gate through an algorithm strategy for the loss value of the current data. When information is discarded by the Sigmoid function, the smaller the error between the current loss value and the loss value of the previous batch, the more stable the model, and the less information is discarded; the larger the error, the greater the model fluctuation, and the more information is discarded. Among them, the input of the current batch is expressed as: Among them, x i represents the input at time t, represents the input weight, E represents the forgetting factor, h t―1 represents the output at time t-1, b f represents the bias term of the forgetting gate, and the calculation method of E is: Among them, L t represents the loss value at time t, and L t―1 represents the loss value at time t - 1.
10. A method for identifying perimeter sound signals based on distributed optical fiber sensing according to claim 9, characterized in that, The network adaptively adjusting the forgetting degree according to the loss value of the previous batch of data includes dynamically adjusting the weights of the forget gate through an improved algorithm strategy, so that the model flexibly controls the forgetting degree of information according to the difference between the loss values of the current batch and the previous batch. When the error between the loss values of the current batch and the previous batch is less than the set value, the model tends to be stable, and the information discarded at this time is at a low level, thus ensuring the high-fidelity recognition of the stable signal by the model; when the error is greater than the set value, the model fluctuates greatly and the information discarded at this time is at a high level, which is conducive to quickly adjusting the model state and avoiding the decline in recognition accuracy caused by overfitting.