A target detection method based on high-resolution distributed fiber optic sensing data

By constructing a spatial-frequency domain hybrid network and utilizing wavelet transform and state-dependent modules to process high-resolution fiber optic sensing data, the problems of low resolution and noise interference in traditional fiber optic sensing technology are solved, and high-precision target detection in complex scenarios is achieved.

CN120974098BActive Publication Date: 2026-03-06INST OF ADVANCED TECH UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511065727.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2026-03-06
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Traditional distributed fiber optic sensing technology is limited by low resolution and noise interference, making it difficult to achieve robust and complete target detection in complex scenarios.

Method used

A target detection method based on high-resolution distributed optical fiber sensing data is adopted. By constructing a spatial-frequency domain hybrid network, wavelet transform fusion module, spatial state dependency module and spatial-frequency domain branch fusion module are used, combined with loss function to train the model to extract and reconstruct target features.

Benefits of technology

Robust, complete, and high-precision target detection was achieved under noise interference, improving the integrity and noise resistance of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974098B_ABST
    Figure CN120974098B_ABST
Patent Text Reader

Abstract

This invention discloses a target detection method based on high-resolution distributed fiber optic sensing data. First, high-resolution raw images are acquired. Then, a target detection model based on a spatial-frequency domain hybrid network is constructed. The target detection model includes an encoder-decoder structure comprising three wavelet transform fusion modules, four spatial state dependency modules, and a spatial-frequency domain branch fusion module. The three wavelet transform fusion modules continuously downsample the high-resolution raw image, while the four spatial state dependency modules learn at different scales to model the interdependencies in long-distance space. The encoder encodes the learned features, and the decoder progressively decodes to generate the target detection result. Finally, target detection is performed based on the trained target detection model to obtain the target detection result. This invention achieves robust, complete, and high-precision target detection under complex noise interference by mining and modeling the interdependencies of target features in the global image space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed optical fiber sensor data processing technology, specifically a target detection method based on high-resolution distributed optical fiber sensor data. Background Technology

[0002] Traditional distributed fiber optic sensing methods utilize the Rayleigh scattering effect of optical signals propagating in optical fibers to monitor strain or acoustic changes caused by external dynamic interference. However, this method faces a fundamental bottleneck in practical applications: limited by the current sensing method of fiber optic technology (mainly relying on Rayleigh scattering light) and inherent noise levels (from the fiber itself and external dynamic interference), the acquired spatiotemporal data images often have low effective resolution to ensure real-time performance. At low resolution, weak linear features are easily completely submerged and covered by environmental noise, making some features difficult or even impossible to detect and extract, severely restricting its performance and reliability in complex scenarios. Current distributed fiber optic sensing technology is limited by hardware demodulation capabilities (such as pulse width and detection bandwidth) and environmental physical propagation characteristics, resulting in limitations in spatial resolution and signal strength of the output spatiotemporal data images. In open and rapidly changing scenarios, the effective strain or vibration signals generated by the target object are usually very weak and localized. When the sensing resolution is insufficient (e.g., large spatial sampling point intervals), the acquired data points are sparse and the signals are weak. These weak linear features are easily submerged in strong background noise. These noise sources are diverse and complex, including inherent noise of the fiber optic sensing system itself (such as laser intensity noise and photodetector noise), as well as unavoidable external environmental interference, such as random vibrations inherent in the system's environment and disturbances caused by environmental factors like temperature and humidity changes. In such low signal-to-noise ratio (SNR) data images, the energy of the effective target signal cannot be significantly higher than the noise background, causing key features to be completely covered by a "noise blanket," resulting in partial interruption or complete absence of the target's shape.

[0003] Traditional methods relying solely on local signal strength or simple low-resolution data processing models fall short when faced with signals completely obscured by noise. It becomes difficult to reliably determine the presence of a valid target or recover complete target morphological features based on a single or small number of contaminated data points. The advantage of high-resolution sensor data lies in its significantly expanded global spatial view and refined spatiotemporal sampling. Using higher-resolution data, a single frame can capture a wider spatial range, allowing linear targets to form sufficiently long, identifiable patterns with continuous target morphological features (such as continuous signal variation stripes) throughout the monitoring process, rather than scattered isolated points. With a high-resolution global view, even if some local points have low signal-to-noise ratios, blurred features, or temporarily disappear due to strong noise, inference and prediction can still be made using the target's large-scale continuity in the image space and its strong interdependence with preceding and following frames (temporal dimension) and adjacent points (spatial dimension). For example, prior information or contextual constraints such as the clear trajectory of the preceding feature, the patterns of other related targets in the same scene, and the typical response of that location in a low-noise region collectively constitute a powerful information redundancy network. This enables the detection to "overcome" the interference of local noise at high resolution. Based on the coherence and continuity of spatial patterns and potential physical constraints, it infers and reconstructs the existence of linear features that are masked by noise, greatly improving the integrity and robustness of target detection. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a target detection method based on high-resolution distributed optical fiber sensing data, which achieves robust, complete and high-precision target detection under complex noise interference by mining and modeling the interdependence of target features in the global image space.

[0005] The technical solution of this invention is as follows:

[0006] A target detection method based on high-resolution distributed fiber optic sensing data specifically includes the following steps:

[0007] (1) Acquire the original spatiotemporal data stream through a high-resolution distributed optical fiber sensing device, and then extract the envelope in the original spatiotemporal data stream to obtain a two-dimensional high-resolution original image, which constitutes the original sensing dataset.

[0008] (2) Construct a target detection model based on a spatial-frequency hybrid network. The target detection model includes an encoder-decoder structure consisting of three wavelet transform fusion modules, four spatial state-dependent modules, and a spatial-frequency branch fusion module.

[0009] The three wavelet transform fusion modules use discrete wavelet transform to perform three consecutive downsampling operations on the input high-resolution original image, generating three images with progressively lower resolution.

[0010] Four spatial state dependency modules learn from the input high-resolution original image and three progressively reduced resolution images at different scales, respectively, model the interdependencies in long-distance space, and output four images.

[0011] The encoder based on the spatial-frequency branch fusion module is responsible for encoding and learning features at different scales and on different branches. The decoder based on the spatial-frequency branch fusion module receives the features obtained from the encoding process in a step-by-step connection manner and gradually decodes to generate target detection results.

[0012] (3) Construct a loss function to train the target detection model based on the spatial frequency domain hybrid network to obtain a trained target detection model. Then, input the high-resolution original image to be detected into the trained target detection model for detection and obtain the target detection result.

[0013] The specific processing procedure of the wavelet transform fusion module is shown in the following equation (1):

[0014] (1);

[0015] In equation (1), The input image represents the wavelet transform fusion module. The output image represents the wavelet transform fusion module. Represents discrete wavelet transform. Represents a 3×3 convolution kernel. Represents a 1×1 convolution kernel. This represents a splicing operation. , , , These represent the four frequency components obtained after performing a discrete wavelet transform on the input image. Represents low-frequency characteristics. Represents high-frequency characteristics, Represents the inverse discrete wavelet transform;

[0016] The output image of the first wavelet transform fusion module is used as the input image of the second wavelet transform fusion module, and the output image of the second wavelet transform fusion module is used as the input image of the third wavelet transform fusion module. The input images of the four spatial state-dependent modules are the high-resolution original images. The output image of the first wavelet transform fusion module The output image of the second wavelet transform fusion module The output image of the third wavelet transform fusion module .

[0017] The specific processing procedure of the spatial state-dependent module is shown in the following formula (2):

[0018] (2);

[0019] In equation (2), The input image represents the spatial state-dependent module. The output image represents the spatial state-dependent module. Represents a 7×7 convolution kernel. Represents the Swish activation function. Represents the state-space reorganization mechanism. Represents image features, The features obtained after global learning modeling of the state-space recombination mechanism represent the characteristics. Representative group normalization, This represents a splicing operation. Represents a 1×1 convolution kernel. Represents a multi-resolution fusion module. This represents the features output after processing by the multi-resolution fusion module. This represents the LeakyReLU activation function.

[0020] The state space reorganization mechanism described above processes as follows: For the input image, it is uniformly divided into four sub-images, and then the four sub-images are fully permuted to obtain twenty-four different sequences. Each sequence is independently modeled by the discrete state space equation. Finally, the twenty-four sequences are merged and reshaped into image features of the original size by summing.

[0021] The discrete state-space equation is shown in equation (3) below:

[0022] (3);

[0023] In equation (3), Represents a discrete-time index; The matrix representing the state variables at time k; The matrix representing the state variables at time k+1; The input matrix represents time k; This is the output matrix; The system matrix represents the discrete form of the system, describing the dynamic characteristics within the state space and the influence of the current state on the state at the next moment. The discrete form of the input control matrix describes the effect of the current input on the state at the next time step. This represents the output control matrix, describing the contribution of the current state to the output; This represents the feedforward matrix, which describes the feedforward effect of the current input on the current output.

[0024] The discrete form of the system matrix The discrete form of the input control matrix is ​​obtained after discretization using equation (4). After discretization using the following formula (5), the result is obtained;

[0025] (4);

[0026] (5);

[0027] In equations (4) and (5), Represents the system matrix. Represents the input control matrix. Represents the learnable time scale parameter. Represents the identity matrix.

[0028] The processing procedure of the multi-resolution fusion module is shown in the following formula (6):

[0029] (6);

[0030] In equation (6), The input features representing the multi-resolution fusion module, This represents the output characteristics of the multi-resolution fusion module. Represents depthwise separable convolution. This represents a double downsampling operation. This represents a fourfold downsampling operation. , , Features representing three different scales. This represents a splicing operation. This represents a 1×1 convolution kernel.

[0031] The encoder's processing procedure is shown in the following formula (7):

[0032] (7);

[0033] The processing procedure of the decoder is shown in the following formula (8):

[0034] (8);

[0035] In equations (7) and (8), Represents the high-resolution original image as input. This represents the output image of the first wavelet transform fusion module. This represents the output image of the second wavelet transform fusion module. This represents the output image of the third wavelet transform fusion module. Represents the spatial and frequency domain branch fusion module. Represents a spatial state-dependent module. This indicates that downsampling is performed using pixel unshuffle. This indicates that upsampling is performed using pixel shuffle. Represents a 3×3 convolution kernel. Represents a 7×7 convolution kernel. Represents the Swish activation function. , , , These represent the output features of the four spatial-frequency-domain branch fusion modules during the encoding process. , , This represents the output characteristics of the three spatial and frequency domain branch fusion modules during the decoding process. This represents the target detection results generated through progressive decoding.

[0036] The processing procedure of the spatial-frequency domain branch fusion module is shown in the following equation (9):

[0037] (9);

[0038] In equation (9), The input features represent the spatial-frequency domain branch fusion module. This represents the output characteristics of the spatial-frequency domain branch fusion module. Represents depthwise separable convolution. Represents a 3×3 convolution kernel. Represents the Swish activation function. Represents local features in the spatial domain at the pixel level. Represents discrete wavelet transform. , , , These represent the four sub-bands at different frequencies obtained after discrete wavelet transform processing. Represents the inverse discrete wavelet transform. Represents the characteristics of wavelet frequency domain regions. Represents Fast Fourier Transform, , These represent the real and imaginary parts obtained after the Fast Fourier Transform, respectively. Represents a 1×1 convolution kernel. Represents the LeakyReLU activation function. Represents the Sigmoid activation function. and These represent the features after activation of the two pathways, Represents the inverse fast Fourier transform. Representative group normalization, Represents the global features in the Fourier frequency domain. Represents the channel attention mechanism. This represents a splicing operation. This represents element-wise multiplication.

[0039] The loss function mentioned is the MAE loss function, as shown in the following equation (10):

[0040] (10);

[0041] In equation (10), Represents the MAE loss function. This represents the predicted value of a target detection model based on a spatial-frequency domain hybrid network. Represents the actual value.

[0042] During the training process of the target detection model based on the spatial frequency domain hybrid network, at step t of period i, the learning rate is... The calculation method is shown in the following formula (11):

[0043] (11);

[0044] In equation (11), This represents the total step size of the current period i. This represents the number of steps that have been executed in the current cycle i. and These represent the minimum and initial values ​​of the learning rate, respectively. Represents a linear upper limit decay coefficient sequence;

[0045] The learning rate follows a cosine function over one period. From the initial value Smooth descent to That is, when At the end of each cycle, the learning rate is reset. The upper limit of the learning rate is based on the settings. The decay process is performed while maintaining the parameters in the target detection model, and the step size required for the next cycle is updated. , The growth factor representing the period step size, and the initial period length. and the growth coefficient of periodic compensation and linear upper limit decay sequence As a hyperparameter setting.

[0046] Advantages of this invention:

[0047] (1) The wavelet transform fusion module WTF in the target detection model of the present invention uses the near lossless discrete wavelet transform DWT to replace the traditional step convolution for downsampling, thereby explicitly separating and fusing the low-frequency features (carrying global structure) and high-frequency features (containing directional details) of the image, effectively avoiding the loss of key spatial information in the downsampling process, and laying a more robust frequency domain foundation for the subsequent preservation and extraction of weak vehicle trajectory features under noise interference. In addition, three wavelet transform fusion modules are used to perform three consecutive downsamplings to generate three images with progressively reduced resolution as multi-scale inputs.

[0048] (2) The spatial state dependency module SSD in the target detection model of the present invention integrates the key state space recombination mechanism SSR and resolution fusion module MRF. SSR generates multiple scanning path sequences by dividing the input image into multiple sub-images and performing full permutations, and applies discrete state space equations to independently model each path. This mechanism can efficiently capture long-distance spatial context dependencies in the image, so that even if the trajectory in some areas is completely submerged by noise, the model can infer and reconstruct the trajectory segments covered by noise based on the continuity of the trajectory in time and space (previous and subsequent time points, adjacent spatial positions) and global structural information, which significantly improves the integrity of trajectory extraction and noise resistance. The resolution fusion module MRF learns the features obtained after SSR processing at different resolutions, effectively aggregating features from different scales while maintaining computational efficiency to obtain richer spatial information, so that the features modeled by SSR can be propagated more efficiently in the model.

[0049] (3) The encoder-decoder structure of the present invention is constructed based on the spatial-frequency branch fusion module. The spatial-frequency branch fusion module SSBF contains three complementary branches: pixel-level spatial domain branch (extracting local spatial features), wavelet frequency domain branch (extracting regional frequency features and reconstructing), and Fourier frequency domain branch (modeling global frequency domain context). These three branches focus on local, regional and global information respectively, and are deeply fused under the guidance of the channel attention mechanism, so that the model can simultaneously utilize spatial details and frequency domain information of different frequency ranges (local-regional-global), breaking through the limitations of modeling in a single domain or a single receptive field.

[0050] (4) In the model training process, the present invention uses random initialization to assign initial values ​​to the parameters of each layer of the model, and proposes a scheduling algorithm that combines cosine annealing, periodic restart, linear upper limit decay and gradient clipping to help the model jump out of local optima and thus improve its generalization ability. After processing with the above scheduling algorithm, the learning rate will decrease at a relatively fast speed in the early stage of training to ensure that the model can converge quickly in the early stage of training, that is, the warm-up stage. Then, the periodic restart mechanism causes the learning rate to suddenly increase, which helps the model jump out of the possible local optima. The magnitude of the sudden increase in the learning rate can be controlled by the linear upper limit decay sequence. As the training cycle increases, the model completes the early rapid exploration. The training step size of each cycle will double, so that the model can be finely tuned with a longer cycle when entering the later stage. Attached Figure Description

[0051] Figure 1 This is a network framework diagram of the target detection model based on a spatial-frequency domain hybrid network according to the present invention.

[0052] Figure 2 This is a network framework diagram of the wavelet transform fusion module of the present invention.

[0053] Figure 3 This is the network framework diagram of the spatial state-dependent module of the present invention.

[0054] Figure 4 This is a network framework diagram of the multi-resolution fusion module of the present invention.

[0055] Figure 5 This is a network framework diagram of the spatial and frequency domain branch fusion module of the present invention.

[0056] Figure 6 This is a network framework diagram of pixel-level spatial domain branching in the spatial domain-frequency domain branching fusion module of the present invention.

[0057] Figure 7 This is a network framework diagram of the wavelet frequency domain branch in the spatial-frequency domain branch fusion module of the present invention.

[0058] Figure 8 This is a network framework diagram of the Fourier frequency domain branch in the spatial-frequency domain branch fusion module of the present invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] A target detection method based on high-resolution distributed fiber optic sensing data specifically includes the following steps:

[0061] (1) Taking the high-resolution distributed optical fiber sensing equipment for vehicle detection on highways as an example: vehicle detection is carried out by communication optical cables deployed in the shallow underground of the median strip of the highway. The communication optical fiber is eight kilometers long. The data is divided into 16,000 high-density monitoring points along the length of the optical fiber, with a spacing of 0.5 meters between each monitoring point, thereby realizing distributed fine sensing of road vibration response and obtaining high-resolution data.

[0062] This embodiment relies on readily available high-resolution distributed fiber optic sensing equipment. Its core working mechanism is based on the principle of phase-sensitive optical time-domain reflectometry (Φ-OTDR). This equipment integrates a highly coherent pulsed light source and a high-sensitivity optical detector. At the end of the sensing fiber, a fiber coupler is used to form a closed-loop optical path system that includes light pulse emission and backscattered signal reception. When the pulsed light propagates in the fiber, the inherent Rayleigh scattering effect of the fiber material produces back-propagating scattered light. This scattered light optically interferes with the original incident pulse. The detector accurately converts the captured interference light signal into an electrical signal, which is then transmitted through a data interface. The data is transmitted to the main control computer. The pressure from passing vehicles causes micro-vibrations on the ground, which are then transmitted to the laid optical fibers, causing slight deformation of the fibers. This deformation directly modulates the physical parameters of the local area of ​​the optical fiber, including its refractive index and axial stress state. This modulation effect causes a nonlinear shift in the phase and intensity of the backscattered Rayleigh light, which is reflected in the baseband waveform of the interference signal finally acquired by the detector as a distinctive feature distortion. The entire high-resolution distributed optical fiber sensing equipment hardware is installed in a standard communication cabinet along the highway and operates at a high-performance 2kHz sampling frequency to ensure real-time capture of rapidly changing vibrations.

[0063] The aforementioned high-resolution distributed fiber optic sensing equipment is used to measure road surface vibration signals in real time. The acquired raw spatiotemporal data stream (including target vibration signals caused by vehicle movement and various background noises) is processed using Hilbert transform to extract the envelope, obtaining a two-dimensional high-resolution raw image, which constitutes the raw sensing dataset. Labels are obtained through manual annotation, thus yielding vehicle trajectory labels. That is, the true value, with 70% of the original sensor dataset used as the training set and 30% as the validation set;

[0064] (2) See Figure 1 A target detection model based on a spatial-frequency hybrid network is constructed. The target detection model includes an encoder-decoder structure with three wavelet transform fusion modules (WTF), four spatial state-dependent modules (SSD), and a spatial-frequency branch fusion module (SSBF).

[0065] S21, see Figure 2 The specific processing procedure of the wavelet transform fusion module WTF is shown in the following equation (1):

[0066] (1);

[0067] In equation (1), The input image represents the wavelet transform fusion module. The output image represents the wavelet transform fusion module. Represents discrete wavelet transform. Represents a 3×3 convolution kernel. Represents a 1×1 convolution kernel. This represents a splicing operation. , , , The discrete wavelet transform of the input image yields four frequency components, each with a spatial resolution half that of the original input. It represents low-frequency features, allowing the image to retain its global structure. It represents high-frequency features and captures detailed information in the horizontal, vertical, and diagonal directions. Represents the inverse discrete wavelet transform;

[0068] Three wavelet transform fusion modules process the input high-resolution original image (dimension 1). The system performs three consecutive downsampling operations to generate three progressively lower resolution images. The output image of the first wavelet transform fusion module (dimension 1) is... The input image of the second wavelet transform fusion module is used as the input image, and the output image of the second wavelet transform fusion module is (dimension: ...). The input image for the third wavelet transform fusion module is the high-resolution original image, and the input images for the four spatial state-dependent modules are the high-resolution original images, respectively. The output image of the first wavelet transform fusion module The output image of the second wavelet transform fusion module The output image of the third wavelet transform fusion module (dimension is) );

[0069] S22, see Figure 3 The specific processing procedure of the SSD space state-dependent module is shown in the following formula (2):

[0070] (2);

[0071] In equation (2), The input image represents the spatial state-dependent module. The output image represents the spatial state-dependent module. Represents a 7×7 convolution kernel. Represents the Swish activation function. Represents the state-space reorganization mechanism. Represents image features, The features obtained after global learning modeling of the state-space recombination mechanism represent the characteristics. Representative group normalization, This represents a splicing operation. Represents a 1×1 convolution kernel. Represents a multi-resolution fusion module. This represents the features output after processing by the multi-resolution fusion module. Represents the LeakyReLU activation function;

[0072] S221. The processing of the State Space Reorganization Mechanism (SSR) is as follows: For the input image, it is uniformly divided into four sub-images, and then the four sub-images are fully permuted to obtain twenty-four different sequences. The twenty-four different sequences obtained by the full permutation are regarded as scanning paths of the image from different directions and at different angles. Each sequence is independently modeled by the discrete state space equation. Finally, the twenty-four sequences are merged and reshaped into image features of the original size by summing.

[0073] The discrete state-space equation is shown in equation (3) below:

[0074] (3);

[0075] In equation (3), Represents a discrete-time index; The matrix representing the state variables at time k; The matrix representing the state variables at time k+1; The input matrix represents time k; This is the output matrix; The system matrix represents the discrete form of the system, describing the dynamic characteristics within the state space and the influence of the current state on the state at the next moment. The discrete form of the input control matrix describes the effect of the current input on the state at the next time step. This represents the output control matrix, describing the contribution of the current state to the output; This represents the feedforward matrix, which describes the feedforward effect of the current input on the current output.

[0076] Discrete form of system matrix After discretization using equation (4), the discrete form of the input control matrix is ​​obtained. After discretization using the following formula (5), the result is obtained;

[0077] (4);

[0078] (5);

[0079] In equations (4) and (5), Represents the system matrix. Represents the input control matrix. Represents the learnable time scale parameter. Represents the identity matrix;

[0080] S222, see Figure 4 Multi-resolution fusion module The processing procedure is shown in the following formula (6):

[0081] (6);

[0082] In equation (6), The input features representing the multi-resolution fusion module, This represents the output characteristics of the multi-resolution fusion module. Represents depthwise separable convolution. This represents a double downsampling operation. This represents a fourfold downsampling operation. , , The features represent three different scales (preserving resolution, downsampling by 2x, and downsampling by 4x, and then processed by depthwise separable convolution). This represents a splicing operation. Represents a 1×1 convolution kernel;

[0083] S23. The encoder based on the spatial-frequency branch fusion module is responsible for encoding and learning features at different scales and on different branches. The decoder based on the spatial-frequency branch fusion module receives the features obtained in the encoding process in a step-by-step connection manner and gradually decodes to generate target detection results.

[0084] S231. The encoder processing procedure is shown in the following formula (7):

[0085] (7);

[0086] S232. The decoding process is shown in the following formula (8):

[0087] (8);

[0088] In equations (7) and (8), Represents the high-resolution original image as input. This represents the output image of the first wavelet transform fusion module. This represents the output image of the second wavelet transform fusion module. This represents the output image of the third wavelet transform fusion module. Represents the spatial and frequency domain branch fusion module. Represents a spatial state-dependent module. This indicates that downsampling is performed using pixel unshuffle. This indicates that upsampling is performed using pixel shuffle. Represents a 3×3 convolution kernel. Represents a 7×7 convolution kernel. Represents the Swish activation function. , , , These represent the output features of the four spatial-frequency-domain branch fusion modules during the encoding process. , , This represents the output characteristics of the three spatial and frequency domain branch fusion modules during the decoding process. Represents the target detection results generated by progressive decoding (dimension: );

[0089] See Figures 5-8 The spatial-frequency branch fusion module SSBF contains three complementary branches: pixel-level spatial domain branch, wavelet frequency domain branch, and Fourier frequency domain branch. The processing procedure is shown in the following equation (9):

[0090] (9);

[0091] In equation (9), The input features represent the spatial-frequency domain branch fusion module. This represents the output characteristics of the spatial-frequency domain branch fusion module. Represents depthwise separable convolution. Represents a 3×3 convolution kernel. Represents the Swish activation function. Represents local features in the spatial domain at the pixel level. Represents discrete wavelet transform. , , , These represent the four sub-bands at different frequencies obtained after discrete wavelet transform processing. Represents the inverse discrete wavelet transform. Represents the characteristics of wavelet frequency domain regions. Represents Fast Fourier Transform, , These represent the real and imaginary parts obtained after the Fast Fourier Transform, respectively. Represents a 1×1 convolution kernel. Represents the LeakyReLU activation function. Represents the Sigmoid activation function. and These represent the features after activation of the two pathways, Represents the inverse fast Fourier transform. Representative group normalization, Represents the global features in the Fourier frequency domain. Channel Attention is a representative channel attention mechanism. This represents a splicing operation. This represents element-wise multiplication;

[0092] (3) Construct the MAE loss function (see Equation 10 below) to train the target detection model based on the spatial frequency domain hybrid network, and then input the high-resolution original image to be detected into the trained target detection model for detection to obtain the target detection result.

[0093] (10);

[0094] In equation (10), Represents the MAE loss function. This represents the predicted value of a target detection model based on a spatial-frequency domain hybrid network. Represents the true value;

[0095] During the training process of the target detection model based on the spatial frequency domain hybrid network, random initialization is used to assign initial values ​​to the parameters of each layer of the model, and a scheduling algorithm is constructed to help the model escape local optima and thus improve its generalization ability.

[0096] Specifically: During the training process, at step t of period i, the learning rate... The calculation method is shown in the following formula (11):

[0097] (11);

[0098] In equation (11), This represents the total step size of the current period i. This represents the number of steps that have been executed in the current cycle i. and These represent the minimum and initial values ​​of the learning rate, respectively. Represents a linear upper limit decay coefficient sequence;

[0099] The learning rate follows a cosine function over one period. From the initial value Smooth descent to That is, when At the end of each cycle, the learning rate is reset. The upper limit of the learning rate is based on the settings. The decay process is performed while maintaining the parameters in the target detection model, and the step size required for the next cycle is updated. , The growth factor representing the period step size, and the initial period length. and the growth coefficient of periodic compensation and linear upper limit decay sequence As a hyperparameter setting;

[0100] After each training round, the object detection model is evaluated using a validation set, and performance metrics such as mIoU and SSIM are calculated. The hyperparameters are dynamically adjusted based on the model's performance on the validation set to improve its generalization ability. After the object detection model is trained, the model parameters are saved for use in detecting and extracting target-vehicle trajectory information on noisy images at higher resolutions.

[0101] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A target detection method based on high-resolution distributed fiber sensing data, characterized in that: Specifically comprising the following steps: (1) Collecting original space-time data stream through high-resolution distributed optical fiber sensing equipment, then extracting envelope in the original space-time data stream, obtaining two-dimensional high-resolution original image, and constituting original sensing data set; (2) Constructing target detection model based on spatial frequency domain hybrid network, the target detection model comprising three wavelet transform fusion modules, four spatial state dependent modules and encoder-decoder structure based on space domain frequency branch fusion module; The three wavelet transform fusion modules adopt discrete wavelet transform to continuously downsample the input high-resolution original image three times, generating three images with gradually reduced resolution; The four spatial state dependent modules learn the input high-resolution original image and the three images with gradually reduced resolution on different scales, model the mutual dependence relationship on long-distance space, and output four images; The encoder based on the space domain frequency branch fusion module is responsible for encoding and learning features on different scales and different branches, and the decoder based on the space domain frequency branch fusion module receives the features obtained in the encoding process in a cross-step connection manner, and gradually decodes to generate the target detection result; The processing process of the encoder is shown in formula (7): (7); The processing process of the decoder is shown in formula (8): (8); in formula (7) and formula (8), represent the input high-resolution raw image, represent the output image of the first wavelet transform fusion module, represent the output image of the second wavelet transform fusion module, represent the output image of the third wavelet transform fusion module, represent the spatial-frequency branch fusion module, represent the spatial state dependent module, represent down-sampling using pixel inverse rearrangement Pixel UnShuffle, represent up-sampling using pixel rearrangement PixelShuffle, represent 3x3 convolution kernel, represent 7x7 convolution kernel, represent Swish activation function, 、 、 、 respectively represent the output features of the four spatial-frequency branch fusion modules in the encoding process, 、 、 represent the output features of the three spatial-frequency branch fusion modules in the decoding process, represent the target detection result generated by step-by-step decoding; (3) Constructing loss function to train the target detection model based on the spatial frequency domain hybrid network, obtaining the trained target detection model, then inputting the high-resolution original image to be detected into the trained target detection model for detection, and obtaining the target detection result.

2. The target detection method based on high-resolution distributed optical fiber sensing data according to claim 1, characterized in that: The processing process of the wavelet transform fusion module is shown in formula (1): (1); In formula (1), representing the input image of the wavelet transform fusion module, representing the output image of the wavelet transform fusion module, representing the discrete wavelet transform, representing a 3x3 convolution kernel, representing a 1x1 convolution kernel, representing a concatenation operation, , , , representing the four frequency components obtained after the discrete wavelet transform of the input image, representing low-frequency features, representing high-frequency features, representing the inverse discrete wavelet transform; The output image of the first wavelet transform fusion module is taken as the input image of the second wavelet transform fusion module, the output image of the second wavelet transform fusion module is taken as the input image of the third wavelet transform fusion module, and the input images of the four spatial state-dependent modules are the input high-resolution original image , the output image of the first wavelet transform fusion module , the output image of the second wavelet transform fusion module , and the output image of the third wavelet transform fusion module .

3. The target detection method based on high-resolution distributed fiber sensing data according to claim 1, characterized in that: The processing process of the spatial state dependent module is shown in formula (2): (2); in formula (2), representing an input image of a spatial state-dependent module, representing an output image of a spatial state-dependent module, representing a 7x7 convolution kernel, representing a Swish activation function, representing a state space reorganization mechanism, representing an image feature, representing a feature obtained after global learning modeling of the state space reorganization mechanism, representing group normalization, representing a splicing operation, representing a 1x1 convolution kernel, representing a multi-resolution fusion module, representing a feature output after processing of the multi-resolution fusion module, representing a LeakyReLU activation function.

4. The target detection method based on high-resolution distributed optical fiber sensing data according to claim 3, characterized in that: The processing process of the state space recombination mechanism is that the input image is uniformly divided into four subgraphs, then the four subgraphs are fully arranged to obtain twenty-four different sequences, each sequence is independently modeled by a discrete state space equation, and finally the twenty-four sequences are combined and restored to the original size of the image features by summation; The discrete state space equation is shown in formula (3): (3); in equation (3), represents a discrete time index; represents a state variable matrix at time k; represents a state variable matrix at time k+1; represents an input matrix at time k; is an output matrix; represents a system matrix in discrete form, describing the internal dynamics of the state space and the influence of the current state on the next time state; represents an input control matrix in discrete form, describing the influence of the current input on the next time state; represents an output control matrix, describing the contribution of the current state to the output; represents a feedforward matrix, describing the feedforward effect of the current input on the current output; The system matrix in the discrete form is obtained by discretization processing through the following formula (4); The input control matrix in the discrete form is obtained by discretization processing through the following formula (5); (4); (5); In formula (4) and formula (5), denotes a system matrix, denotes an input control matrix, denotes a learnable time scale parameter, denotes an identity matrix.

5. The target detection method based on high-resolution distributed fiber sensing data according to claim 3, characterized in that: The processing process of the multi-resolution fusion module is shown in formula (6): (6); in formula (6), input features representing the multi-resolution fusion module, output features representing the multi-resolution fusion module, representing a depthwise separable convolution, representing a two times down-sampling operation, representing a four times down-sampling operation, 、 、 features representing three different scales, representing a concatenation operation, representing a 1 x 1 convolution kernel.

6. The target detection method based on high-resolution distributed fiber sensing data according to claim 1, characterized in that: The processing process of the space domain frequency branch fusion module is shown in formula (9): (9); in formula (9), represent input features of the spatial-frequency domain branch fusion module, represent output features of the spatial-frequency domain branch fusion module, represent depth separable convolution, represent 3x3 convolution kernels, represent Swish activation function, represent pixel-level spatial domain local features, represent discrete wavelet transform, 、 、 、 represent four subbands of different frequencies obtained after discrete wavelet transform processing, represent inverse discrete wavelet transform, represent wavelet frequency domain region features, represent fast Fourier transform, 、 represent real and imaginary parts obtained after fast Fourier transform, represent 1x1 convolution kernels, represent LeakyReLU activation function, represent Sigmoid activation function, and represent two activated features, represent inverse fast Fourier transform, represent group normalization, represent Fourier frequency domain global features, represent channel attention mechanism, represent concatenation operation, represent element-wise multiplication.

7. The target detection method based on high-resolution distributed fiber sensing data according to claim 1, characterized in that: The loss function is MAE loss function, and is shown in formula (10): (10); In formula (10), represents the MAE loss function, represents the predicted value of the target detection model based on the spatial frequency domain hybrid network, represents the true value.

8. The target detection method based on high-resolution distributed fiber sensing data according to claim 7, characterized in that: In the training process of the target detection model based on the spatial frequency domain hybrid network, at the tth step of the i th period, the learning rate The calculation method of the learning rate is shown in the following formula (11): (11); In formula (11), represents the total step length of the current cycle i, represents the number of steps that have been performed in the current cycle i, and respectively represent the minimum value and the initial value of the learning rate, represents a linear upper limit decay coefficient sequence; The learning rate follows a cosine function over one period. From the initial value Smooth descent to That is, when At the end of a cycle, the learning rate is reset. The upper limit of the learning rate depends on the settings. The decay process is performed while maintaining the parameters in the target detection model, and the step size required for the next cycle is updated. , The growth factor representing the period step size, and the initial period length. and the growth coefficient of periodic compensation and linear upper limit decay sequence As a hyperparameter setting.

Citation Information

Patent Citations

  • Infrared weak and small target detection method based on wavelet guidance

    CN118587507A

  • Multi-domain unmanned aerial vehicle infrared image super-resolution dividing and conquering method based on Mama

    CN120374383A