Distributed acoustic sensing signal processing method and system based on FPGA, medium and product
By employing an FPGA-based distributed acoustic sensing signal processing method, utilizing complex coefficient partitioned storage and time-division multiplexing operation units, combined with pulse repetition period timing scheduling and delay compensation, the resource waste and real-time issues of centralized processing architecture are resolved. This achieves efficient filtering and timing alignment of multi-band signals, improving the system's processing efficiency and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZHONGTUO XINYUAN TECH CO LTD
- Filing Date
- 2026-04-13
- Publication Date
- 2026-05-12
AI Technical Summary
Existing distributed acoustic sensing signal processing methods suffer from problems such as excessive processing unit burden, waste of storage resources, and difficulty in guaranteeing the real-time performance and accuracy of signal processing due to centralized processing architecture.
A distributed acoustic sensing signal processing method based on FPGA is adopted. By storing complex coefficients in partitions and using the same set of computing units in a time-division multiplexing manner across multiple frequency bands, combined with strict timing scheduling of pulse repetition cycles and a hierarchical parallel accumulation structure, it is ensured that each frequency band completes filtering within a specified time window, and the timing alignment of signal output is achieved through a delay compensation mechanism.
It significantly saves digital signal processing and storage resources, supports multi-band complex coefficient finite-length unit impulse response filtering, meets the real-time processing requirements under high sampling rates, improves system robustness and adaptability, and ensures phase consistency and frame integrity of multi-band data output.
Smart Images

Figure CN122018849A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal processing, and in particular to a distributed acoustic sensing signal processing method, system, medium, and product based on FPGA. Background Technology
[0002] With the rapid development of information technology, distributed acoustic sensing technology has demonstrated enormous application potential in numerous fields. It enables real-time monitoring and analysis of acoustic signals over large areas, playing a crucial role in fields such as oil and gas exploration, perimeter security, and structural health monitoring. Distributed acoustic sensing signal processing technology, as its core component, directly impacts the performance and application effectiveness of the entire system. Efficient and accurate signal processing methods can extract valuable information from complex acoustic signals, providing reliable data for decision-making in various fields.
[0003] In existing technologies, the processing of distributed acoustic sensing signals typically employs a traditional centralized processing architecture. This architecture concentrates all signal processing tasks in a single processing unit. When processing signals across multiple frequency bands, it often requires allocating processing resources separately for each band and storing and processing the coefficients of all bands uniformly.
[0004] Existing distributed acoustic sensing signal processing methods have significant drawbacks. Centralized processing architectures overburden processing units, failing to meet the demands of large-scale signal processing. The unified storage and processing of frequency band coefficients fails to fully utilize the characteristics of different types of memory, resulting in wasted storage resources. Furthermore, the lack of effective processing timing scheduling mechanisms makes it difficult to guarantee the real-time performance and accuracy of signal processing, thus failing to meet the requirements of efficient and precise signal processing in practical applications. Summary of the Invention
[0005] This application provides a distributed acoustic sensing signal processing method, system, medium, and product based on FPGA, which effectively utilizes FPGA resources, improves processing efficiency, achieves efficient filtering processing of multi-band signals, and ensures timing alignment of output signals.
[0006] In a first aspect, this application provides a distributed acoustic sensing signal processing method based on FPGA, the method comprising: Acquire the distributed acoustic sensor signal to be processed, and generate the corresponding valid data signal and frame end signal based on the pulse repetition period of the distributed acoustic sensor signal; Based on frequency band priority, the preset order complex coefficients of multiple frequency bands are partitioned and configured in the on-chip heterogeneous memory of the FPGA. The priority is determined based on the signal energy ratio of each frequency band. The preset order complex coefficients of the first frequency band with the highest signal energy ratio are stored in the first type of memory, and the preset order complex coefficients of the second frequency band other than the first frequency band are stored in the second type of memory. The access speed of the first type of memory is higher than that of the second type of memory. Based on the pulse repetition period, the system processing timing is determined. According to the timing and the storage location of the preset order complex coefficient partition configuration, the finite-length unit impulse response filtering operations of the multiple frequency bands are scheduled to the same group of complex multiplication operation units in a time-division manner. This ensures that within the operation time window corresponding to the target frequency band, the preset order complex coefficients are read from the corresponding first type memory or second type memory according to the storage location of the target frequency band and input to the complex multiplication operation unit. The target frequency band is any one of the multiple frequency bands. Within the computation time window, using the preset order complex coefficients input to the complex multiplication unit, complex multiplication is performed on the input signal to obtain multiple multiplication accumulation results. Then, multi-level parallel accumulation is performed on the multiple multiplication accumulation results. Each accumulation stage pairs and sums the results output from the previous stage to compress the number of nodes until the final filtering result of the target frequency band is obtained. The data validity signal and the frame end signal are delayed by a delay compensation mechanism to generate a control signal that is time-aligned with the output of the final filtered result, and the final output is controlled based on the pulse repetition period and the control signal.
[0007] By employing the above technical solution, through partitioned storage of complex coefficients and time-division multiplexing of the same set of computing units across multiple frequency bands, digital signal processing and storage resources are significantly saved, supporting the implementation of multi-band complex coefficient finite-length unit impulse response (FIR) filtering on a single FPGA. Based on strict timing scheduling of the pulse repetition period and a hierarchical parallel accumulation structure, filtering of each frequency band is ensured within a specified time window, with controllable system latency, meeting the real-time processing requirements of distributed acoustic sensing (DAS) at high sampling rates. A configurable delay chain is used to synchronously compensate for the valid data signal and the frame end signal, ensuring strict alignment between the control signal and the filtering result, guaranteeing phase consistency and frame integrity of multi-band data output. Dynamic allocation of storage and computing resources based on signal energy distribution supports priority scheduling of different frequency bands, adapting to the time-varying characteristics of DAS signals in complex monitoring scenarios, and improving the overall robustness and adaptability of the system.
[0008] In some embodiments, determining the system processing timing based on the pulse repetition period includes: Based on the system sampling rate and the pulse repetition period, the time length of a single pulse repetition period and the corresponding total number of sampling points are determined. The time length is divided according to the preset calculation weights of the multiple frequency bands, and a continuous time segment is assigned to each frequency band. All time segments are connected end to end and their sum is equal to the pulse repetition period. Establish a mapping relationship between system clock counting and frequency band operation, so that when the system clock count falls within the time segment allocated by the third frequency band, the filtering operation of the third frequency band is automatically triggered. The third frequency band is any one of the multiple frequency bands. The system processing timing is determined based on the time segment and the mapping relationship, so that the complex multiplication unit processes a signal of one frequency band at each moment, and each frequency band is processed in sequence.
[0009] By employing the above technical solution, filtering operations across multiple frequency bands are time-divisionally mapped to the same group of complex multiplication units, avoiding the need for independent deployment of computing hardware for each frequency band. This allows for parallel processing of multiple frequency bands without increasing DSP resources, significantly improving hardware utilization. Precise time window division based on the pulse repetition period and system sampling rate ensures that all frequency bands complete a full cycle of processing within one pulse repetition period, without data backlog or frame loss, meeting the high requirements of DAS systems for timing predictability. By pre-setting computation weights to allocate time segments, the computation duration can be dynamically adjusted according to the importance of different frequency bands or the amount of data processed, achieving differentiated resource allocation and enhancing the system's adaptability to changing scenarios. A deterministic mapping relationship is established between system clock counting and frequency band operations; corresponding frequency band operations can be automatically triggered simply by counting comparisons, eliminating the need for complex arbitration or synchronization circuits, reducing control logic resource consumption and timing design complexity.
[0010] In some embodiments, the plurality of multiplicative accumulation results include a preset number of real part multiplication results and a preset number of imaginary part multiplication results, and the multi-level hierarchical parallel accumulation of the plurality of multiplicative accumulation results includes: The product results of the real parts of the preset order are paired in order to form real part node pairs, and the remaining single product results of the real parts are recorded as real part remainders. Based on the real part node pairs and the real part remainders, a set of real part input nodes is constructed. The product results of the imaginary parts of the preset order are paired in order to form imaginary part node pairs, and the remaining single product results of the imaginary parts are recorded as imaginary part remainders. Based on the imaginary part node pairs and the imaginary part remainders, a set of imaginary part input nodes is constructed. A preset number of parallel adders are used to perform synchronous summation on all pairs of real part nodes in the real part input node set to generate a real part compression result set, and the real part remainder is added to the real part compression result set. A preset number of parallel adders are used to perform synchronous summation on all pairs of imaginary part nodes in the imaginary part node set to generate an imaginary part compression result set, and the imaginary part remainder is added to the imaginary part compression result set. Use the real part compression result set as the new real part input node set, and the imaginary part compression result set as the new imaginary part input node set. Repeat the previous step until the real part nodes are compressed into a single real part result and the imaginary part nodes are compressed into a single imaginary part result. The single real part result and the single imaginary part result are output as the final real part filtering result and the final imaginary part filtering result of the target frequency band, respectively.
[0011] By employing the above technical solution, a parallel accumulation structure of "pairing together and compressing step by step" reduces the multiple additions required for serial accumulation to only a few pipeline stages, significantly shortening the accumulation path delay of FIR filtering and improving the system's real-time response capability. Each stage uses a preset number of parallel adders, with the number of adders decreasing progressively with each stage, avoiding resource waste caused by deploying a large number of adders at once. This achieves an optimized balance between speed and area, making it particularly suitable for the rational allocation of DSP resources in FPGAs. A strategy of retaining residual terms and incorporating them into the next round of compression ensures that all product results participate in the final accumulation without data loss. Simultaneously, the step-by-step compression structure facilitates insertion into pipeline registers, controlling data overflow and timing deviations during accumulation. It is not dependent on a fixed order and can adapt to different FIR order requirements by adjusting the number of stages and the number of adders per stage, possessing excellent parameterization and reconfigurability characteristics, facilitating flexible deployment in different DAS application scenarios.
[0012] In some embodiments, the step of delaying the valid data signal and the end-of-frame signal through a delay compensation mechanism to generate a control signal that is time-aligned with the output of the final filtered result specifically includes: Based on the total number of delay levels accumulated in parallel at multiple levels, a data validity delay chain and a frame end signal delay chain are constructed respectively. Both the data validity delay chain and the frame end signal delay chain include the total number of cascaded triggers. The data valid signal is input into the data valid delay chain and passed through each stage for a total delay level of several clock cycles to generate a regenerated synchronous data valid signal. The frame end signal is input into the frame end signal delay chain and passed through each stage for a total delay level of several clock cycles to generate a regenerated synchronous frame end signal. During the effective period when the regenerated synchronization data valid signal is high, the real and imaginary parts of the final filtering result are continuously sampled to obtain sampled data. By verifying the phase continuity of the sampled data during the effective period, it is possible to check whether the phase alignment accuracy between the regenerated synchronization data valid signal and the final filtering result meets the requirements. When the phase alignment accuracy meets the requirements, the valid signal of the regenerated synchronization data and the end signal of the regenerated synchronization frame are used as control signals and output synchronously with the final filtering result.
[0013] By employing the above technical solution, a delay chain that is strictly matched to the depth of the accumulation pipeline is constructed, ensuring that the delay levels of the valid data signal and the frame end signal are completely consistent with the total delay of the filtering operation. This guarantees precise alignment between the control edge and the data output time, fundamentally preventing timing misalignment. The delay-compensated frame end signal accurately identifies the boundary of the valid data segment within each pulse cycle, preventing cross-frame data confusion or truncation caused by filtering pipeline delays, and ensuring the independence and integrity of each frame's data. By actively verifying the phase continuity of the output data within the valid time period, synchronization accuracy can be monitored in real time, and alignment can be determined to meet system requirements, providing a built-in verification method for system debugging, online monitoring, and fault diagnosis. The deterministic delay method using cascaded triggers is less affected by factors such as process technology, voltage, and temperature, resulting in stable and predictable output timing. Simultaneously, the alignment verification mechanism ensures correct system operation even under slight timing fluctuations.
[0014] In some embodiments, verifying the phase continuity of the sampled data within the effective time period to check whether the phase alignment accuracy between the regenerated synchronization data effective signal and the final filtering result meets the requirements specifically includes: The target adder with the largest delay in the multi-level parallel accumulation is selected as the monitoring point, and the stable establishment time margin of the output signal of the target adder from data validity to sampling by the next level register is monitored in real time. When the stable setup time margin is lower than the dynamic safety threshold, the number of delay clock cycles that need to be additionally compensated is calculated based on the difference between the stable setup time margin and the dynamic safety threshold, as well as the system clock cycle, and is used as the delay compensation increment. Based on the delay compensation increment, by controlling the data selector set at the input of each stage of the trigger in the data validity delay chain and the frame end signal delay chain, the triggers at a preset level are bypassed or connected to adjust the effective depth of the data validity delay chain and the frame end signal delay chain. After adjusting the effective depth, the step of restoring the verification phase continuity is performed.
[0015] By employing the above technical solution and selecting the critical path (maximum delay adder) to monitor the setup time margin, the system can detect timing drift caused by environmental changes or hardware aging in real time. Based on the difference, it calculates compensation and achieves closed-loop adaptive adjustment, significantly improving the system's adaptability to timing fluctuations. When the setup time margin falls below the dynamic safety threshold, a compensation mechanism is automatically triggered, effectively preventing data sampling errors or system failures caused by timing violations. This is particularly suitable for long-term stable operation under complex conditions such as large industrial temperature variations and power supply fluctuations. Through dynamic bypassing or triggering via the data selector, the delay chain depth can be programmably adjusted, avoiding resource waste or insufficient compensation caused by using a fixed-depth delay chain, allowing delay resources to be precisely configured according to actual needs.
[0016] In some embodiments, the step of partitioning and configuring the complex coefficients of multiple frequency bands of a preset order based on frequency band priority in the on-chip heterogeneous memory of the FPGA includes: By using a preset weighted scoring function, the historical average signal energy of each frequency band, the frequency of signal mutations within a preset historical period, and the preset system-level importance weights are weighted and summed to generate a comprehensive priority score. All frequency bands are sorted in descending order according to the comprehensive priority score. The upper limit of the number of frequency bands that the first type of memory can accommodate is calculated according to the ratio of the total capacity of the first type of memory to the amount of complex coefficient data in a single frequency band. The first frequency band is defined according to the upper limit of the number of frequency bands. The preset order complex coefficients of the first frequency band are preloaded into the first type of memory. The preset order complex coefficients of the remaining frequency bands are stored in the second type of memory. The first frequency band in the first type of memory is sorted a second time according to the comprehensive priority score, and stored in the memory blocks with increasing access latency in descending order of comprehensive priority score; A mapping table of frequency band ID, memory type, and storage block address is generated and written to the configuration memory of the FPGA. When the system is powered on and initialized, the preset order complex coefficients of each frequency band are loaded from external storage to the corresponding heterogeneous memory partition on the FPGA chip according to the mapping table.
[0017] By employing the above technical solution, priority is dynamically calculated by combining multi-dimensional parameters such as signal energy, mutation frequency, and system importance, and high-frequency bands are allocated to high-speed memory accordingly. This ensures minimal access latency for coefficients in critical frequency bands, optimizing storage access performance at the system level. The number of high-priority frequency bands that can be accommodated is precisely determined based on the actual capacity limit of the high-speed memory, avoiding resource over-allocation or waste. Simultaneously, the mapping table mechanism supports dynamic adjustment of the number and priority of frequency bands, facilitating system upgrades and scenario adaptation. Introducing historical signal mutation frequency as an evaluation factor enables the storage allocation strategy to identify and prioritize frequency bands with significant transient characteristics, improving system response performance and reliability in scenarios such as event detection and anomaly warning. A pre-generated mapping table manages storage allocation relationships uniformly, automatically loading coefficients and initializing partitions upon system power-up, reducing manual configuration workload and error probability, and supporting rapid switching and reconstruction under different application scenarios.
[0018] In some embodiments, controlling the final output based on the pulse repetition period and the control signal includes: Based on the time segment allocated to each frequency band, the number of filtered results that need to be buffered in the current pulse repetition period for each frequency band is calculated in real time, and an independent ring buffer is allocated to each frequency band based on the number of filtered results. Each circular buffer is configured with independent write and read pointers, and these pointers are managed using a circular incrementing method. Within the computation time window corresponding to each frequency band, the time-division write enable signal generated by the computation time window is used as a write trigger to control the write pointer of the corresponding ring buffer to increment, and the filtering result of the current frequency band is written to the storage address pointed to by the write pointer. At the end of each pulse repetition cycle, a global read enable pulse is generated using the frame end signal; In response to the global read enable pulse, the read pointers of all circular buffers are reset to the starting address, and a synchronous read process is started to read all buffered data from each circular buffer in sequence. The real and imaginary data of all frequency bands read at the same time are spliced and assembled into a multi-band output data frame in a preset order. A frame valid identification signal is generated that is synchronized with the transmission process of the multi-band output data frame. The high-level valid window of the frame valid identification signal determines the final output to the external system.
[0019] By employing the above technical solution, an independent circular buffer is allocated to each frequency band, and the write pointers are controlled by a time-division write enable signal, completely eliminating write address conflicts between different frequency bands. Reading is triggered uniformly at the end of the pulse repetition cycle, ensuring synchronous output of data from all frequency bands and guaranteeing data integrity and timing consistency. The buffer depth is dynamically allocated based on the actual output data volume of each frequency band, avoiding storage waste or insufficiency caused by uniform allocation. The circular buffer structure, combined with circular pointer management, enables cyclic reuse of storage space, significantly reducing the static occupation of critical resources such as BRAM. The frame end signal is used as a global synchronization trigger, ensuring that read operations in all buffers start at the same time, eliminating output delay differences between frequency bands, providing strictly aligned multi-frequency band data frames for subsequent processing modules, and improving system-level timing accuracy.
[0020] In a second aspect, embodiments of this application provide a computer system including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the steps of the method described in any possible implementation of the first aspect.
[0021] Thirdly, embodiments of this application provide a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of the method described in any possible implementation of the first aspect.
[0022] Fourthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the method described in any possible implementation of the first aspect.
[0023] It is understood that the computer system provided in the second aspect, the storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are all used to execute the method provided in this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0024] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By partitioning and storing complex coefficients according to the proportion of signal energy (the core frequency band is stored in high-speed memory and the edge frequency band is stored in ordinary memory), combined with the address mapping of "frequency band ID + order", the single-cycle fast reading of multi-frequency band coefficients is realized, ensuring the real-time processing of key frequency bands and optimizing the utilization efficiency of storage resources. 2. Based on the pulse repetition period, multiple frequency bands are time-division scheduled so that the same group of complex multiplication units can process different frequency bands in sequence according to time slices. This solves the problem of insufficient DSP resources when filtering multiple frequency bands in parallel. Multi-band complex coefficient FIR filtering is implemented on a single FPGA, reducing system cost and power consumption. 3. By adopting a multi-level parallel accumulation structure (pairwise pairing and summation, stepwise compression), the accumulation operation of long-order FIR is decomposed into a multi-level parallel pipeline, which significantly reduces the accumulation latency and greatly improves the operation speed and system real-time performance while ensuring controllable timing. 4. By constructing a delay chain that is depth-matched with the accumulation pipeline, synchronous compensation is performed on the valid data signal and the frame end signal, ensuring that the control signal and the filtering result are strictly aligned, avoiding frame misalignment and data loss, and improving the consistency of multi-band output timing and system reliability. 5. By combining the pulse repetition period and the synchronized control signal, the multi-band filtering results are uniformly buffered and synchronously output, ensuring the integrity and timing alignment of the data in each frequency band within the pulse repetition period, and providing a stable and consistent data interface for subsequent processing. Attached Figure Description
[0025] Figure 1 This is a flowchart illustrating a distributed acoustic sensing signal processing method based on FPGA in an embodiment of this application. Figure 2 This is a schematic diagram of the core steps of a distributed acoustic sensing signal processing method based on FPGA in an embodiment of this application. Figure 3 This is a schematic diagram of the system framework flow in the embodiments of this application; Figure 4 This is a schematic diagram of an exemplary hardware structure of a computer system in an embodiment of this application. Detailed Implementation
[0026] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.
[0027] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0028] The following is combined Figure 1 The method of the embodiments of this application will be described below.
[0029] Please see Figure 1 This is a flowchart illustrating a distributed acoustic sensing signal processing method based on FPGA, as described in an embodiment of this application. The distributed acoustic sensing signal processing method based on FPGA includes the following steps: S101. Obtain the distributed acoustic sensing signal to be processed, and generate a corresponding data valid signal and frame end signal based on the pulse repetition period of the distributed acoustic sensing signal. S102. Based on frequency band priority, the preset order complex coefficients of multiple frequency bands are partitioned and configured in the on-chip heterogeneous memory of the FPGA. The priority is determined based on the signal energy ratio of each frequency band. The preset order complex coefficients of the first frequency band with the highest signal energy ratio are stored in the first type of memory, and the preset order complex coefficients of the second frequency band other than the first frequency band are stored in the second type of memory. The access speed of the first type of memory is higher than that of the second type of memory. S103. Based on the pulse repetition period, determine the system processing timing. According to the timing and the storage location of the preset order complex coefficient partition configuration, schedule the finite-length unit impulse response filtering operations of the multiple frequency bands to the same group of complex multiplication operation units in a time-division manner. This is so that within the operation time window corresponding to the target frequency band, the preset order complex coefficients are read from the corresponding first type memory or second type memory according to the storage location of the target frequency band and input to the complex multiplication operation unit. The target frequency band is any one of the multiple frequency bands. S104. Within the operation time window, using the preset order complex coefficients input to the complex multiplication operation unit, perform complex multiplication operation on the input signal to obtain multiple multiplication accumulation results, and perform multi-level parallel accumulation on the multiple multiplication accumulation results. Each level of accumulation pairs and sums the results output by the previous level to compress the number of nodes until the final filtering result of the target frequency band is obtained. S105. The data valid signal and the frame end signal are delayed by a delay compensation mechanism to generate a control signal that is time-aligned with the output of the final filtering result, and the final output is controlled based on the pulse repetition period and the control signal.
[0030] Figure 2 This is a schematic flowchart illustrating the core steps of a distributed acoustic sensing signal processing method based on FPGA, as described in an embodiment of this application. Figure 2 As shown, the core steps include complex coefficient partitioning storage, multi-band time-sharing scheduling, hierarchical parallel accumulation, timing synchronization compensation, and multi-band synchronous output.
[0031] Figure 3 This is a schematic diagram of the system framework flow in the embodiments of this application, such as... Figure 3 As shown, the system includes a control and clock module, a time-sharing scheduling control module, a complex coefficient storage module, a hierarchical accumulation operation module, a timing synchronization module, and a multi-band buffer module. The FPGA processing unit, as the core processing unit, integrates the complex coefficient storage module, the time-sharing scheduling control module, the hierarchical parallel accumulation operation module, the timing synchronization module, and the multi-band dynamic ring buffer module. The entire logic design is implemented using Verilog hardware description language. All modules operate within the system's preset clock domain to ensure synchronization between computation and data transmission. The signal input unit receives the digital signal after photoelectric conversion and analog-to-digital conversion from the DAS system. The signal bandwidth covers the target frequency band, and it synchronously provides data validity signals and frame end signals to identify valid data segments and pulse frame boundaries. Data Buffer Unit: A hierarchical storage architecture is constructed using on-chip BRAM and UltraRAM on the FPGA, divided into a coefficient storage area (including the UltraRAM core area and BRAM edge area), a multi-band operation window buffer area (an independent partition matching the number of frequency bands), and an output buffer area. The total storage depth adapts to the number of sampling points corresponding to one pulse repetition cycle, and the bit width covers the coefficient bit width, input data bit width, intermediate result bit width, and output data bit width. Clock and Control Unit: Composed of a phase-locked loop (PLL), a timing counter, window partitioning logic, and a parameter register group. The PLL generates a stable system clock signal (phase noise meets design requirements), the timing counter monitors the pulse repetition cycle in real time, the window partitioning logic generates a time-division enable signal, and the parameter register group stores configurable parameters (such as frame end delay levels, buffer depth, and frequency band ID mapping relationships). Connection relationships between units: The signal input unit is connected to the port of the FPGA processing unit through a parallel data bus; the data buffer unit realizes bidirectional data interaction (coefficient reading, intermediate result writing, and output data reading) with the FPGA processing unit through a dual-port data bus; the clock and control unit distributes the system clock to all modules through the clock signal line, and outputs module enable, read / write control, reset signals and parameter configuration signals through the control signal line to ensure that all units work together.
[0032] The embodiments of this application are described in detail below: The system receives the digital signal (i.e., the distributed acoustic sensing signal to be processed) from the DAS system after photoelectric conversion and analog-to-digital conversion. This signal bandwidth covers the system's preset target monitoring frequency band (e.g., the infrasound band for geological disaster monitoring or the vibration band for perimeter security). The system synchronously generates a valid data signal and a frame end signal. The valid data signal identifies the valid sampled data segment within each pulse width. Its high-level duration strictly corresponds to the width of the emitted laser pulse, ensuring that only sampling points containing valid backscattered Rayleigh signals are sent to subsequent processing. The frame end signal marks the end of a complete pulse repetition cycle. This signal generates a clock-cycle-width pulse at the end of each PRI, serving as a boundary marker for the data frame and is crucial for maintaining frame synchronization in multi-band processing. Based on the DAS system's pre-calibration of the target monitoring frequency band or real-time spectrum analysis, the average signal energy percentage of each frequency band in historical monitoring data is calculated. The frequency band with the highest energy percentage is defined as the core frequency band (i.e., the first frequency band), and the rest are edge frequency bands (i.e., the second frequency band). Core band coefficients: Stored in Type I memory (e.g., UltraRAM on the FPGA chip). UltraRAM features large capacity, high bandwidth, and low latency, enabling single-cycle, high-speed coefficient retrieval for the core band, ensuring real-time processing of critical signals. Edge band coefficients: Stored in Type II memory (e.g., ordinary block random access memory (BRAM)). BRAM has abundant resources but slightly higher access latency than UltraRAM, used to store coefficients of relatively lower importance, achieving a balance between performance and resources. Addressing is performed using a two-dimensional composite address of "band ID + order index". For example, the address width is 11 bits, where the high 3 bits [10:8] represent the band ID, and the low 8 bits [7:0] represent the tap order index of the FIR filter within that band. The address decoder determines whether the target frequency band is a core or edge band based on the high-order bit of the frequency band ID, thereby selecting the UltraRAM or the corresponding BRAM storage block and directly mapping the order index to the address lines of the storage unit. This enables precise reading of the specified frequency band and specified order coefficients within a single clock cycle, eliminating address calculation latency during multi-band switching. Timing reference establishment: Based on the system sampling rate (Fs) and the known pulse repetition period (PRI), the total number of sampling points Ntotal = Fs × PRI within a single PRI is calculated. Operation window division: The time-sharing scheduling control module embeds a loop counter with a counting range of 0 to (Ntotal-1). This counter divides the entire PRI cycle into several consecutive and non-overlapping operation time windows, each allocated to a specific frequency band. The sum of the window lengths for all frequency bands is strictly equal to Ntotal, ensuring that all frequency bands receive processing opportunities within a PRI without losing sampling points. Operation unit multiplexing: A unique time-sharing enable signal is generated for each time window.When the system clock count falls within the window of a certain frequency band (e.g., frequency band A), the time-division enable signal of frequency band A is set to high. At this time: the address decoder reads all the preset-order complex coefficients of frequency band A sequentially from its corresponding memory area (UltraRAM or BRAM) according to the ID of frequency band A; the same set of complex multiplication units (configured by the FPGA's DSPSlice) is enabled, receiving the read coefficients and the current input signal sample value. Through this time-division multiplexing mechanism, DSP multiplication resources are shared by multiple frequency bands, breaking through the hardware bottleneck caused by each frequency band needing to occupy a dedicated set of DSP resources in traditional schemes, making it possible to implement multi-band complex coefficient FIR filtering on FPGAs with limited DSP resources. The complex multiplication unit performs parallel multiplication of the input signal sample value and the read complex coefficients, generating a preset-order (e.g., 128-order) real part product result and the same number of imaginary part product results. Hierarchical parallel accumulation: a tree-like compressed structure is used to sum the product results, rather than the traditional long-chain serial accumulation. First-stage compression: The product results of the real parts (excluding the initial zero value or the historical value of the previous sampling point) of (preset order - 1) are paired up, and the sum of each pair is calculated simultaneously by a set of parallel adders. If the order is odd, the remaining product result is directly passed to the next stage as the remainder. The imaginary part is processed synchronously. The intermediate results (sum value and possible remainder) output from the previous stage are used as a new set of input nodes, and the pairwise sums are calculated again. With each stage, the number of nodes to be accumulated is approximately halved. After several stages of compression (the number of stages is determined by the preset order, for example, 128th order requires about 7 stages), the real and imaginary parts are compressed into only two remaining intermediate values. Finally, the final summation of the real and imaginary parts is completed by two adders respectively, and the final complex numerical filtering result (real and imaginary parts) of the current sampling point in this frequency band is output. The output of each accumulation stage is inserted into a pipeline register. While increasing the system clock frequency, a synchronous reset design ensures that each operation is completed and cleared within a window, preparing for the next frequency band. The register width increases by 1-2 bits at each stage to accommodate the data width increase during accumulation and prevent overflow. Data validity delay chain: A chain of D flip-flops of length L (L equals the total number of pipeline stages for multi-stage accumulation) is constructed. The original valid input data signal passes through this L-stage flip-flop sequentially, generating a regenerated synchronous valid data signal delayed by L clock cycles. Frame end signal delay chain: Similarly, a chain of D flip-flops of length L is constructed, delaying the input frame end signal by the same amount to generate a regenerated synchronous frame end signal. Through the above compensation, the edges of the control signal are precisely aligned in time with the final filtered result output after L cycles. An independent circular buffer is allocated in the BRAM for each frequency band, with a depth exactly equal to the number of output results corresponding to the operation window of that frequency band within one PRI cycle.Within the computation window of each frequency band, the time-division multiplexing enable signal is used as the write enable to sequentially write the calculated final filtering result (real and imaginary parts) into the corresponding circular buffer for that frequency band. A cyclically incrementing write pointer is used for management. At the end of each PRI cycle, a global read enable pulse is triggered by the delay-compensated regenerated synchronization frame end signal. This pulse simultaneously resets the read pointers of all circular buffers to their starting addresses and initiates synchronous read operations. Data synchronously read from the buffers of all frequency bands is concatenated into a complete multi-band output data frame according to a preset frequency band order, combining the real and imaginary parts. Simultaneously, the regenerated synchronization data validity signal is used as the frame validity identifier signal for this output data frame, indicating its valid transmission window. Finally, this packaged data frame and its control signals are sent to subsequent demodulation, positioning, or host computer processing modules to complete the entire signal processing flow.
[0033] In some embodiments, determining the system processing timing based on the pulse repetition period includes: Based on the system sampling rate and the pulse repetition period, the time length of a single pulse repetition period and the corresponding total number of sampling points are determined. The time length is divided according to the preset calculation weights of the multiple frequency bands, and a continuous time segment is assigned to each frequency band. All time segments are connected end to end and their sum is equal to the pulse repetition period. Establish a mapping relationship between system clock counting and frequency band operation, so that when the system clock count falls within the time segment allocated by the third frequency band, the filtering operation of the third frequency band is automatically triggered. The third frequency band is any one of the multiple frequency bands. The system processing timing is determined based on the time segment and the mapping relationship, so that the complex multiplication unit processes a signal of one frequency band at each moment, and each frequency band is processed in sequence.
[0034] The pulse repetition interval (PRI) of a known DAS system is a fixed value \(T_{pri}\) (e.g., 100 μs), which determines the maximum detection distance and the repeated detection rate of the system. The sampling rate (\(F_s\)) of the system is known (e.g., 100 MHz), which determines the time resolution. The total number of sampling points \(N_{total}\) contained within a single pulse repetition interval is determined by the formula: \(N_{total}=F_s\times T_{pri}\). This \(N_{total}\) defines the "time budget" (in sampling clock cycles) that the system must complete for one round of processing all frequency bands and is the basis for all subsequent timing designs. Not all frequency bands are equally important or require the same processing intensity, and preset operation weights are used to quantify the relative processing requirements of each frequency band. This weight can be based on the signal energy ratio, target event characteristics, and filter order. Signal energy ratio: Frequency bands with high energy usually require higher processing accuracy or more priority responses. Target event characteristics: Key characteristic frequency bands of specific monitoring targets (such as pipeline leaks, intrusion events). Filter order: FIR filters for different frequency bands may have different orders. The higher the order, the greater the single-point operation volume. For example, for three frequency bands (A, B, C), their preset weights may be [0.5, 0.3, 0.2]. Allocate continuous time segments: According to the weights of each frequency band, divide the total time budget \(N_{total}\) proportionally. The length of the time segment for frequency band A: \(L_A = N_{total}\times weight_A\); The length of the time segment for frequency band B: \(L_B = N_{total}\times weight_B\); The length of the time segment for frequency band C: \(L_C = N_{total}\times weight_C\). Key constraint: \(L_A + L_B + L_C = N_{total}\). This ensures the completeness of the division, that is, a pulse repetition interval is fully utilized without idle gaps (or gaps are precisely reserved for control, transmission, etc.) and there is no risk of timeout. These time segments are connected end-to-end on the time axis to form a continuous, non-overlapping sequence of operation windows. For example: [0, \(L_A - 1\)] is assigned to frequency band A, [\(L_A\), \(L_A + L_B - 1\)] is assigned to frequency band B, and [\(L_A + L_B\), \(N_{total}-1\)] is assigned to frequency band C. Design a free-running system clock counter with a counting range from 0 to (\(N_{total}-1\)) and automatically reset to 0 after reaching (\(N_{total}-1\)), and its reset signal is synchronized with the input frame end signal. Implement a combinational logic mapping circuit in hardware (usually within a time-sharing scheduling control module). This circuit takes the current value of the counter as input and outputs the enable selection signals for each frequency band. Mapping rule: By comparing the current value of the counter (CNT) with the boundary addresses of each time segment, the enable signal for the corresponding frequency band is activated. If \(0\leq CNT < L_A\), then \(enable_A = 1\), and the enables for other frequency bands are 0. If \(L_A\leq CNT < L_A + L_B\), then \(enable_B = 1\). If \(L_A + L_B\leq CNT\leq (N_{total}-1)\), then \(enable_C = 1\).This mechanism ensures that when the system clock count falls within the time segment allocated to any frequency band (i.e., the third frequency band), the filter enable signal for that frequency band will automatically and deterministically go high, triggering its operation. The above mapping relationship generates a set of mutually exclusive frequency band enable signals (i.e., only one is high at any given time). This set of signals directly controls the ownership of the same set of complex multiplication operation units (and their associated coefficient read paths, accumulator initialization, etc.). At time t (a certain clock cycle), only the enabled frequency band can occupy the shared operation unit. As the clock count increments, the service objects of the operation unit cyclically switch in the order of frequency band A → frequency band B → frequency band C → frequency band A → ... Within a complete pulse repetition cycle (Ntotal clock cycles), each frequency band precisely occupies the computing resources within its allocated continuous time segment and completes the filter calculations for all its sampling points within that cycle.
[0035] In some embodiments, the plurality of multiplicative accumulation results include a preset number of real part multiplication results and a preset number of imaginary part multiplication results, and the multi-level hierarchical parallel accumulation of the plurality of multiplicative accumulation results includes: The product results of the real parts of the preset order are paired in order to form real part node pairs, and the remaining single product results of the real parts are recorded as real part remainders. Based on the real part node pairs and the real part remainders, a set of real part input nodes is constructed. The product results of the imaginary parts of the preset order are paired in order to form imaginary part node pairs, and the remaining single product results of the imaginary parts are recorded as imaginary part remainders. Based on the imaginary part node pairs and the imaginary part remainders, a set of imaginary part input nodes is constructed. A preset number of parallel adders are used to perform synchronous summation on all pairs of real part nodes in the real part input node set to generate a real part compression result set, and the real part remainder is added to the real part compression result set. A preset number of parallel adders are used to perform synchronous summation on all pairs of imaginary part nodes in the imaginary part node set to generate an imaginary part compression result set, and the imaginary part remainder is added to the imaginary part compression result set. Use the real part compression result set as the new real part input node set, and the imaginary part compression result set as the new imaginary part input node set. Repeat the previous step until the real part nodes are compressed into a single real part result and the imaginary part nodes are compressed into a single imaginary part result. The single real part result and the single imaginary part result are output as the final real part filtering result and the final imaginary part filtering result of the target frequency band, respectively.
[0036] The output from the complex multiplication unit is a set of N complex products of a predetermined order (e.g., N=128). Each product can be split into a real part and an imaginary part. Therefore, the input to this step is N real part product results (denoted as R[0], R[1], ..., R[N-1]) and N imaginary part product results (denoted as I[0], I[1], ..., I[N-1]). Usually, the first invalid product corresponding to the initial zero value or historical data is removed in the FIR operation, so the actual number of valid terms participating in the accumulation is N-1 (e.g., 127). For simplicity, we will still take N terms as an example here, but the principle is the same. Initial node set construction (first level input): Pair the N real part product results in order. That is, (R[0], R[1]) is the first pair, (R[2], R[3]) is the second pair, and so on. Processing Remainders: If N is even, all terms are paired; if N is odd, the last product result R[N-1] cannot be paired and is marked separately as a real part remainder (R_remainder). All paired "real part node pairs" and possible "real part remainders" are combined to form the first-level real part input node set. This set contains approximately N / 2 node pairs to be summed and 0 or 1 remainders. The same operation is performed on the N imaginary part product results, forming an imaginary part input node set consisting of imaginary part node pairs and possible imaginary part remainders (I_remainder). The processing of the real and imaginary parts is completely independent and parallel. MR1 parallel adders are configured for the real part input node set, where MR1 equals the number of node pairs in the set. Each adder synchronously and independently computes the sum of one node pair. For example, adder 1 calculates Sum_R1[0] = R[0] + R[1], adder 2 calculates Sum_R1[1] = R[2] + R[3], and so on, with all additions completed within one clock cycle. The same operation is performed on the imaginary part, using MI1 adders (MI1 = MR1) to synchronously calculate the sum of each imaginary part node pair, resulting in Sum_I1[0], Sum_I1[1], ... . The sums of all the parallel-calculated real part node pairs are combined into a new set. Key operation: The real part remainder R_remainder (if it exists) left over from the previous stage is directly merged into this new set. At this time, the elements of the new real part compression result set are: {Sum_R1[0], Sum_R1[1], ..., Sum_R1[MR1-1], (R_remainder)}. Similarly, a new imaginary part compression result set is obtained: {Sum_I1[0], Sum_I1[1], ..., Sum_I1[M_I1-1], (I_remainder)}. The current real part compression result set and imaginary part compression result set are checked. If the number of elements in either set is greater than 1, compression needs to continue. The entire real part compression result set obtained in the previous step is used as the new real part input node set.For this new input node set, repeat the above steps: pair its elements sequentially to form new node pairs, process any new remainder terms, and simultaneously sum the node pairs using a suitable number of parallel adders (approximately half the size of the current set). Combine the sum with the remainder terms to form the updated real part compressed result set. The imaginary part undergoes the same iterative process. After each level of this operation, the number of nodes to be accumulated for both real and imaginary parts is reduced by approximately half. After K levels of iterative compression (K≈log2(N)), the number of elements in the real part compressed result set is compressed to 1, resulting in a single real part result (Final_R). Simultaneously, the imaginary part compressed result set is also compressed to a single imaginary part result (Final_I). Output Final_R as the final real part filtering result for the current sampling point in the current target frequency band. Output Final_I as the final imaginary part filtering result. Final_R and Final_I together constitute a complete complex number, representing the instantaneous response of the input signal after passing through the complex coefficient FIR filter of this frequency band.
[0037] In some embodiments, the step of delaying the valid data signal and the end-of-frame signal through a delay compensation mechanism to generate a control signal that is time-aligned with the output of the final filtered result specifically includes: Based on the total number of delay levels accumulated in parallel at multiple levels, a data validity delay chain and a frame end signal delay chain are constructed respectively. Both the data validity delay chain and the frame end signal delay chain include the total number of cascaded triggers. The data valid signal is input into the data valid delay chain and passed through each stage for a total delay level of several clock cycles to generate a regenerated synchronous data valid signal. The frame end signal is input into the frame end signal delay chain and passed through each stage for a total delay level of several clock cycles to generate a regenerated synchronous frame end signal. During the effective period when the regenerated synchronization data valid signal is high, the real and imaginary parts of the final filtering result are continuously sampled to obtain sampled data. By verifying the phase continuity of the sampled data during the effective period, it is possible to check whether the phase alignment accuracy between the regenerated synchronization data valid signal and the final filtering result meets the requirements. When the phase alignment accuracy meets the requirements, the valid signal of the regenerated synchronization data and the end signal of the regenerated synchronization frame are used as control signals and output synchronously with the final filtering result.
[0038] As shown above, the multi-stage parallel accumulation operation adopts a pipelined structure. Assume that this structure requires a total of L clock cycles from input complex number multiplication and accumulation to output the final complex number result. This value L is predetermined and fixed by the hardware design, including the number of accumulator stages and the pipeline registers of each stage. L is the baseline for delay compensation. For example, L=7 means that it takes 7 clock cycles for data to travel from the input of the accumulation module to the output. Data validity delay chain: Consists of L cascaded D flip-flops (DFFs). The data input of the first flip-flop receives the original data validity signal. The clock input of each flip-flop is connected to the system master clock, and the reset input is connected to the global synchronization reset signal. The output of the previous flip-flop is connected to the input of the next flip-flop. Frame end signal delay chain: Uses the exact same structure, also consisting of L cascaded D flip-flops, but its data input receives the original frame end signal. Each D flip-flop latches the input data and outputs it on the rising edge of the clock, generating a delay of 1 clock cycle. L flip-flops cascaded together produce a precise delay of L clock cycles for the input signal. Regenerated Synchronous Data Valid Signal: The original data valid signal is passed through L stages of the data validity delay chain and output from the output of the last flip-flop. The valid level (high level) window of this new signal is shifted backward by L clock cycles on the time axis. Its rising and falling edges are aligned with the output time of the final filtered result. Regenerated Synchronous Frame End Signal: The original frame end signal (usually a single-cycle pulse) also passes through the frame end signal delay chain, and its pulse occurrence time is also precisely delayed by L clock cycles to align with the end time of the last valid data output, accurately marking the boundary of the delayed data frame. During the entire period when the regenerated synchronous data valid signal is high, the system does not passively output, but actively and continuously samples the real and imaginary parts of the final filtered result being output at this time. Verification Logic: The core of the verification is to determine whether the data stream is smooth and continuous, without any "glitch," "undefined state," or unexpected jumps caused by timing misalignment. Specific methods may include: within the effective window, the difference between adjacent sampling points should be within the expected range based on the signal's physical characteristics; the sign bit of the data should not exhibit frequent and irregular flips within a short period; if the effective data segment contains a known preamble or a specific pattern, the integrity and correctness of the pattern can be checked. In this embodiment, "phase" refers not only to the carrier phase of the signal but also, more broadly, to the continuity of the data sample sequence on the time axis and the consistency of the waveform. High alignment accuracy means that the effective signal window is precisely fitted onto a clean and stable data waveform. The verification circuit generates an alignment-ready flag. When the sampled data passes the phase continuity check within several consecutive pulse repetition cycles, the phase alignment accuracy between the regenerated synchronous data effective signal and the final filtered result is considered to meet the system requirements.The system is only allowed to execute the final step when the alignment-ready flag is valid (e.g., high): combining the regenerated synchronization data valid signal, the regenerated synchronization frame end signal, and the final filtering result into a single synchronized data packet and outputting it to subsequent modules. If the check fails, the system can trigger an alarm, maintain the current output disabled state, or attempt to fine-tune the delay chain (if the design supports dynamic adjustment).
[0039] In some embodiments, verifying the phase continuity of the sampled data within the effective time period to check whether the phase alignment accuracy between the regenerated synchronization data effective signal and the final filtering result meets the requirements specifically includes: The target adder with the largest delay in the multi-level parallel accumulation is selected as the monitoring point, and the stable establishment time margin of the output signal of the target adder from data validity to sampling by the next level register is monitored in real time. When the stable setup time margin is lower than the dynamic safety threshold, the number of delay clock cycles that need to be additionally compensated is calculated based on the difference between the stable setup time margin and the dynamic safety threshold, as well as the system clock cycle, and is used as the delay compensation increment. Based on the delay compensation increment, by controlling the data selector set at the input of each stage of the trigger in the data validity delay chain and the frame end signal delay chain, the triggers at a preset level are bypassed or connected to adjust the effective depth of the data validity delay chain and the frame end signal delay chain. After adjusting the effective depth, the step of restoring the verification phase continuity is performed.
[0040] In the hardware circuit of "multi-level hierarchical parallel accumulation", different adders have different degrees of timing tightness of their output signals due to different arrival times of input signals and different lengths of combinational logic paths. Target adder: Determine in advance, through static timing analysis (STA) or simulation, the adder with the largest delay and the tightest setup time. This path is usually the critical path of the entire accumulation module, and its stability directly determines whether the system can operate stably at the preset clock frequency. Insert a timing monitoring unit between the output of the target adder and the next-level register that samples it. The timing monitoring unit can measure in real time the time interval from when the data output from the adder is valid (logically stable) to the arrival of the next clock rising edge, that is, the "stable setup time margin (T_slack)". T_slack is a real-time and dynamic metric. T_slack > 0 indicates that the timing meets the requirements, and the larger the value, the more sufficient the timing margin. T_slack will decrease as the chip temperature rises, the supply voltage drops, or the transistors age. The dynamic safety threshold (T_threshold) is not a fixed value, but a safety margin greater than zero preset according to the system reliability requirements. For example, set to 10% of the clock period (T_clk), which represents the minimum timing margin allowed by the system. Continuously compare the real-time monitored T_slack with T_threshold. If T_slack ≥ T_threshold: It is considered that the current timing is healthy, the delay chain depth L is still appropriate, and no adjustment is required; if T_slack < T_threshold: Trigger an alarm! It indicates that the timing margin of the critical path has fallen below the safety line, and the currently used delay compensation amount (L cycles) may not be sufficient to fully align the data and control signals because the actual delay of the data path may have drifted slightly. Calculate the delay compensation increment (ΔL): When T_slack < T_threshold, the system automatically calculates the additional compensation delay required. Calculation formula: ΔL = ceil((T_threshold - T_slack) / T_clk). Calculate how many complete system clock cycles (T_clk) the currently shorted timing margin (T_threshold - T_slack) is equivalent to, and round up (ceil) to ensure sufficient compensation. For example, T_clk = 10 ns, T_threshold = 1 ns, and T_slack = 0.2 ns are measured, then ΔL = ceil((1 - 0.2) / 10) = ceil(0.08) = 1. It means that an additional 1 clock cycle of delay needs to be added for compensation. Configurable delay chain structure: The traditional fixed series flip-flop chain is modified, and a 1-of-2 data selector (MUX) is inserted before the data input terminal of each flip-flop.The MUX has two inputs: Input 0 (bypass path): directly from the output of the previous stage flip-flop (or the original input signal for the first stage); Input 1 (normal path): from the output of the previous stage flip-flop, but after passing through an additional buffer unit with precisely controllable delay (or directly considered as passing through the current stage flip-flop). The MUX selection is driven by a unified depth control logic that generates a control word based on the calculated ΔL. Initial state: The control word causes all MUXs to select "Input 1," and the delay chain behaves as a standard L-stage flip-flop chain with an effective depth of L. Adding compensation (when ΔL > 0): The depth control logic is reconfigured. It might switch the last ΔL MUXs in the delay chain to select "Input 0." However, this is usually inaccurate. A more typical strategy is: in the delay chain, L stages are fixed as the basic compensation; an additional series of optional additional delay units (also composed of flip-flops and MUXs) are provided; when additional compensation is needed, the control logic "connects" these additional units in series to the end of the main delay chain, making the total effective depth L + ΔL. Adjustment in effect: After the new control word is loaded, the signal path through the delay chain changes, increasing the total delay by ΔL*T_clk. The generation of the regenerated synchronization data valid signal and the regenerated synchronization frame end signal is therefore delayed by an additional ΔL cycles. After adjusting the delay chain depth, the system does not immediately assume synchronization has been restored; it automatically re-executes the previously described "phase continuity check" step. Under the new delay settings, the final data is sampled again within the regenerated synchronization data valid signal window to verify its phase continuity. If the check passes, the system operates stably under the new delay configuration. If the check still fails, and the detected T_slack is still insufficient, the control logic may attempt further fine-tuning of ΔL (e.g., adding another cycle), and then check again, forming a closed-loop "monitoring-calculation-adjustment-verification" feedback loop until the synchronization accuracy meets the requirements again.
[0041] In some embodiments, the step of partitioning and configuring the complex coefficients of multiple frequency bands of a preset order based on frequency band priority in the on-chip heterogeneous memory of the FPGA includes: By using a preset weighted scoring function, the historical average signal energy of each frequency band, the frequency of signal mutations within a preset historical period, and the preset system-level importance weights are weighted and summed to generate a comprehensive priority score. All frequency bands are sorted in descending order according to the comprehensive priority score. The upper limit of the number of frequency bands that the first type of memory can accommodate is calculated according to the ratio of the total capacity of the first type of memory to the amount of complex coefficient data in a single frequency band. The first frequency band is defined according to the upper limit of the number of frequency bands. The preset order complex coefficients of the first frequency band are preloaded into the first type of memory. The preset order complex coefficients of the remaining frequency bands are stored in the second type of memory. The first frequency band in the first type of memory is sorted a second time according to the comprehensive priority score, and stored in the memory blocks with increasing access latency in descending order of comprehensive priority score; A mapping table of frequency band ID, memory type, and storage block address is generated and written to the configuration memory of the FPGA. When the system is powered on and initialized, the preset order complex coefficients of each frequency band are loaded from external storage to the corresponding heterogeneous memory partition on the FPGA chip according to the mapping table.
[0042] Historical Average Signal Energy (E_avg): This calculates the average energy of the received signal in each frequency band over a past period (e.g., the last 1000 pulse repetition cycles). Higher energy usually indicates that the band contains more target signals or environmental background signals, making it a priority for filtering and requiring faster access speeds to ensure real-time performance. Signal Transient Frequency (F_transient): This calculates the number of rapid changes (transients) in signal amplitude exceeding a specific threshold within a preset historical period for each frequency band. High-frequency transients may indicate abnormal events (such as intrusion or leakage), and the corresponding frequency bands should be prioritized to ensure sensitivity and response speed in event detection. System-Level Importance Weight (W_sys): This is a weight statically set in advance by the system designer or user. It reflects the inherent importance of certain frequency bands based on prior knowledge. For example, in bridge health monitoring, frequency bands reflecting the inherent frequencies of the structure may be assigned a higher fixed weight; in security systems, frequency bands reflecting human movement characteristics may have a higher weight. Design a scoring function, for example: Priority_Score_i = α × E_avg_i + β × F_transient_i + γ × W_sys_i. i Representing the iEach frequency band has configurable weighting coefficients α, β, and γ to balance the contributions of the three factors (e.g., α=0.5, β=0.3, γ=0.2). This function calculates a quantified overall priority score for each frequency band. A higher score indicates greater overall importance of the frequency band across historical performance, event sensitivity, and system design goals. All frequency bands are sorted in descending order of their overall priority scores, resulting in a list of frequency bands from highest to lowest priority. The total storage capacity of the first type of memory (such as UltraRAM) is limited (e.g., 4MB). The amount of complex coefficient data per frequency band = preset order × coefficient bit width (real part + imaginary part) × 2 (if real and imaginary parts are stored separately). For example, with 128 orders and 16-bit coefficients, the amount of data per frequency band is approximately 128 × 16 × 2 / 8 = 512 bytes. The upper limit for capacity is calculated as: N_ultraram = floor(total UltraRAM capacity / amount of data per frequency band). Assuming a total UltraRAM capacity of 2MB, it can hold coefficients for floor (2 × 1024 × 1024 / 512) = 4096 frequency bands. Note: This calculation is theoretical; in practice, the data volume of a single frequency band may be larger, and memory addressing granularity, partitioning overhead, etc., must be considered. Based on the calculated N_ultraram, the first N_ultraram of the highest priority frequency bands in the sorted list are designated as the first frequency band. The pre-order complex coefficients of these first frequency bands are pre-loaded into the first type of memory during system initialization. All remaining frequency bands in the list are designated as the second frequency band, and their coefficients are stored in the second type of memory. Even within the first type of memory, the access latency of different memory blocks may have slight differences (e.g., due to wiring distance, memory bank structure). This group of frequency bands already designated as the first frequency band is then sorted again according to their overall priority score. The physical layout follows the principle of "higher score, better storage location": the coefficients of the highest-scoring frequency band are placed in the storage block with the lowest access latency and highest bandwidth in UltraRAM (such as Bank0, or a location close to the DSP unit); the coefficients of the second-highest-scoring frequency band are placed in blocks with slightly higher access latency, and so on. This achieves a positive correlation between access latency and frequency band priority within the first type of memory, further maximizing performance potential. Based on all the above decisions, the system (or the accompanying configuration software) automatically generates a "frequency band ID - memory type - storage block base address" mapping table. This table is a core configuration data structure. For example: ; This mapping table is either programmed into the FPGA's configuration memory (such as Flash) or sent by the host computer during startup. During the system power-on initialization phase, the configuration management logic within the FPGA reads this mapping table. Based on the instructions in the mapping table, it moves the complex coefficient data for each frequency band from external mass storage (such as DDR) to the designated heterogeneous memory partitions (specific addresses of UltraRAM or BRAM) on the FPGA chip. After loading is complete, the entire storage system is ready with the optimized layout, awaiting operation.
[0043] In some embodiments, controlling the final output based on the pulse repetition period and the control signal includes: Based on the time segment allocated to each frequency band, the number of filtered results that need to be buffered in the current pulse repetition period for each frequency band is calculated in real time, and an independent ring buffer is allocated to each frequency band based on the number of filtered results. Each circular buffer is configured with independent write and read pointers, and these pointers are managed using a circular incrementing method. Within the computation time window corresponding to each frequency band, the time-division write enable signal generated by the computation time window is used as a write trigger to control the write pointer of the corresponding ring buffer to increment, and the filtering result of the current frequency band is written to the storage address pointed to by the write pointer. At the end of each pulse repetition cycle, a global read enable pulse is generated using the frame end signal; In response to the global read enable pulse, the read pointers of all circular buffers are reset to the starting address, and a synchronous read process is started to read all buffered data from each circular buffer in sequence. The real and imaginary data of all frequency bands read at the same time are spliced and assembled into a multi-band output data frame in a preset order. A frame valid identification signal is generated that is synchronized with the transmission process of the multi-band output data frame. The high-level valid window of the frame valid identification signal determines the final output to the external system.
[0044] Each frequency band operates and produces filtered results only within its allocated computation time slice within a single pulse repetition cycle. The length of this time slice (in clock cycles) is the number of filtered results (N_band_i) that the frequency band needs to buffer within that pulse repetition cycle. For example, if a 50-clock-cycle window is allocated to frequency band A, then N_band_A = 50. A separate circular buffer is allocated to each frequency band, typically implemented using on-chip BRAM in the FPGA. The depth (i.e., the number of memory cells) of each circular buffer is configured to N_band_i (or slightly larger to allow for margin). For example, a BRAM block with a depth of 50 is allocated to frequency band A. The memory spaces are logically contiguous. When data is written to the end, the next write address wraps back to the beginning. This allows for the continuous and efficient use of memory resources by circularly buffering periodic data streams with a fixed-size memory space. A separate pair of write pointers (Wr_Ptr) and read pointers (Rd_Ptr) is maintained for each circular buffer. The pointers are essentially counters. When incrementing is required (the write pointer increments after writing, and the read pointer increments after reading), the counter increments by 1. When the count reaches the buffer depth (e.g., 50), it automatically resets to zero in the next clock cycle, thus implementing a loop. During initialization, all read and write pointers are reset to 0 (pointing to the starting address of the buffer). The write operation is precisely controlled by the time-division multiplexing write enable signal. This signal is active high, and its active window is completely synchronized with the operation time window of the corresponding frequency band. When the time-division multiplexing write enable signal for frequency band A is active, the current write pointer (Wr_Ptr_A) value of the circular buffer for frequency band A is used as the write address of the BRAM. The final complex filtering result (real and imaginary parts) of frequency band A, calculated at the current time, is written to this address in the BRAM. After the write operation is completed (usually at the rising edge of the clock), the write pointer (Wr_Ptr_A) is incremented by 1 cyclically, pointing to the next free unit, preparing to receive the next result. Since the write enable signals for each frequency band are mutually exclusive in time, data from different frequency bands are written to their respective independent storage spaces, and there is no write address conflict. The read operation is not continuous, but batch-processed and synchronously triggered. At the end of each pulse repetition cycle, a global read enable pulse is generated by a frame end signal that is accurately aligned after delay compensation. This pulse indicates that all the latest data in all frequency bands within a complete pulse repetition cycle is ready and buffered. In response to the global read enable pulse, the system simultaneously resets the read pointers (Rd_Ptr) of all circular buffers to their starting address (0). This ensures that each read starts from the oldest (or the first) data in each buffer. After the reset, the system initiates a continuous read process. For example, in a high-speed clock domain, data is read sequentially or pipelined from all buffers. In the first clock cycle, data is read simultaneously (or sequentially) from address 0 of all buffers.In the second clock cycle, all read pointers are incremented by 1, and data is then read from address 1 of all buffers. This process is repeated until exactly N_band_i data have been read from each buffer (i.e., all data generated in that frequency band within a complete pulse repetition cycle has been read). The data read from the same address offset in each buffer is assembled. For example, in the k-th clock cycle of reading, the k-th result of frequency bands A, B, C... is obtained simultaneously. The real and imaginary parts of these results are concatenated into a wider data word, called a multi-band output data frame, according to a preset, fixed order (e.g., ascending order by frequency band ID). For example, if frequency bands A, B, and C each output 16 real bits + 16 imaginary bits, then the width of the concatenated data frame is (16 + 16) × 3 = 96 bits. Simultaneously with the start of synchronous reading, the system generates a frame validity flag (e.g., out_valid). The high-level valid window of this signal is completely synchronized with the continuous output process of the multi-band output data frame. That is, the signal remains high from the start of outputting the first spliced frame until the completion of outputting the last spliced frame. During each clock cycle when out_valid is high, the corresponding spliced data frame is driven onto the output bus and sent to an external system (such as a CPU, DSP, or transmission interface). Simultaneously, the regenerated synchronization frame end signal can serve as a precise indication of the falling edge of out_valid, or as a separate frame synchronization pulse output, providing the receiver with a clear frame boundary.
[0045] The above describes a distributed acoustic sensing signal processing method based on FPGA in the embodiments of this application. The computer system in the embodiments of this application will be described in detail below with reference to the above-described distributed acoustic sensing signal processing method based on FPGA.
[0046] Please see Figure 4 This is a schematic diagram of an exemplary hardware structure of a computer system in an embodiment of this application.
[0047] In some embodiments, the computer system 400 includes a computer device, which may be a terminal device. The computer device includes a processor 401, a memory 402, a sensor module 403, a communication module 404, an input device 405, and an output device 406 connected via a system bus. The processor 401 of the computer device provides computing and control capabilities. The memory 402 of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database is used to store data.
[0048] Those skilled in the art will understand that Figure 4The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0049] In some embodiments of this application, a computer-readable storage medium is provided, including instructions that, when executed on a computer system 400, cause the computer system 400 to perform an FPGA-based distributed acoustic sensing signal processing method according to an embodiment of this application.
[0050] In some embodiments of this application, a computer program product is also provided, which, when run on a computer system 400, causes the computer system 400 to execute an FPGA-based distributed acoustic sensing signal processing method according to an embodiment of this application.
[0051] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0052] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0053] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A distributed acoustic sensing signal processing method based on FPGA, characterized in that, include: Acquire the distributed acoustic sensor signal to be processed, and generate the corresponding valid data signal and frame end signal based on the pulse repetition period of the distributed acoustic sensor signal; Based on frequency band priority, the preset order complex coefficients of multiple frequency bands are partitioned and configured in the on-chip heterogeneous memory of the FPGA. The priority is determined based on the signal energy ratio of each frequency band. The preset order complex coefficients of the first frequency band with the highest signal energy ratio are stored in the first type of memory, and the preset order complex coefficients of the second frequency band other than the first frequency band are stored in the second type of memory. The access speed of the first type of memory is higher than that of the second type of memory. Based on the pulse repetition period, the system processing timing is determined. According to the timing and the storage location of the preset order complex coefficient partition configuration, the finite-length unit impulse response filtering operations of the multiple frequency bands are scheduled to the same group of complex multiplication operation units in a time-division manner. This ensures that within the operation time window corresponding to the target frequency band, the preset order complex coefficients are read from the corresponding first type memory or second type memory according to the storage location of the target frequency band and input to the complex multiplication operation unit. The target frequency band is any one of the multiple frequency bands. Within the computation time window, using the preset order complex coefficients input to the complex multiplication unit, complex multiplication is performed on the input signal to obtain multiple multiplication accumulation results. Then, multi-level parallel accumulation is performed on the multiple multiplication accumulation results. Each accumulation stage pairs and sums the results output from the previous stage to compress the number of nodes until the final filtering result of the target frequency band is obtained. The data validity signal and the frame end signal are delayed by a delay compensation mechanism to generate a control signal that is time-aligned with the output of the final filtered result, and the final output is controlled based on the pulse repetition period and the control signal.
2. The FPGA-based distributed acoustic sensing signal processing method according to claim 1, characterized in that, The process of determining the system processing timing based on the pulse repetition period includes: Based on the system sampling rate and the pulse repetition period, the time length of a single pulse repetition period and the corresponding total number of sampling points are determined. The time length is divided according to the preset calculation weights of the multiple frequency bands, and a continuous time segment is assigned to each frequency band. All time segments are connected end to end and their sum is equal to the pulse repetition period. Establish a mapping relationship between system clock counting and frequency band operation, so that when the system clock count falls within the time segment allocated by the third frequency band, the filtering operation of the third frequency band is automatically triggered. The third frequency band is any one of the multiple frequency bands. The system processing timing is determined based on the time segment and the mapping relationship, so that the complex multiplication unit processes a signal of one frequency band at each moment, and each frequency band is processed in sequence.
3. The FPGA-based distributed acoustic sensing signal processing method according to claim 1, characterized in that, The multiple multiplicative accumulation results include a preset number of real part multiplication results and a preset number of imaginary part multiplication results. The multi-level, hierarchical, parallel accumulation of the multiple multiplicative accumulation results includes: The product results of the real parts of the preset order are paired in order to form real part node pairs, and the remaining single product results of the real parts are recorded as real part remainders. Based on the real part node pairs and the real part remainders, a set of real part input nodes is constructed. The product results of the imaginary parts of the preset order are paired in order to form imaginary part node pairs, and the remaining single product results of the imaginary parts are recorded as imaginary part remainders. Based on the imaginary part node pairs and the imaginary part remainders, a set of imaginary part input nodes is constructed. A preset number of parallel adders are used to perform synchronous summation on all pairs of real part nodes in the real part input node set to generate a real part compression result set, and the real part remainder is added to the real part compression result set. A preset number of parallel adders are used to perform synchronous summation on all pairs of imaginary part nodes in the imaginary part node set to generate an imaginary part compression result set, and the imaginary part remainder is added to the imaginary part compression result set. Use the real part compression result set as the new real part input node set, and the imaginary part compression result set as the new imaginary part input node set. Repeat the previous step until the real part nodes are compressed into a single real part result and the imaginary part nodes are compressed into a single imaginary part result. The single real part result and the single imaginary part result are output as the final real part filtering result and the final imaginary part filtering result of the target frequency band, respectively.
4. The FPGA-based distributed acoustic sensing signal processing method according to claim 1, characterized in that, The step of delaying the valid data signal and the end-of-frame signal through a delay compensation mechanism to generate a control signal that is time-aligned with the final filtered output specifically includes: Based on the total number of delay levels accumulated in parallel at multiple levels, a data validity delay chain and a frame end signal delay chain are constructed respectively. Both the data validity delay chain and the frame end signal delay chain include the total number of cascaded triggers. The data valid signal is input into the data valid delay chain and passed through each stage for a total delay level of several clock cycles to generate a regenerated synchronous data valid signal. The frame end signal is input into the frame end signal delay chain and passed through each stage for a total delay level of several clock cycles to generate a regenerated synchronous frame end signal. During the effective period when the regenerated synchronization data valid signal is high, the real and imaginary parts of the final filtering result are continuously sampled to obtain sampled data. By verifying the phase continuity of the sampled data during the effective period, it is possible to check whether the phase alignment accuracy between the regenerated synchronization data valid signal and the final filtering result meets the requirements. When the phase alignment accuracy meets the requirements, the valid signal of the regenerated synchronization data and the end signal of the regenerated synchronization frame are used as control signals and output synchronously with the final filtering result.
5. The FPGA-based distributed acoustic sensing signal processing method according to claim 4, characterized in that, The step of verifying the phase continuity of the sampled data within the effective time period to check whether the phase alignment accuracy between the regenerated synchronization data effective signal and the final filtering result meets the requirements specifically includes: The target adder with the largest delay in the multi-level parallel accumulation is selected as the monitoring point, and the stable establishment time margin of the output signal of the target adder from data validity to sampling by the next level register is monitored in real time. When the stable setup time margin is lower than the dynamic safety threshold, the number of delay clock cycles that need to be additionally compensated is calculated based on the difference between the stable setup time margin and the dynamic safety threshold, as well as the system clock cycle, and is used as the delay compensation increment. Based on the delay compensation increment, by controlling the data selector set at the input of each stage of the trigger in the data validity delay chain and the frame end signal delay chain, the triggers at a preset level are bypassed or connected to adjust the effective depth of the data validity delay chain and the frame end signal delay chain. After adjusting the effective depth, the step of restoring the verification phase continuity is performed.
6. The FPGA-based distributed acoustic sensing signal processing method according to claim 1, characterized in that, The step of partitioning and configuring the complex coefficients of multiple frequency bands into the on-chip heterogeneous memory of the FPGA based on frequency band priority includes: By using a preset weighted scoring function, the historical average signal energy of each frequency band, the frequency of signal mutations within a preset historical period, and the preset system-level importance weights are weighted and summed to generate a comprehensive priority score. All frequency bands are sorted in descending order according to the comprehensive priority score. The upper limit of the number of frequency bands that the first type of memory can accommodate is calculated according to the ratio of the total capacity of the first type of memory to the amount of complex coefficient data in a single frequency band. The first frequency band is defined according to the upper limit of the number of frequency bands. The preset order complex coefficients of the first frequency band are preloaded into the first type of memory. The preset order complex coefficients of the remaining frequency bands are stored in the second type of memory. The first frequency band in the first type of memory is sorted a second time according to the comprehensive priority score, and stored in the memory blocks with increasing access latency in descending order of comprehensive priority score; A mapping table of frequency band ID, memory type, and storage block address is generated and written to the configuration memory of the FPGA. When the system is powered on and initialized, the preset order complex coefficients of each frequency band are loaded from external storage to the corresponding heterogeneous memory partition on the FPGA chip according to the mapping table.
7. The FPGA-based distributed acoustic sensing signal processing method according to claim 2, characterized in that, The control of the final output based on the pulse repetition period and the control signal includes: Based on the time segment allocated to each frequency band, the number of filtered results that need to be buffered in the current pulse repetition period for each frequency band is calculated in real time, and an independent ring buffer is allocated to each frequency band based on the number of filtered results. Each circular buffer is configured with independent write and read pointers, and these pointers are managed using a circular incrementing method. Within the computation time window corresponding to each frequency band, the time-division write enable signal generated by the computation time window is used as a write trigger to control the write pointer of the corresponding ring buffer to increment, and the filtering result of the current frequency band is written to the storage address pointed to by the write pointer. At the end of each pulse repetition cycle, a global read enable pulse is generated using the frame end signal; In response to the global read enable pulse, the read pointers of all circular buffers are reset to the starting address, and a synchronous read process is started to read all buffered data from each circular buffer in sequence. The real and imaginary data of all frequency bands read at the same time are spliced and assembled into a multi-band output data frame in a preset order. A frame valid identification signal is generated that is synchronized with the transmission process of the multi-band output data frame. The high-level valid window of the frame valid identification signal determines the final output to the external system.
8. A computer system comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1-7.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1-7.