Distributed general waveform analysis architecture
Through the distributed general waveform analysis architecture, GP Compiler and GP Accelerator are used to solve the problems of low development efficiency and resource competition in the centralized architecture, and realize efficient waveform analysis services.
Patent Information
- Application Number
- CN202510867426.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-10
AI Technical Summary
The existing waveform analysis architecture suffers from low FS accelerator development efficiency, performance degradation caused by resource competition, and low data transmission efficiency, which are particularly evident in the centralized architecture.
A distributed general waveform analysis architecture is adopted, and GP Compiler is used to map the waveform analysis kernel to the distributed GP Accelerator. The data flow execution mode is implemented through a three-layer structure, and the GP Compiler compiler is used for automatic mapping and resource optimization.
It significantly improves waveform analysis development efficiency and service efficiency, avoids data transmission delays and resource competition, and maintains high throughput performance.
Smart Images

Figure CN120761683A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of waveform analysis accelerator, and particularly relates to a distributed general waveform analysis architecture. BACKGROUND
[0002] Compared with traditional oscilloscopes, waveform recorders play a key role in long-term recording of multi-dimensional signals. Taking the automotive motor test scene as an example, it is difficult for an oscilloscope to synchronously measure non-traditional physical quantities such as motor vibration, torque and acceleration for a long time. The waveform analysis capability of the waveform recorder further realizes in-situ signal processing after acquisition, such as deriving real-time motor power calculation through voltage and current measurement, or performing spectral analysis on torque and vibration signals to identify noise sources, which significantly improves the efficiency of problem diagnosis and solution.
[0003] Figure 1 The waveform analysis architecture of an industrial-grade waveform recorder MR-6000 is shown. The device contains multiple data acquisition boards specially designed for long-term acquisition and storage of various physical quantities such as voltage, current, charge, etc. The main control board is responsible for rendering the graphical user interface (GUI), receiving user-triggered data acquisition instructions, and sending control signals through the PXIe interface to coordinate synchronous data acquisition and storage between the acquisition boards. Users can initiate multi-channel waveform analysis requests for acquired data from the acquisition board through the user graphical interface, and the relevant data will be transmitted to the main control board through the PXIe interface. To optimize processing efficiency, common waveform analysis kernels (such as FIR filtering, FFT spectral analysis, minimum value detection, etc.) are accelerated through dedicated hardware accelerators (Function-Specific Accelerator, FS Accelerator); and unconventional or user-defined waveform analysis kernels are processed by the central processing unit (CPU) to maintain system flexibility.
[0004] However, the existing waveform analysis architecture fails to fully focus on the development and service efficiency of waveform analysis. Consistent with most researches on FPGA-accelerated specific algorithms, waveform recorder manufacturers design function FS Accelerators for common waveform analysis kernels and deploy them on the centralized FPGA of the main control board. This method has the following defects: (I) The design process of FS Accelerator for a single waveform analysis kernel is time-consuming, and user-defined kernels are still limited to CPU execution; (II) FS Accelerators of different kernels cause excessive use of FPGA resources, leading to exponential growth of compilation time and performance degradation of all accelerators; (III) The centralized architecture requires a large amount of data transmission from the acquisition board to the main control board, causing a significant efficiency bottleneck. SUMMARY
[0005] The purpose of this invention is to design a new waveform analysis architecture. Its waveform analysis accelerator is a general-purpose accelerator (GP Accelerator). To address the problems of low development efficiency of FS accelerators in the existing technology, centralized deployment leading to resource competition and performance degradation, and low data transmission efficiency, a distributed general-purpose waveform analysis architecture is proposed.
[0006] The present invention proposes a distributed universal waveform analysis architecture, comprising: Main control board, integrated GP Compiler (general purpose accelerator compiler); Multiple data acquisition boards, each with a distributed and independently running GP Accelerator (general waveform analysis accelerator); The GP Compiler automatically maps the high-level language description of the waveform analysis kernel to the hardware configuration of the GP Accelerator; wherein the waveform analysis kernel is executed locally by the GP Accelerator on the data acquisition board, avoiding the transmission of the original data to the main control board.
[0007] Furthermore, the GP Accelerator adopts a three-layer structure, namely: The computing unit structure includes routing registers, configurable crossbar switches, multi-function arithmetic units, and loop execution control modules; The interconnected network structure dynamically configures the computing functions and interconnection directions of each unit to achieve data flow execution modes for different waveform analysis cores; Distributed memory structure, the peripheral unit integrates a multi-bank double-buffered memory controller and interacts with the external memory through DMA.
[0008] Furthermore, the configurable crossbar switch includes: a routing crossbar switch for controlling data flow; and a functional crossbar switch for configuring computing functions.
[0009] Furthermore, the multifunctional operation unit supports vectorized SIMD arithmetic / logic / storage operations.
[0010] Furthermore, the mapping process of the GP Compiler includes: The front end converts the waveform analysis kernel written in C / C++ into an intermediate representation (IR) through the LLVM framework and extracts the data flow graph (DFG); The back-end GP Accelerator Mapper (the general accelerator mapper, which is the back-end of the GP compiler) uses a hierarchical scheduling strategy: First, we use the Kahn algorithm to perform topological sorting to determine the computation priority. Then, we build a spatiotemporal extended MRRG model and use the A* algorithm improved by the breadth-first algorithm (BFS) to search for paths, thus achieving spatiotemporal mapping from DFG nodes to different tiles of the GP Accelerator. The A* (A Star) algorithm is a heuristic search algorithm widely used in path optimization. Its unique feature is that it incorporates global information when examining each possible node on the shortest path. It estimates the distance from the current node to the destination and uses this information as a measure of the likelihood that the node is on the shortest path.
[0011] When resources conflict, the compiler expands the time dimension resources by dynamically increasing the startup interval (II) to achieve collaborative optimization of computing / storage / communication resources.
[0012] Furthermore, the FPGA of each data acquisition board integrates: Data acquisition chain: connect the interface driver module, low-pass filter, calibration module, trigger module, and trigger delay module in sequence; Accelerator controller: Works with the GP Accelerator's double-buffering mechanism to request computational data from external memory through the memory controller to execute the waveform analysis kernel.
[0013] Furthermore, the double buffering mechanism is implemented through FPGA block memory (BRAM), the computing unit array adopts DSP48E to perform arithmetic operations, the configurable crossbar switch is implemented using a lookup table (LUT), and the configuration memory consumes flip-flop (FF) resources.
[0014] Furthermore, the time multiplexing mechanism of the GP Accelerator enables fixed hardware resources to dynamically host different cores, and its compilation time is linearly related to the number of cores.
[0015] The beneficial effects of the present invention are: 1. The GP Accelerator of this invention is a general-purpose waveform accelerator that supports automated acceleration of built-in and user-defined waveform analysis kernels without any manual design, thereby significantly improving waveform analysis development efficiency. Furthermore, the GPAccelerator is distributed across various acquisition boards, providing sufficient computing power for various waveform analysis kernels while avoiding the need to transmit large amounts of data to the main control board, thereby significantly improving service efficiency.
[0016] 2、The GP Accelerator of the present application completely eliminates the "development" link in the traditional sense through its coarse-grained architecture to predefine the computing unit and interconnection reconstruction mechanism. The mapping of any built-in or user-defined kernel only needs to use the GPCompiler to reconfigure the function and interconnection relationship of the computing unit according to the kernel data flow graph (DFG), thereby avoiding the iterative synthesis and physical design phase of the FPGA.
[0017] 3、The compilation time of the GP Compiler of the present application is in linear proportion to the number of kernels, which reflects its low complexity by several orders of magnitude compared with the FPGA tool chain under the transistor-level reconstruction burden. Finally, the GP Compiler exhibits significant superiority in total development efficiency over the labor-intensive Verilog workflow and the intermediate solution of HLS through the combination of zero design investment and minimum compilation overhead.
[0018] 4、The kernel of the present application is directly executed on the on-board accelerator of the acquisition board, and the data localization processing eliminates the PXIe transmission delay and avoids cross-accelerator resource competition. Through the decentralization of computation and the optimization of data locality, the GP Accelerator achieves high throughput regardless of the size of the load, and finally leads the CPU and the FSAccelerato by more than 27.5 times. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is the waveform analysis architecture of the industrial waveform recorder MR-6000; Figure 2 is a schematic diagram of the GP Accelerator of the present application distributed in each acquisition board; Figure 3 is the GP Accelerator of the present application and its corresponding GP Compiler; wherein, Figure 3 a is the GPAccelerator architecture diagram; Figure 3 b is the workflow diagram of the GP compiler; Figure 4 is a development time comparison diagram of the GPCompiler of the present application compared with traditional Verilog and HLS; Figure 5 is a service time comparison diagram of the GPAccelerator of the present application compared with the FSAccelerator and the RK3588 CPU; Figure 6 is a physical schematic diagram of the present application; wherein Figure 6 a is a waveform recorder system diagram; Figure 6b shows that the FPGA on each data acquisition board is distributedly integrated with the GP Accelerator; Figure 6 c is the FPGA internal controller architecture diagram. DETAILED DESCRIPTION
[0020] Specific implementation method 1: This implementation method proposes a distributed general waveform analysis architecture. In order to cope with the challenges of the existing centralized FS Accelerator waveform analysis architecture, a distributed GP Accelerator waveform analysis architecture is proposed, such as Figure 2 As shown in the figure, the FS Accelerator is removed from the main control board and replaced with a GP Accelerator distributed across the data acquisition boards. In conjunction with the corresponding GP Compiler, any built-in or user-defined waveform analysis kernel can be quickly and automatically mapped to the GP Accelerator for execution. Furthermore, since the GP Accelerator is deployed on each type of data acquisition board, waveform analysis can fully utilize local computing power, reducing PXIe data transmission and preventing computing resource competition between waveform analysis kernels, thereby significantly improving service efficiency.
[0021] GP Accelerator is essentially a coarse-grained reconfigurable architecture (CGRA), which pre-integrates the fine-grained hardware resources (LUT, FF, DSP, BRAM) of FPGA into a programmable coarse-grained computing unit array. Figure 3 As shown in Figure 1, the architecture utilizes a three-layer structure: compute units + interconnection network + distributed memory. Each compute unit (tile) includes routing registers, a configurable crossbar switch (routing crossbar switch and function crossbar switch), a multi-function arithmetic unit (supporting vectorized SIMD arithmetic, logic, and storage operations), and a loop execution control module. By dynamically configuring the arithmetic functions and interconnection directions of each unit, this architecture enables flexible data flow execution modes for different waveform analysis cores. Specifically, the peripheral unit integrates a multi-bank double-buffered memory controller, interacting with external memory via DMA, effectively minimizing data transfer latency. The entire architecture utilizes a time-multiplexing mechanism, enabling fixed hardware resources to be dynamically mapped to any waveform analysis core, significantly improving development efficiency while maintaining acceleration performance comparable to that of the FS Accelerator.
[0022] The GP Compiler automates the mapping from the high-level language of the waveform analysis kernel to the GP Accelerator. Its core lies in a heuristic scheduling algorithm based on the Modular Routing Resource Graph (MRRG). The front-end uses the LLVM framework to convert waveform analysis kernels written in C / C++ into an intermediate representation (IR) and extract a data flow graph (DFG). The back-end GP AcceleratorMapper implements a hierarchical scheduling strategy: first, topological sorting is performed using the Kahn algorithm to determine computational priorities. Next, a spatiotemporal MRRG model is constructed, and a path search is performed using the A* algorithm, modified from the breadth-first search (BFS). This achieves the spatiotemporal mapping of DFG nodes to different tiles of the GP Accelerator. When resource conflicts arise, the compiler dynamically increases the initiation interval (II) to expand the time dimension, enabling the coordinated optimization of compute, storage, and communication resources. This compiler overcomes the limitations of traditional HLS tools in fine-grained resource scheduling. By transforming the mapping problem into a constrained graph embedding optimization problem, the hardware implementation time of complex waveform analysis kernels is reduced from manual weeks to minutes, significantly improving waveform analysis development efficiency.
[0023] In terms of development efficiency, this paper compares the development time of GP Compiler with the Verilog and High-Level Synthesis (HLS) methods used for FSAccelerator development in traditional waveform analysis systems. In terms of service efficiency, this paper compares the service time (TTS) of GP Accelerator with the FSAccelerator or CPU (Rockchip RK3588) methods in traditional waveform analysis systems. The results are as follows: Figure 4 shown.
[0024] Figure 4Experimental results show that as the number of waveform analysis cores increases, the total development time scalability of different methods varies significantly. Both the Verilog and HLS methods exhibit exponential growth, while the GP Compiler maintains a linear scaling trajectory. This discrepancy stems from fundamental differences in resource management paradigms: Verilog and HLS operate on the fine-grained logic level of the FPGA, and adding cores intensifies resource contention. When the number of cores exceeds a critical threshold (12 for HLS and 16 for Verilog), the FPGA compilation process enters a phase of explosive latency growth due to routing congestion and placement failures, which manifests as a steep inflection point on the time curve. It is worth noting that HLS's time overhead exceeds that of Verilog when there are more than 12 cores. This is because its automated RTL conversion generates redundant logic, which in turn exacerbates resource competition and adds to the complexity of placement. In sharp contrast, GP Accelerator completely eliminates the traditional "development" link through its coarse-grained architecture, pre-defined computing units and interconnection reconstruction mechanism. The mapping of any built-in or user-defined core only requires GP Compiler to reconfigure the functionality and interconnection relationships of the computing units according to the core data flow graph (DFG), thereby circumventing the iterative synthesis and physical design stages of the FPGA. As a result, GP Compiler's compilation time scales linearly with the number of cores, reflecting its complexity, which is several orders of magnitude lower than the FPGA tool chain under the burden of transistor-level reconstruction. Ultimately, GP The Compiler demonstrates significant superiority in overall development efficiency over labor-intensive Verilog workflows and HLS intermediate solutions through the combined advantages of zero design investment and minimized compilation overhead.
[0025] like Figure 5Figure 2 shows the time-to-service (TTS) trends for three processing platforms (RK3588 CPU, FS Accelerator, and GP Accelerator) as the number of waveform analysis cores increases from 4 to 32. Limited by the parallel processing capabilities of its fixed quad-core architecture, the CPU's TTS increases linearly, from 7.9 seconds with 4 waveform analysis cores to 63.2 seconds with 32 cores. In stark contrast, the FS Accelerator's performance deteriorates exponentially, with its TTS rising sharply from 0.82 seconds with 4 waveform analysis cores to over 52.4 seconds with 32 cores. This catastrophic degradation stems from FPGA resource contention and frequency limitations. When the number of cores reaches 32, routing congestion during the compilation phase causes the accelerator frequency to plummet to 10 MHz, completely eroding its computational advantage and bringing its performance closer to that of the CPU baseline. Although the GP Accelerator was slightly inferior to the FS Accelerator when using four waveform analysis cores (0.93 seconds vs. 0.82 seconds), its TTS demonstrated excellent stability across the full core count range, maintaining a mere 2.3 seconds even when loaded with 32 waveform analysis cores. This resilience is due to its distributed architecture design: cores execute directly on the acquisition board's accelerator, and data localization eliminates PXIe transmission latency while avoiding cross-accelerator resource competition. By decentralizing computation and optimizing data locality, the GP Accelerator achieves high throughput regardless of workload size, ultimately maintaining a lead of over 27.5 times over both the CPU and FS Accelerator.
[0026] Specific implementation method 2: This implementation method proposes a waveform recorder system, such as Figure 6 As shown in Figure a, the system consists of a PXIe chassis, a display unit, a main control board, and 14 data acquisition boards. The main control board is interconnected with multiple data acquisition boards via the PXIe bus, and the GP Accelerator processing results of each acquisition board are transmitted back to the main control board via the XMDA module. Figure 6 As shown in Figure 2b, the FPGA (model XC7K325T-2FFG900I) on each data acquisition board is distributedly integrated with a GP Accelerator, providing sufficient computing power for waveform analysis. Calculations are performed directly on the data acquisition board without the need to request data from the main control board, significantly improving the efficiency of waveform analysis services. Figure 6Figure c details the FPGA's internal controller architecture: Data acquired by the analog-to-digital converter (ADC) is first received and controlled by the interface driver module, then processed by the low-pass filter (LPF) module for noise suppression and calibration. Trigger conditions configured by the user through the GUI (such as a specific voltage threshold) are transmitted via PXIe to the trigger module. The trigger delay module verifies whether the calibrated signal meets the preset conditions and initiates data storage only when the trigger conditions are met. To handle high-frequency data streams, the calibrated signal is downsampled and envelope extracted by the DDR controller before being stored in external memory.
[0027] The accelerator controller works in conjunction with the GP Accelerator's double-buffering mechanism, requesting computational data from external memory via the memory controller to execute the waveform analysis kernel. The processed results are then transmitted to the main control board CPU via the XMDA module for graphical rendering in the user interface. The resource utilization report provided by the FPGA toolchain reveals the implementation details of the GP Accelerator: the double-buffering mechanism uses FPGA block memory (BRAM), the computational unit array employs DSP48Es for arithmetic operations, the configurable crossbar switch is implemented using lookup tables (LUTs), and the configuration memory consumes flip-flop (FF) resources. This architecture, through tight hardware and software collaboration, enables concurrent execution of data acquisition, real-time built-in / custom waveform analysis, and visualization.
[0028] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement of the above embodiments made according to the technical essence of the present invention, within the spirit and principles of the present invention, without departing from the content of the technical solution of the present invention, shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A distributed general waveform analysis architecture, characterized in that: include: Main control board, integrated GP Compiler; Multiple data acquisition boards, each with a distributed and independently running GP Accelerator; The GP Compiler automatically maps the high-level language description of the waveform analysis kernel to the hardware configuration of the GP Accelerator; wherein the waveform analysis kernel is executed locally by the GP Accelerator on the data acquisition board, avoiding the transmission of the original data to the main control board.
2. A distributed universal waveform analysis architecture according to claim 1, characterized in that: The GPAccelerator adopts a three-layer structure, namely: The computing unit structure includes routing registers, configurable crossbar switches, multi-function arithmetic units, and loop execution control modules; The interconnected network structure dynamically configures the computing functions and interconnection directions of each unit to achieve data flow execution modes for different waveform analysis cores; Distributed memory structure, the peripheral unit integrates a multi-bank double-buffered memory controller and interacts with the external memory through DMA.
3. A distributed universal waveform analysis architecture according to claim 2, characterized in that: The configurable crossbar switch comprises: Routing crossbar switch, used to control data flow; Function crossbar switch, used to configure computing functions.
4. A distributed universal waveform analysis architecture according to claim 2, characterized in that: The multifunctional operation unit supports vectorized SIMD arithmetic / logic / storage operations.
5. A distributed universal waveform analysis architecture according to claim 1, characterized in that: The mapping process of the GPCompiler includes: The front end converts the waveform analysis kernel written in C / C++ into an intermediate representation through the LLVM framework and extracts the data flow graph; The backend GP Accelerator Mapper uses a hierarchical scheduling strategy: First, the Kahn algorithm is used to perform topological sorting to determine the calculation priority. Then, a spatiotemporal extended MRRG model is constructed. The A* algorithm improved by the breadth-first algorithm is used for path search to achieve spatiotemporal mapping from DFG nodes to different tiles of the GP Accelerator. When resources conflict, the compiler expands the time dimension resources by dynamically increasing the startup interval to achieve coordinated optimization of computing / storage / communication resources.
6. A distributed universal waveform analysis architecture according to claim 1, characterized in that: The FPGA of each data acquisition board integrates: Data acquisition chain: connect the interface driver module, low-pass filter, calibration module, trigger module, and trigger delay module in sequence; Accelerator controller: Works with the GP Accelerator's double-buffering mechanism to request computational data from external memory through the memory controller to execute the waveform analysis kernel.
7. A distributed universal waveform analysis architecture according to claim 6, characterized in that: The double buffer mechanism is implemented through FPGA block memory, the computing unit array uses DSP48E to perform arithmetic operations, the configurable crossbar switch is implemented using a lookup table, and the configuration memory consumes trigger resources.
8. A distributed universal waveform analysis architecture according to claim 1, characterized in that: The time multiplexing mechanism of the GPAccelerator enables fixed hardware resources to dynamically host different cores, and its compilation time is linearly related to the number of cores.