FPGA-based device fault diagnosis and monitoring hardware acceleration method and device

By optimizing the data flow architecture and algorithm structure of FPGA, and adopting row buffering, window buffering, loop swapping and double buffering techniques, the problems of resource optimization and parallel computing in FPGA fault diagnosis are solved, and efficient and real-time fault diagnosis acceleration is achieved.

CN120849126BActive Publication Date: 2025-12-05HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511341961.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-12-05
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Existing FPGAs suffer from problems such as long development cycles, difficulty in resource optimization, and lack of dedicated hardware architecture design in equipment fault diagnosis, making it difficult to fully leverage their parallel computing advantages.

Method used

The data flow architecture is optimized by using row caching and window caching techniques, data dependencies are eliminated by refactoring the loop structure, a parallel computing unit array is designed, and double buffering technology is used to realize the parallel execution of data transmission and computation, thus constructing an FPGA-based fault diagnosis hardware accelerator.

Benefits of technology

It significantly improves the performance of FPGA fault diagnosis system, realizing efficient, real-time, and low-power edge fault diagnosis, achieving a speedup of tens of times compared to traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849126B_ABST
    Figure CN120849126B_ABST
Patent Text Reader

Abstract

This application relates to a hardware acceleration method and apparatus for device fault diagnosis and monitoring based on FPGA. The method loads the parameters of a trained fault diagnosis model into a high-level synthesis tool to construct the basic hardware architecture; it optimizes data flow using row caching and window caching techniques, and achieves parallel data extraction through shift registers, significantly reducing memory accesses and pipeline startup intervals; it achieves computational parallelization through loop structure reconstruction and dependency optimization, eliminates data dependencies using loop swapping techniques, and designs a parallel computing unit array to improve processing efficiency; and it employs double buffering technology to achieve parallel execution of data transmission and computation, further improving system throughput. This method, through a four-step progressive optimization strategy, solves key technical problems of FPGA neural network accelerators, achieves significant performance improvements, provides a complete solution for the hardware deployment of fault diagnosis neural networks, and has significant engineering application value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of hardware acceleration, and relates to a device fault diagnosis and monitoring hardware acceleration method and device based on FPGA. BACKGROUND

[0002] With the continuous improvement of modern industrial automation, various key devices play an increasingly important role in the production process. Unexpected device failure not only causes production line downtime, resulting in huge economic losses, but also may cause safety accidents and threaten personnel safety. Therefore, timely and accurate diagnosis of device failure and implementation of preventive maintenance have become the core demand of modern industrial management. The traditional device maintenance mode mainly relies on periodic maintenance and post-failure maintenance. This passive maintenance method not only has high cost, but also cannot effectively prevent sudden failures. In contrast, the predictive maintenance technology based on real-time monitoring and intelligent diagnosis can identify device abnormalities before failure occurs and arrange maintenance plans in advance, thereby significantly reducing maintenance costs and improving device availability and production efficiency.

[0003] At present, device fault diagnosis and monitoring technology mainly includes two categories of traditional methods based on signal processing and intelligent methods based on machine learning. Traditional fault diagnosis methods mainly rely on signal processing techniques such as frequency domain analysis, time domain analysis, and wavelet transform, and determine device status by extracting characteristic parameters of vibration, temperature, current, and other sensor signals. This method has the advantages of solid theoretical foundation and strong interpretability, but has limitations in dealing with complex working conditions and multi-variable coupling problems, and is difficult to adapt to the increasingly complex operating environment of modern devices. In recent years, with the rapid development of artificial intelligence technology, machine learning and deep learning-based fault diagnosis methods have gradually become a research hotspot. This method can automatically learn complex patterns in data and perform well in handling nonlinear and multi-dimensional fault features.

[0004] With the development of industrial Internet of Things and edge computing technology, deploying fault diagnosis algorithms on edge devices has become an inevitable trend. Edge deployment can realize the on-site processing of data, reduce network transmission delay, improve the real-time performance and reliability of the system, and reduce the dependence on cloud computing resources. Among the many edge computing platforms, field programmable gate array (FPGA) is attracting attention due to its unique technical characteristics. Compared with traditional CPU and GPU, FPGA has the following advantages: first, FPGA has high parallel processing capability. It contains a large number of configurable logic units inside, which can realize true parallel computing, especially suitable for processing matrix operations, convolution operations and other computationally intensive tasks in signal processing and machine learning algorithms. Second, FPGA has low power consumption. Compared with high-performance computing platforms such as GPU, FPGA provides comparable computing power while significantly reducing power consumption, which is particularly important for resource-constrained edge devices. Third, FPGA has deterministic real-time response capability. Due to its hardware implementation characteristics, FPGA can provide predictable delay and deterministic execution time, meeting the strict real-time requirements of industrial applications. Finally, FPGA has flexible reconfigurable characteristics, which can dynamically adjust the hardware architecture according to different application requirements, realizing the rapid iteration and optimization of algorithms.

[0005] Although FPGA has obvious advantages in edge fault diagnosis applications, how to fully utilize its parallel computing capability still faces many challenges. Traditional FPGA development methods are mainly based on hardware description language (HDL), with long development cycles, high technical threshold, and difficulty in quickly adapting to the evolving fault diagnosis algorithms. In terms of resource optimization, FPGA's logic resources, storage resources, and DSP resources are limited, and how to achieve efficient implementation of algorithms under resource constraints is another important challenge. This requires finding the best balance point between computing accuracy, processing speed, and resource consumption. In addition, existing technologies lack specialized hardware architecture design for fault diagnosis application characteristics. Fault diagnosis algorithms usually involve a combination of various signal processing and machine learning techniques, requiring the design of hardware architecture that can efficiently support these hybrid computing modes. SUMMARY

[0006] To solve the problems in the above traditional methods, the present application proposes a FPGA-based device fault diagnosis and monitoring hardware acceleration method and device, which can fully utilize the parallel computing advantages of FPGA and realize an efficient, real-time, and low-power edge fault diagnosis system.

[0007] To achieve the above purpose, the embodiments of the present application adopt the following technical solutions:

[0008] On the one hand, a FPGA-based device fault diagnosis and monitoring hardware acceleration method is provided, which comprises the following steps:

[0009] Step 1: Load the trained fault diagnosis neural network model parameters into the high-level synthesis tool in the form of a header file, build a basic hardware architecture based on FPGA, and obtain a baseline hardware accelerator.

[0010] Step 2: Optimize the data flow architecture of the baseline hardware accelerator using line buffering and window buffering techniques to obtain a data flow optimized hardware architecture.

[0011] Step 3: On the basis of the data flow optimized hardware architecture, reconstruct the neural network algorithm using a loop structure to eliminate data dependencies in the calculation, and design a parallel computing unit array to realize simultaneous calculation of multiple output neurons.

[0012] Step 4: Use double buffering technology to realize parallel execution of data transmission and calculation for the hardware architecture optimized in step 3, further improving the overall throughput of the system.

[0013] On the other hand, a FPGA-based device fault diagnosis and monitoring hardware acceleration device is also provided, comprising steps:

[0014] A basic hardware architecture construction module is used to load the trained fault diagnosis neural network model parameters into the high-level synthesis tool in the form of a header file, build a basic hardware architecture based on FPGA, and obtain a baseline hardware accelerator.

[0015] A data flow optimization module is used to optimize the data flow architecture of the baseline hardware accelerator using line buffering and window buffering techniques to obtain a data flow optimized hardware architecture.

[0016] An algorithm structure reconstruction and efficient parallelization module is used to reconstruct the neural network algorithm using a loop structure to eliminate data dependencies in the calculation on the basis of the data flow optimized hardware architecture, and design a parallel computing unit array to realize simultaneous calculation of multiple output neurons.

[0017] A system throughput improvement module is used to realize parallel execution of data transmission and calculation for the hardware architecture optimized by the algorithm structure reconstruction and efficient parallelization module using double buffering technology, further improving the overall throughput of the system.

[0018] One of the above technical solutions has the following advantages and beneficial effects:

[0019] The FPGA-based device fault diagnosis and monitoring hardware acceleration method and device load the trained fault diagnosis model parameters into a high-level synthesis tool to construct a basic hardware architecture, use row buffer and window buffer technology to optimize data flow, realize data parallel extraction through a shift register, significantly reduce memory access times and pipeline start interval, realize calculation parallelization through loop structure reconstruction and dependency optimization, eliminate data dependency using loop exchange technology, design a parallel computing unit array to improve processing efficiency, and use double buffering technology to realize parallel execution of data transmission and calculation, further improving system throughput. The method solves the key technical problems of FPGA neural network accelerator through a four-step progressive optimization strategy, realizes significant performance improvement, provides a complete solution for fault diagnosis neural network hardware deployment, and has important engineering application value. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] Figure 1 A flowchart of the FPGA-based device fault diagnosis and monitoring hardware acceleration method in one embodiment;

[0022] Figure 2 A specific flowchart of the FPGA-based device fault diagnosis and monitoring hardware acceleration method in one embodiment. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing the specific embodiments and are not intended to limit the present application.

[0025] It is noted that reference herein to "embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. Those skilled in the art will recognize that embodiments described herein can be combined with other embodiments in accordance with the application. As used herein the term "and / or" means any combination of one or more of the associated listed items and includes all possible combinations.

[0026] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0027] In one embodiment, as shown in Figure 1 a FPGA-based device fault diagnosis and monitoring hardware acceleration method can include the following processing steps S1 to S4:

[0028] Step 1: Load the trained fault diagnosis neural network model parameters into the high-level synthesis tool in the form of a header file, build a basic hardware architecture based on FPGA, and get a baseline hardware accelerator.

[0029] Specifically, the trained fault diagnosis model parameters are loaded into the high-level synthesis tool in the form of a header file, and a basic hardware architecture including data preprocessing, feature extraction, and classification decision is built.

[0030] The high-level synthesis tool can be, but is not limited to, Vitis HLS.

[0031] The baseline hardware accelerator refers to an acceleration system that is preliminarily implemented on the basis of the hardware architecture and has not been deeply optimized for hardware characteristics or algorithm characteristics, and is mainly used as a comparison benchmark for subsequent performance optimization. The hardware accelerator refers to a special acceleration unit implemented on a platform such as FPGA, which is used for efficient processing of specific tasks.

[0032] The basic hardware architecture is the underlying structure and resource configuration required to implement the accelerator, including the division of each functional module, the data flow mode, and the necessary hardware resources (such as ARM core, FPGA, AXI bus, etc.), which is the basic platform for all acceleration optimization.

[0033] In short, the basic hardware architecture refers to the underlying hardware platform and resource layout, and is not concerned with the specific implementation of functions, while the baseline hardware accelerator is the overall hardware system designed to implement basic diagnosis acceleration, which can include multiple modules or even cover the entire FPGA resource.

[0034] Step 1 realizes the loading of model parameters and the construction of basic hardware architecture. Specifically, first, the parameters of the fault diagnosis neural network model trained with PyTorch are converted into a format suitable for FPGA loading and imported into the FPGA platform through the Vitis HLS tool. The system basic architecture adopts a pipeline design, which is divided into three main modules: data preprocessing, feature extraction, and classification decision, respectively responsible for signal processing, feature extraction, and final fault type judgment.

[0035] Step 2: The data flow architecture of the baseline hardware accelerator is optimized using row buffer and window buffer technology to obtain the optimized hardware architecture.

[0036] Specifically, to solve the data access bottleneck problem in step 1, the row buffer and window buffer technology are used to optimize the data flow architecture of the baseline hardware accelerator. Through the use of shift registers, data parallel extraction is achieved, significantly reducing the number of memory access and pipeline start interval.

[0037] Step 2 realizes the optimization of the data flow architecture. Specifically, to solve the data access bottleneck problem in the convolution operation, the system introduces row buffer and window buffer technology. These cache mechanisms can significantly reduce the number of external memory accesses and improve data processing efficiency and throughput by parallel extraction of data windows.

[0038] Step 3: Based on the optimized hardware architecture of the data flow, the neural network algorithm is reconstructed using a loop structure to eliminate data dependency in the calculation, and a parallel computing unit array is designed to realize the simultaneous calculation of multiple output neurons.

[0039] Specifically, based on step 2, algorithm loop structure reconstruction and calculation dependency optimization are used to achieve efficient parallelization of the computing unit. The loop exchange technique is used to eliminate data dependencies, and a parallel computing unit array is designed to improve processing efficiency.

[0040] The optimized hardware architecture of the data flow is not the same as the baseline hardware accelerator. Each optimization of the algorithm or system will bring corresponding changes to the hardware architecture, which is a significant feature of the flexibility and reconfigurability of FPGA. That is, each optimization will form a new hardware architecture, and the next optimization will continue on the basis of the hardware architecture obtained from the previous optimization. Correspondingly, each optimization will produce a new hardware accelerator with optimized functions, achieving more efficient acceleration. Therefore, the baseline hardware accelerator is only the initial implementation, and each optimized accelerator and hardware architecture after each optimization can also be used as a stage of achievement for comparison and analysis.

[0041] Step 3: Algorithm structure reconstruction and efficient parallelization are realized on the basis of the hardware architecture after data stream optimization. Specifically, on the basis of data stream optimization, the calculation cycle structure of the neural network is optimized to eliminate the data dependency between the calculation units. At the same time, a parallel calculation unit is designed, so that the calculation of multiple neurons can be performed simultaneously, greatly improving the calculation speed of core components such as fully connected layers.

[0042] Step 4: Double buffering technology is used to realize the parallel execution of data transmission and calculation on the hardware architecture optimized in step 3, further improving the overall throughput of the system.

[0043] Specifically, double buffering technology is used to realize the parallel execution of data transmission and calculation, and time overlap is achieved through ping-pong operation mode and intelligent scheduling control, further improving the system throughput.

[0044] Step 4 uses double buffering technology to improve system throughput. Specifically, a double buffering storage architecture is used, and through "ping-pong" switching mode, data transmission and calculation can be performed simultaneously, further improving the overall processing speed and throughput capacity of the system.

[0045] In one specific embodiment, the specific process is as shown in Figure 2 The original vibration signal data of the faulty motor is first collected, and the vibration sensor continuously collects the vibration signal during the operation of the motor at a sampling frequency of 10 kHz. The collected time domain vibration signal is converted into SDP image format, and the image size is 224x224 pixels, which is used as the input feature of the fault diagnosis neural network. A convolutional neural network model is constructed under the PyTorch framework, and the training data set contains normal, rotor bar fracture fault, bow fault, bearing fault, rotor misalignment fault, rotor imbalance fault, single-phase open circuit fault, etc. 8 kinds of motor states. After the model training is completed, the trained fault diagnosis model is first deployed to the ARM processor of the PYNQ platform for verification, and the single diagnosis time on the ARM Cortex-A9 processor is 210.19 ms, which verifies the effectiveness of the model.

[0046] The FPGA-based device fault diagnosis and monitoring hardware acceleration method loads the trained fault diagnosis model parameters into a high-level synthesis tool, constructs a basic hardware architecture, optimizes data flow by using line buffer and window buffer technology, realizes data parallel extraction through a shift register, significantly reduces the number of memory access times and pipeline start intervals, realizes calculation parallelization through loop structure reconstruction and dependency optimization, eliminates data dependency by using loop exchange technology, designs a parallel computing unit array to improve processing efficiency, and realizes parallel execution of data transmission and calculation by using double buffering technology, and further improves the system throughput. The method solves the key technical problems of FPGA neural network accelerator through a four-step progressive optimization strategy, realizes significant performance improvement, provides a complete solution for hardware deployment of fault diagnosis neural network, and has important engineering application value.

[0047] In one embodiment, the trained fault diagnosis neural network model parameters include: convolution layer weight matrix, fully connected layer weight vector, bias parameter and batch normalization parameter; step 1 includes: converting the trained fault diagnosis neural network model parameters into single-precision floating-point format and packaging them into a header file, and loading them into Vitis HLS through the header file; a pipeline design is used to construct a basic FPGA hardware implementation architecture; the basic FPGA hardware implementation architecture includes: a data preprocessing unit, a feature extraction unit and a classification decision unit; the data preprocessing unit is responsible for input signal normalization and format conversion; the feature extraction unit extracts fault features through convolution calculation and pooling operation; the classification decision unit realizes fault type discrimination through a fully connected layer; each computing unit is implemented by using Xilinx floating-point IP core, and the IP core is imported into Vivado for circuit design, and the obtained bitstream file is loaded into PYNQ; the Xilinx floating-point IP core includes a floating-point multiplier, an accumulator and an activation function calculation unit.

[0048] In one embodiment, the convolution kernel weight adopts a four-dimensional tensor form is stored in a memory, wherein is the number of output channels, is the number of input channels, and are the height and width of the convolution kernel, respectively.

[0049] The fully connected layer weight adopts a two-dimensional matrix form is stored in a memory, wherein is the number of output neurons, is the input feature dimension.

[0050] Specifically, step 1 specifically includes:

[0051] Step 1.1: The fault diagnosis model parameters are loaded into Vitis HLS in the form of a header file;

[0052] The trained fault diagnosis neural network model parameters, including the convolution layer weight matrix, fully connected layer weight vector, bias parameter, and batch normalization parameter, are converted into single-precision floating-point format and packaged into a header file. The parameter organization stores the convolution kernel weights in the form of a four-dimensional tensor , where is the number of output channels, is the number of input channels, and are the height and width of the convolution kernel, respectively. The fully connected layer weights are stored in the form of a two-dimensional matrix , where is the number of output neurons, is the input feature dimension.

[0053] Specifically, the first layer convolution layer weight tensor size is W

[32] [3][3][3], the second layer is W

[64]

[32] [3][3], the third layer is W

[128]

[64] [3][3], and the fourth layer is W

[256]

[128] [3][3]. The fully connected layer weight matrix sizes are W

[512]

[2048] and W [8]

[512] , corresponding to the final 8-class fault classification output.

[0054] Step 1.2: Build the basic FPGA hardware implementation architecture to obtain the baseline hardware accelerator;

[0055] The basic hardware architecture adopts a pipeline design, including data preprocessing unit, feature extraction unit, and classification decision unit. The data preprocessing unit is responsible for input signal normalization and format conversion; the feature extraction unit extracts fault features through convolution calculation and pooling operation; the classification decision unit realizes fault type discrimination through fully connected layers. Each calculation unit adopts Xilinx floating-point IP core implementation, including floating-point multiplier, accumulator, and activation function calculation unit, and the IP core is imported into Vivado for circuit design. The obtained bitstream file is loaded into PYNQ, and the fault diagnosis model is deployed in the PL of PYNQ, i.e., the FPGA part. The time performance analysis of the baseline architecture is as follows:

[0056] Convolution layer calculation time:

[0057]

[0058] Fully connected layer computation time:

[0059]

[0060] Total computation delay:

[0061]

[0062] where, is the output feature map size, is the multiplier delay, is the accumulator delay, is the adder delay.

[0063] The baseline hardware accelerator implemented in the PL part of the PYNQ platform has a single fault diagnosis time of 12.94 ms, which is 16.24 times faster than the ARM processor of 210.19 ms, verifying the correctness and preliminary acceleration effect of the hardware implementation.

[0064] In one embodiment, step 2 includes: using a row buffer mechanism to store continuous K row input data; wherein K is the convolution kernel height; the row buffer mechanism is used to implement a shift register structure, when new data is input, the oldest data is removed, realizing the pipeline update of data; on the basis of the row buffer, a window buffer is constructed, realizing K × K parallel extraction of convolution window data; wherein the window buffer is implemented by a shift register array, which can output K × K data elements simultaneously for convolution calculation at each clock cycle; the update strategy of the window buffer uses a sliding window mechanism, and when moving horizontally by one pixel position each time, only K new data elements need to be updated.

[0065] Specifically, step 2 specifically includes:

[0066] Step 2.1: row buffer optimization;

[0067] In view of the frequent memory access problem in convolution calculation, a row buffer mechanism is designed to store continuous K row input data, wherein K is the convolution kernel height. The row buffer is implemented by a shift register structure, when new data rows are input, the oldest data rows are removed, realizing the pipeline update of data. The storage capacity of the row buffer is bits, wherein W is the input data width, D is the data bit width. ​​

[0068] The row buffer optimization time analysis is as follows:

[0069] The data access time after optimization is:

[0070]

[0071] The memory access times are reduced from times per pixel to 1 time per pixel.

[0072] In the present embodiment, for a 3x3 convolution kernel, K = 3, W = 224, D = 32, and the row buffer storage capacity is:

[0073]

[0074] The data access time after optimization is:

[0075]

[0076] The memory access times are reduced from times per pixel to 1 time per pixel, and the access efficiency is improved by 9 times.

[0077] Step 2.2: Window buffer optimization.

[0078] The window buffer is constructed on the basis of the row buffer, and the K x K parallel extraction of convolution window data is achieved. The window buffer is implemented by a shift register array, and K x K data elements can be output simultaneously for convolution calculation every clock cycle. The update strategy of the window buffer adopts a sliding window mechanism, and only K new data elements need to be updated when moving horizontally by one pixel position each time.

[0079] Window buffer optimization time analysis:

[0080] The pipeline start interval is optimized from cycles to 1 cycle

[0081] The calculation time of the convolution layer after optimization is:

[0082]

[0083] The time improvement ratio is:

[0084]

[0085] In the present embodiment, the pipeline start interval is reduced from The period is optimized to 1 period, and after step 2 optimization, the fault diagnosis time is reduced from 12.94 ms to 6.18 ms, the data stream optimization effect is remarkable, and the performance is improved by 2.09 times.

[0086] In one embodiment, step 3 comprises: reconstructing the nested order of the full connection layer calculation cycle by using the loop interaction technology to obtain the loop interaction optimization result; and designing a parallel multiplication accumulation unit array based on the loop interaction optimization result to realize simultaneous calculation of multiple output neurons.

[0087] In one embodiment, the parallel degree of the parallel multiplication accumulation unit array is:

[0088] ;

[0089] Wherein, is the parallel degree of the parallel multiplication accumulation unit array, is the number of output neurons, is the number of available DSP resources, is the number of DSP resources required by a single multiplier.

[0090] Specifically, step 3 specifically comprises:

[0091] Step 3.1: Loop exchange optimization strategy;

[0092] In view of the data dependency problem existing in the full connection layer calculation, the nested order of the calculation cycle is reconstructed by using the loop exchange technology. The original full connection layer calculation adopts an output-priority loop structure, that is, the outer loop traverses the output neurons, and the inner loop performs accumulation calculation, resulting in data dependency between adjacent accumulation operations. Time analysis before loop exchange:

[0093] Accumulation operation interval time:

[0094]

[0095] Total time of full connection layer:

[0096]

[0097] Step 3.2: Parallel calculation unit design;

[0098] Based on the loop exchange optimization, a parallel multiplication accumulation unit array is designed to realize simultaneous calculation of multiple output neurons. The selection of parallel degree P needs to balance the calculation performance and hardware resource consumption, and the calculation formula is , wherein is the number of output neurons, is the number of available DSP resources, is the number of DSP resources required by a single multiplier.

[0099] In this embodiment , , Therefore . The time analysis after the loop exchange optimization is as follows:

[0100] The accumulated operation interval time is:

[0101] one clock cycle

[0102] The time of the full connection layer after optimization:

[0103]

[0104] The full connection layer time improvement ratio:

[0105]

[0106] After step 3 optimization, the fault diagnosis time is further reduced from 6.18 ms to 3.42 ms, and the parallelization optimization effect is obvious, and the performance is improved by 1.81 times.

[0107] In one embodiment, step 4 includes: constructing a double buffering storage system, using a double buffering working mechanism for calculation, after the calculation is completed, the roles of the two cache areas are interchanged, realizing the time overlap of data transmission and calculation; the double buffering storage system includes: two groups of cache areas Buffer_A and Buffer_B with the same capacity; the cache capacity of each group is 1 times the size of the input data; the double buffering working mechanism adopts a ping-pong operation mode: when the data in Buffer_A is used for fault diagnosis calculation, Buffer_B simultaneously loads the next batch of data; an intelligent scheduling controller is designed to manage the switching timing of double buffering for pipeline scheduling optimization, so that the data transmission time is optimally matched with the calculation time ; when , the data transmission is completely covered by the calculation time, and the system throughput reaches the theoretical maximum value; when , the system throughput is limited by the data transmission speed.

[0108] Specifically, step 4 specifically includes:

[0109] Step 4.1: Double buffering storage architecture design;

[0110] A double buffering system is constructed, which contains two groups of buffer areas Buffer_A and Buffer_B with the same capacity. The capacity of each group is equal to the size of input data. The double buffering mechanism adopts a ping-pong operation mode: when the data in Buffer_A is used for fault diagnosis calculation, Buffer_B simultaneously loads the next batch of data; after the calculation is completed, the roles of the two buffer areas are interchanged, so as to overlap the data transmission and calculation in time. Time analysis of the system before double buffering,

[0111] The total time of single processing is:

[0112]

[0113] The total time of N times of processing is:

[0114]

[0115] Step 4.2: pipeline scheduling optimization.

[0116] An intelligent scheduling controller is designed to manage the switching sequence of double buffering, so as to ensure the optimal matching between data transmission time and calculation time . When , the data transmission is completely covered by the calculation time, and the system throughput reaches the theoretical maximum value; when , the system throughput is limited by the data transmission speed. Time analysis of the system after double buffering optimization, wherein the single processing time is:

[0117]

[0118] The total time of N times of processing is:

[0119]

[0120] The double buffering time improvement ratio is:

[0121]

[0122] The overall system optimization effect analysis is as follows: wherein the overall time improvement ratio is:

[0123]

[0124] When N is large enough:

[0125]

[0126] The total improvement ratio can reach:

[0127]

[0128] In this embodiment, the data transmission time , the calculation time , meets the condition, so . After the double buffering optimization in step 4, the final fault diagnosis time is optimized to 2.24ms, which realizes a performance improvement of 1.53 times compared with 3.42ms in step 3.

[0129] Through the four-step optimization strategy of the application, the performance of the fault motor diagnosis system is significantly improved. The system starts from the initial single diagnosis time of 210.19ms on the ARM processor, and first reduces the diagnosis time to 12.94ms through the construction of the FPGA hardware baseline accelerator, realizing a preliminary acceleration effect of 16.24 times. Subsequently, through the row buffer and window buffer data stream optimization in step 2, the diagnosis time is further reduced to 6.18ms, which is 2.09 times higher than the FPGA baseline; step 3 adopts the loop exchange and parallel computing unit design, and the diagnosis time is optimized to 3.42ms, which is 1.81 times higher than step 2; finally, through the double buffering technology in step 4, the parallel execution of data transmission and calculation is realized, and the diagnosis time reaches 2.24ms, which is 1.53 times higher than step 3. The overall speedup ratio of the entire optimization process is 93.84 times compared with the ARM processor and 5.78 times compared with the FPGA baseline, successfully verifying the effectiveness and practicability of the optimization method in the fault motor diagnosis hardware acceleration.

[0130] Compared with the prior art, the method has the following advantages:

[0131] (1) The method adopts the row buffer and window buffer data stream optimization technology, effectively solves the traditional BRAM access bottleneck problem through the mixed storage architecture and parallel data access, and significantly improves the data processing parallelism.

[0132] (2) The method proposes a loop reconstruction and parallel computing method, which eliminates data dependency through algorithm-level optimization, greatly improves the utilization rate of the computing unit, and fully utilizes the FPGA parallel computing capability.

[0133] (3) The method innovatively adopts a double buffering mechanism to realize system-level optimization, effectively hides the data transmission delay, and realizes the true parallelism of data transmission and calculation while maintaining the calculation accuracy.

[0134] (4) The method constructs a complete FPGA hardware acceleration architecture, and through the organic combination of multi-level optimization strategies, provides an efficient, real-time and low-power hardware acceleration solution for edge device fault diagnosis, which can obtain tens of times performance improvement compared with the traditional software implementation.

[0135] It should be understood that, although the steps are shown in sequence according to the direction of the arrows, the steps are not necessarily executed in the order of the direction of the arrows. Unless explicitly stated herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other orders. Moreover, at least a part of the steps can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the sub-steps or stages is not necessarily sequential, but can be alternately executed with other steps or at least a part of the sub-steps or stages of other steps.

[0136] In one embodiment, a FPGA-based device fault diagnosis and monitoring hardware acceleration apparatus is also provided, comprising the steps of:

[0137] A basic hardware architecture construction module is configured to load the trained fault diagnosis neural network model parameters into a high-level synthesis tool in the form of a header file, construct a basic hardware architecture based on FPGA, and obtain a baseline hardware accelerator.

[0138] A data flow optimization module is configured to optimize the data flow architecture of the baseline hardware accelerator using line buffering and window buffering techniques, and obtain a hardware architecture after data flow optimization.

[0139] An algorithm structure reconstruction and efficient parallelization module is configured to, on the basis of the hardware architecture after data flow optimization, reconstruct the neural network algorithm using a loop structure to eliminate data dependency in calculation, and design an array of parallel computing units to realize simultaneous calculation of multiple output neurons.

[0140] A system throughput improvement module is configured to, for the hardware architecture optimized by the algorithm structure reconstruction and efficient parallelization module, use a double buffering technique to realize parallel execution of data transmission and calculation, and further improve the overall throughput of the system.

[0141] In one embodiment, the trained fault diagnosis neural network model parameters include: a convolutional layer weight matrix, a fully connected layer weight vector, a bias parameter, and a batch normalization parameter; the basic hardware architecture construction module is further configured to uniformly convert the trained fault diagnosis neural network model parameters into single-precision floating-point format and encapsulate them into a header file, and load them into Vitis HLS in the form of a header file; a basic FPGA hardware implementation architecture is constructed by using a pipeline design; the basic FPGA hardware implementation architecture includes: a data preprocessing unit, a feature extraction unit, and a classification decision unit; the data preprocessing unit is configured to be responsible for normalization and format conversion of input signals; the feature extraction unit is configured to extract fault features through convolutional calculation and pooling operation; the classification decision unit is configured to realize fault type discrimination through a fully connected layer; each calculation unit is implemented by using an Xilinx floating-point IP core, and the IP core is imported into Vivado for circuit design, and a bitstream file obtained is loaded into PYNQ; the Xilinx floating-point IP core includes a floating-point multiplier, an accumulator, and an activation function calculation unit.

[0142] In one embodiment, the convolution kernel weight in the basic hardware architecture construction module adopts a four-dimensional tensor form is stored in a memory, where is the number of output channels, is the number of input channels, and are the height and width of the convolution kernel, respectively.

[0143] The fully connected layer weight adopts a two-dimensional matrix form is stored in a memory, where is the number of output neurons, is the input feature dimension.

[0144] In one embodiment, the data flow optimization module is further configured to store continuous K row input data by using a row buffer mechanism; where K is the convolution kernel height; the row buffer mechanism is configured to be implemented by using a shift register structure, when new data is input, the oldest data is removed, and pipeline updating of the data is realized; a window buffer is constructed on the basis of the row buffer, and K × K parallel extraction of convolution window data is realized; where the window buffer is implemented by using a shift register array, and K × K data elements can be simultaneously output for convolution calculation in each clock cycle; the updating strategy of the window buffer adopts a sliding window mechanism, and when a pixel position is horizontally moved each time, only K new data elements need to be updated.

[0145] In one embodiment, the algorithm structure reconstruction and high-efficiency parallelization module is further configured to reconstruct the nested order of the computation loop of the full connection layer by using a loop interaction technique to obtain a loop interaction optimization result; and based on the loop interaction optimization result, design a parallel multiply-accumulate unit array to realize simultaneous computation of multiple output neurons.

[0146] In one embodiment, the parallel degree of the parallel multiply-accumulate unit array is:

[0147] ;

[0148] wherein, is the parallel degree of the parallel multiply-accumulate unit array, is the number of output neurons, is the number of available DSP resources, is the number of DSP resources required by a single multiplier.

[0149] In one embodiment, the system throughput improvement module is further configured to construct a double-buffering storage system and perform computation by using a double-buffering working mechanism, wherein after the computation is completed, the roles of the two buffer areas are interchanged to realize time overlap of data transmission and computation; the double-buffering storage system comprises two groups of buffer areas Buffer_A and Buffer_B with the same capacity; the buffer capacity of each group is 1 times the size of the input data; the double-buffering working mechanism adopts a ping-pong operation mode: when the data in Buffer_A is used for fault diagnosis computation, Buffer_B simultaneously loads the next batch of data; an intelligent scheduling controller is designed to manage the switching timing of the double buffering for pipeline scheduling optimization, so that the data transmission time is optimally matched with the computation time ; when , the data transmission is completely covered by the computation time, and the system throughput reaches the theoretical maximum value; when , the system throughput is limited by the data transmission speed.

[0150] It can be understood that the specific explanations and descriptions of the FPGA-based device fault diagnosis and monitoring hardware acceleration apparatus can refer to the corresponding explanations and descriptions of the FPGA-based device fault diagnosis and monitoring hardware acceleration method in the above embodiments, and will not be repeated here. Each module in the FPGA-based device fault diagnosis and monitoring hardware acceleration apparatus described above can be realized by software, hardware, and combinations thereof, in whole or in part. Each module described above can be embedded in or independent of a device with data processing function in hardware form, or can be stored in the memory of the device in software form, so that the processor can call and execute the operations corresponding to each module. The device can be, but is not limited to, various types of data processing computer devices known in the art.

[0151] Any technical features in the above embodiments can be combined, and for the sake of brevity, not all possible combinations are described above, however, any combination of the technical features is considered to be within the scope of the present disclosure.

[0152] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the protection scope of the present application. It should be pointed out that, for ordinary skilled persons in the art, some modifications and improvements can be made without departing from the concept of the present application, and all of them belong to the protection scope of the present application.

Claims

1. A FPGA-based device fault diagnosis and monitoring hardware acceleration method, characterized in that, The method comprises the steps of: Step 1: loading the trained fault diagnosis neural network model parameters into the high-level synthesis tool in the form of a header file, constructing a basic hardware architecture based on FPGA, and obtaining a baseline hardware accelerator; Step 2: optimizing the data flow architecture of the baseline hardware accelerator by using line buffering and window buffering technology to obtain a data flow optimized hardware architecture; Step 3: on the basis of the data flow optimized hardware architecture, the neural network algorithm is reconstructed to eliminate data dependency in calculation by using a loop structure, and a parallel computing unit array is designed to realize simultaneous calculation of multiple output neurons; Step 4: using double buffering technology to realize parallel execution of data transmission and calculation on the hardware architecture optimized in step 3, and further improving the overall throughput of the system; Step 1 specifically comprises: The trained fault diagnosis neural network model parameters are converted into single-precision floating-point format and packaged into a header file, which is loaded into Vitis HLS in the form of a header file; A basic FPGA hardware implementation architecture is constructed by using a pipeline design; the basic FPGA hardware implementation architecture includes a data preprocessing unit, a feature extraction unit and a classification decision unit; the data preprocessing unit is responsible for input signal normalization and format conversion; the feature extraction unit extracts fault features through convolution calculation and pooling operation; the classification decision unit realizes fault type discrimination through a fully connected layer; Each computing unit is implemented by using a Xilinx floating-point IP core, and the IP core is imported into Vivado for circuit design, and the obtained bitstream file is loaded into PYNQ; the Xilinx floating-point IP core includes a floating-point multiplier, an accumulator and an activation function calculation unit.

2. The FPGA-based device fault diagnosis and monitoring hardware acceleration method according to claim 1, characterized in that, The convolution kernel weight adopts a four-dimensional tensor form is stored, wherein is the output channel number, is the input channel number, and are the height and width of the convolution kernel respectively. The full connection layer weight adopts a two-dimensional matrix form is stored, wherein is the number of output neurons, is the input feature dimension.

3. The FPGA-based device fault diagnosis and monitoring hardware acceleration method according to claim 1, characterized in that, Step 2 comprises: The continuous input data is stored by using a line buffer mechanism K wherein K is the convolution kernel height; the line buffer mechanism is used to realize a shift register structure, when new data is input, the oldest data is removed, realizing pipeline update of the data A window buffer is constructed based on a line buffer, and a convolution calculation is realized K K Parallel extraction of convolution window data; wherein the window buffer is implemented by a shift register array, and K K data elements can be output simultaneously for convolution calculation in each clock cycle; the update strategy of the window buffer adopts a sliding window mechanism, and only K new data elements need to be updated when moving one pixel position horizontally each time.​​ 4. The FPGA-based device fault diagnosis and monitoring hardware acceleration method according to claim 1, characterized in that, Step 3 comprises: The nested order of the full connection layer calculation loop is reconstructed by using loop interaction technology to obtain a loop interaction optimization result; Based on the loop interaction optimization result, a parallel multiplication accumulation unit array is designed to realize simultaneous calculation of multiple output neurons.

5. The FPGA-based device fault diagnosis and monitoring hardware acceleration method according to claim 4, characterized in that, The parallel degree of the parallel multiplication accumulation unit array is: wherein, is the parallelism of the array of parallel multiply-accumulate units, is the number of output neurons, is the number of available DSP resources, is the number of DSP resources required for a single multiplier.

6. The FPGA-based device fault diagnosis and monitoring hardware acceleration method according to claim 4, characterized in that, Step 4 comprises: A double buffering storage system is constructed, and a double buffering working mechanism is used for calculation; after the calculation is completed, the roles of the two buffer areas are interchanged to realize time overlap of data transmission and calculation; the double buffering storage system includes two groups of buffer areas Buffer_A and Buffer_B with the same capacity; the buffer capacity of each group is 1 times the size of the input data; the double buffering working mechanism adopts a ping-pong operation mode: when the data in Buffer_A is used for fault diagnosis calculation, Buffer_B simultaneously loads the next batch of data; The intelligent dispatching controller is designed to manage the switching timing of double buffering for pipeline scheduling optimization, so that the data transmission time is optimally matched with the calculation time ; when , the data transmission is completely covered by the calculation time, and the system throughput reaches the theoretical maximum; when , the system throughput is limited by the data transmission speed.

7. An FPGA-based device fault diagnosis and monitoring hardware acceleration apparatus, characterized in that, The method comprises the steps of: The baseline hardware architecture construction module is used to load the trained fault diagnosis neural network model parameters into the high-level synthesis tool in the form of a header file, construct a baseline hardware accelerator based on an FPGA, and obtain a baseline hardware accelerator. Specifically, the trained fault diagnosis neural network model parameters are converted into single-precision floating-point format and packaged into a header file, and then loaded into Vitis HLS in the form of a header file. The baseline FPGA hardware implementation architecture is constructed using a pipeline design. The baseline FPGA hardware implementation architecture includes a data preprocessing unit, a feature extraction unit, and a classification decision unit. The data preprocessing unit is responsible for input signal normalization and format conversion. The feature extraction unit extracts fault features through convolution calculation and pooling operations. The classification decision unit implements fault type discrimination through a fully connected layer. Each calculation unit is implemented using an Xilinx floating-point IP core, and the IP core is imported into Vivado for circuit design. The obtained bitstream file is loaded into PYNQ. The Xilinx floating-point IP core includes a floating-point multiplier, an accumulator, and an activation function calculation unit. The trained fault diagnosis neural network model parameters include a convolution layer weight matrix, a fully connected layer weight vector, a bias parameter, and a batch normalization parameter. The data flow optimization module is used to optimize the data flow architecture of the baseline hardware accelerator using row buffer and window buffer techniques to obtain a data flow optimized hardware architecture. The algorithm structure reconstruction and efficient parallelization module is used to eliminate data dependency in the calculation by reconstructing the loop structure of the neural network algorithm based on the data flow optimized hardware architecture, and design a parallel computing unit array to realize the simultaneous calculation of multiple output neurons. The system throughput improvement module is used to realize the parallel execution of data transmission and calculation by using double buffering technology for the hardware architecture optimized by the algorithm structure reconstruction and efficient parallelization module, and further improve the overall system throughput.

Citation Information

Patent Citations

  • FPGA fault diagnosis accelerator design method and system

    CN119397235A

  • Equipment real-time fault diagnosis method based on FPGA and SNN

    CN120213431A