Partial dynamic reconstruction componentization method based on distributed FPGA (Field Programmable Gate Array)
By introducing distributed architecture and deep learning models into the FPGA architecture, combined with dynamic partial reconfigurable technology, the uneven resource allocation and performance degradation problems of traditional FPGA architecture under long-term high load operation are solved, and the stability and efficiency of the system are achieved.
Patent Information
- Application Number
- CN202510091569.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-21
AI Technical Summary
Traditional FPGA architectures have problems such as uneven resource allocation, degradation in performance and component flow blocking in long-term high-load operating environments, making it difficult to maintain stability and efficiency.
The distributed FPGA partial dynamic reconstruction componentization method is adopted, and the system status monitoring and prediction is carried out in combination with deep learning models. The FPGA functional components are adjusted and optimized in real time through dynamic partial reconfigurable technology.
Accurate prediction and active optimization of possible blocking, resource conflicts or performance degradation in the future are achieved, and the stability and efficiency of the system under high load and long-term operation are improved.
Smart Images

Figure CN120011304A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and proposes a distributed FPGA partial dynamic reconstruction componentization method. The method is based on FPGA distributed components and combines deep learning models to monitor and predict system status, accurately identify potential performance bottlenecks and blocking trends, and adjust and optimize FPGA functional components in real time through dynamic partial reconfiguration technology to ensure the stability and efficiency of the system under high load and long-term operation. Background Art
[0002] With the growing demand for complex computing tasks and large-scale data processing, traditional FPGA architectures have gradually shown their limitations in long-term high-load operating environments. Traditional methods have significant bottlenecks in resource allocation, task scheduling, and system operating efficiency, which often lead to problems such as component flow blocking, uneven resource allocation, and performance degradation, seriously affecting the overall operating efficiency of the system and even causing task interruptions or failures. This static configuration architecture cannot flexibly adapt to dynamically changing computing needs and is difficult to maintain stability and efficiency in high-intensity, long-term operating scenarios. Therefore, how to achieve efficient dynamic management and resource optimization of FPGA architecture has become an important technical problem that needs to be solved urgently.
[0003] In recent years, deep learning technology has made significant breakthroughs, especially in data prediction, pattern recognition, and adaptive optimization. By training a large amount of historical data, deep learning models can accurately predict system status, identify potential problems, and make optimization decisions in advance. Its adaptive learning and iterative optimization capabilities enable the model to continuously improve performance in complex and dynamic environments. These characteristics make deep learning show great potential in system status monitoring, resource allocation optimization, and fault prediction, and provide new ideas and methods for solving the technical bottlenecks faced by traditional FPGA architectures.
[0004] Combining deep learning with FPGA dynamic partial reconfiguration technology has become a key research direction of current academic and industrial circles. Through real-time monitoring and intelligent evaluation of the FPGA system status through deep learning models, it is possible to accurately predict the blocking and resource conflicts that may occur during the operation of the system, thereby realizing dynamic resource allocation and adaptive adjustment of functional modules. Based on this combination, the system can dynamically reconfigure the FPGA without affecting the overall task operation, and flexibly respond to resource bottlenecks and performance issues in high-intensity and long-term operation scenarios. In addition, the introduction of distributed architecture further improves the collaborative efficiency and resource sharing capabilities of multiple FPGA nodes, providing strong support for large-scale data processing and high-performance computing scenarios. This deep integration can not only provide more efficient computing performance, improve real-time performance and flexibility, but also effectively solve the problems of high power consumption, poor adaptability, and hardware resource waste in existing technologies. This combination has important advantages in practical applications, especially in embedded systems and edge computing scenarios that require high performance, low power consumption, and flexible configuration, which can significantly improve the overall performance and scalability of the system. Summary of the invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a distributed FPGA partial dynamic reconstruction componentization method that combines deep learning technology to perform real-time status monitoring and prediction during system operation, and can accurately predict problems such as congestion, resource conflicts or performance degradation that may occur in the future.
[0006] The object of the present invention is achieved through the following technical solution: a distributed FPGA partial dynamic reconstruction componentization method, comprising the following steps:
[0007] Step 1: Monitor and collect data on the operating status of multiple FPGA servers, and collect historical monitoring data of the servers, including three historical load status data: CPU usage, memory usage, and hard disk usage;
[0008] Step 2: Preprocess the historical monitoring data: First, perform zero-mean normalization on the three data types of CPU usage, memory usage, and hard disk usage; then standardize and crop the data into multiple samples of length m; and generate a sample label for each sample: if the data of a sample does not show full load in the next two hours, set the label of the sample to 0, otherwise set it to 1;
[0009] The normalized expression is as follows:
[0010]
[0011] Among them, X represents the original data, X *represents the normalized data, μ represents the mean of the original data, and δ represents the standard deviation of the original data; the specific expressions of μ and δ are as follows:
[0012]
[0013] n is the number of samples, x i is the i-th sample in the original data;
[0014] Step 3: Use the data obtained in step 2 to train the deep learning network model;
[0015] Step 4: Initialize the distributed FPGA service cluster: First, perform self-check and verification on the FPGA hardware resources to ensure the normal operation of the logic unit, storage unit, clock management module and communication interface; then, the system loads the basic configuration file, initializes the dynamic reconstruction controller, presets the partition of the reconstruction area and the reconfigurable logic module, and establishes the dynamic reconstruction management table;
[0016] Step 5: The system starts the monitoring module to complete the real-time monitoring of the cluster system status and collect the CPU usage, memory usage, and hard disk usage of each server node in the cluster in real time;
[0017] Step 6: After normalizing and standardizing the data collected in real time, input it into the deep learning network model trained in step 3 for prediction, obtain the future system status results and perform abnormal status judgment; when there is no abnormality in the system, return to step 5 for real-time monitoring; when there is an abnormality in the system, find the abnormal server node and generate a dynamic reconstruction plan;
[0018] Step 7: Execute the FPGA dynamic reconstruction solution, use the FPGA dynamic partial reconstruction technology, load the new Bitstream to the specified area or suspend the FPGA component service, and perform self-check on the reconstructed FPGA server to ensure the normal operation of each component.
[0019] The beneficial effects of the present invention are:
[0020] (1) The present invention is capable of intelligent prediction and adaptive optimization;
[0021] The present invention combines deep learning technology to perform real-time status monitoring and prediction during system operation, and can accurately predict problems such as blockages, resource conflicts or performance degradation that may occur in the future. This prediction mechanism relies on the construction of a high-precision system status prediction model through training with historical data and input of real-time monitoring data. Unlike the traditional passive fault repair mode, this method can actively make optimization decisions before potential problems appear, reduce the probability of failure, and avoid the impact of performance bottlenecks on the overall efficiency of the system. At the same time, the system also has adaptive optimization capabilities, dynamically adjusting system parameters according to different task scenarios and operating environments to ensure the optimal configuration of resource utilization. This intelligent adaptive mechanism not only reduces the need for manual intervention, but also improves the stability and response speed of the system, and can maintain continuous and efficient performance output in complex and changeable application environments.
[0022] (2) The present invention has efficient dynamic reconstruction capability;
[0023] The present invention realizes dynamic replacement, reconfiguration and resource adjustment of specific functional modules during system operation by introducing FPGA dynamic partial reconfiguration technology. This reconstruction process can be carried out without interrupting the overall operation of the system, ensuring the continuity and stability of the task. Under long-term high-load operation environment, traditional FPGA architecture often leads to performance bottlenecks or even system crashes due to resource solidification and lack of flexibility. When the invention detects potential problems such as system performance degradation, resource conflicts or component flow blockage, it can quickly trigger the dynamic reconstruction process, reallocate computing resources or replace faulty components to restore the normal operation of the system. In addition, this dynamic reconstruction not only improves the flexibility of the system, but also avoids task failures caused by single point failures, and improves the reliability and stability of the system in complex computing task scenarios. Therefore, this method shows significant advantages when facing application requirements such as high real-time performance, strong adaptability and long-term continuous operation.
[0024] (3) The present invention supports distributed architecture collaborative work.
[0025] The present invention adopts a distributed architecture design to achieve efficient collaboration and resource sharing between multiple FPGA nodes. This architecture allows multiple FPGA devices to cooperate with each other in large-scale computing tasks to form an efficient and flexible computing network. Under the distributed architecture, each node can independently complete a specific computing task, and at the same time, communicate and share resources when necessary, avoiding the occurrence of system crashes caused by exhaustion of resources at a single node. In addition, the distributed collaborative work mechanism enables the system to show stronger flexibility and scalability in scenarios such as high-concurrency tasks, large data volume processing, and cross-node resource scheduling. Even in the case of node failure, the system can still ensure the continuity and stability of the task through dynamic resource scheduling and task transfer. The introduction of this distributed architecture provides strong support for complex computing tasks and can significantly improve the overall performance and stability of the system.
[0026] (4) The present invention supports continuous model iteration and optimization.
[0027] The system state prediction model of the present invention has the ability of self-iteration and continuous optimization, and can perform continuous model training and updating according to the system monitoring data collected recently. This mechanism enables the system to always maintain high prediction accuracy and adaptability during long-term operation. When the system triggers specific conditions, the model will automatically enter the iterative update process, and use the latest monitoring data to adjust the model parameters to better match the current system state and task requirements. This continuous optimization process not only improves the adaptability of the system in complex environments, but also avoids the problems of model aging and prediction inaccuracy. In addition, through the automated model training and deployment process, the system reduces manual intervention and operation and maintenance costs, ensuring that efficient prediction and decision-making capabilities can be maintained in different scenarios. This continuous iteration and optimization mechanism provides a strong guarantee for the stable operation of the system in long-term high-load scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic flow chart of the component-based method for dynamic partial reconfiguration of FPGA based on distributed configuration of the present invention;
[0029] Figure 2 This is a schematic diagram of the hardware composition of the distributed FPGA cluster based on the present invention;
[0030] Figure 3 It is a schematic diagram of the system service composition of the distributed FPGA partial dynamic reconstruction componentization method of the present invention;
[0031] Figure 4 A schematic diagram of a system state prediction network designed for the present invention;
[0032] Figure 5This is a network structure diagram for predicting system status during the specific implementation of the present invention. DETAILED DESCRIPTION
[0033] The present invention proposes a distributed FPGA partial dynamic reconstruction componentization method. The method is based on FPGA distributed components and combines deep learning models to monitor and predict system status, accurately identify potential performance bottlenecks and blocking trends, and use dynamic partial reconfiguration technology to adjust and optimize FPGA functional components in real time to ensure the stability and efficiency of the system under high load and long-term operation.
[0034] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0035] like Figure 1 As shown, a distributed FPGA partial dynamic reconfiguration componentization method of the present invention includes the following steps:
[0036] Step 1: Monitor and collect data on the operating status of multiple FPGA servers, and collect historical monitoring data of the servers, including three historical load status data: CPU usage, memory usage, and hard disk usage. The hardware composition diagram of the distributed FPGA cluster is as follows: Figure 2 As shown in the figure, the system consists of a main server and multiple FPGA servers through a core switch. The main server is equipped with a GPU computing graphics card, and the FPGA server is equipped with multiple FPGA computing boards. Data accesses the cluster system through the core switch. The cluster sends data to each FPGA server for calculation and reasoning according to the data business situation, and merges the calculation results for return.
[0037] Step 2: Preprocess the historical monitoring data: First, perform zero-mean normalization on the three data types of CPU usage, memory usage, and hard disk usage; then standardize and crop the data into multiple samples of length m; and generate a sample label for each sample: if the data of a sample does not show full load in the next two hours, it means there is no load abnormality, and the label of the sample is set to 0, otherwise it is set to 1; the constructed data set contains N samples, the dimension of the data set is N*3*m, and the samples are divided into training set and validation set;
[0038] The normalized expression is as follows:
[0039]
[0040] Among them, X represents the original data, X * represents the normalized data, μ represents the mean of the original data, and δ represents the standard deviation of the original data; the specific expressions of μ and δ are as follows:
[0041]
[0042] n is the number of samples, x i is the i-th sample in the original data;
[0043] Step 3: Use the training set and validation set data obtained in step 2 to train the deep learning network model. Send the training set and validation set together to the model training; the training set is used to train the model; the validation set is used to evaluate the training effect of the model during the training process and does not participate in back propagation. Use the model to predict the input data and obtain the system state. Build an error estimate for the predicted data and the real data, and back propagate the model through the optimizer to update the model parameters. After multiple iterations of training, the model area is stable, the error between the predicted data and the real data becomes very small, and the validation set has a high accuracy rate. The model can stop training and obtain a trained deep learning model. The trained model can accurately predict the future operating status of the system.
[0044] Step 4: Initialize the distributed FPGA service cluster: First, perform self-check and verification on the FPGA hardware resources to ensure the normal operation of the logic unit, storage unit, clock management module and communication interface; then, the system loads the basic configuration file, initializes the dynamic reconstruction controller, presets the partition of the reconstruction area and the reconfigurable logic module, and establishes the dynamic reconstruction management table;
[0045] Step 5: The system starts the monitoring module to complete the real-time monitoring of the cluster system status, collect the CPU usage, memory usage, and hard disk usage of each server node in the cluster in real time, and provide data support for subsequent dynamic reconstruction decisions;
[0046] Step 6: After normalization and standardized trimming of the real-time collected data, the data is input into the deep learning network model trained in step 3 for prediction, and the future system status results are obtained and abnormal status is identified; when there is no abnormality in the system, it returns to step 5 for real-time monitoring; when there is an abnormality in the system, the abnormal server node is found and a dynamic reconstruction plan is generated;
[0047] Step 7: Execute the FPGA dynamic reconstruction solution, use the FPGA dynamic partial reconstruction technology, load the new Bitstream to the specified area or suspend the FPGA component service, and perform self-check on the reconstructed FPGA server to ensure the normal operation of each component.
[0048] Step 8: Continue to return to step 5 for real-time monitoring and sort out the abnormal data. The present invention can also set up timed training. When the conditions are met, the system will sort out the historical data of the recent period and go to step 2 to update the historical data, retrain the model, and continuously iterate and optimize the model for incremental learning, thereby ensuring the longer-term stability of the system.
[0049] The present invention is based on the distributed FPGA partial dynamic reconstruction component method system service composition diagram as shown in Figure 3 As shown, the system main server is deployed with training service, prediction service, deployment service, and monitoring service, and the computing service is deployed in the FPGA server; the training service is used to complete the iterative update of the system status prediction model, the prediction service is used to complete the reasoning of the future system status, the monitoring service is used to complete the status monitoring of the system and each sub-computing node, and the deployment service is used to formulate dynamic reconstruction plans and issue instructions.
[0050] The structure of the system state prediction network designed by the present invention is a three-input single-output network structure, such as Figure 4 The principle of the network structure is to extract features from different input data separately through the convolution module, splice the extracted features using the feature splicing module, continue to use the convolution module to extract high-order features from the spliced features, and then use the fully connected module to integrate the features and map them to the single output feature space. The specific network structure of this embodiment is as follows Figure 5 As shown in the figure, the input size of each input layer is 1*8192*1, and a CNN network feature extraction module composed of network convolution, standardization and pooling is used after each input layer; the CNN network feature extraction module first uses the convolution kernel to extract features from the input feature matrix, superimposes the extraction result and the input data into a matrix and uses standardization, and finally performs pooling; 512, 256, and 128 convolution kernel modules are used to process the three inputs in turn, and 1*2 pooling is performed to obtain three feature matrices of size 1*1024*128; the outputs of the three CNN network feature extraction modules are then spliced through the network feature connection module according to the feature matrix splicing method to obtain a feature matrix of 1*1024*384; the CNN network feature extraction module composed of 64, 128, 256, and 512 convolution kernel modules is continued to be input for feature processing to obtain a feature vector of 1*64*512, and after flattening, a 512-dimensional fully connected module is used for feature convergence, and finally output through a 1-dimensional full connection.
[0051] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A distributed FPGA partial dynamic reconfiguration componentization method, characterized in that: The following steps are involved: Step 1: Monitor and collect data on the operating status of multiple FPGA servers, and collect historical monitoring data of the servers, including three historical load status data: CPU usage, memory usage, and hard disk usage; Step 2: Preprocess the historical monitoring data: First, perform zero-mean normalization on the three data of CPU usage, memory usage, and hard disk usage; then standardize and crop the data into multiple samples of length m; And generate a sample label for each sample: if the data of a sample does not show full load in the next two hours, the label of the sample is set to 0, otherwise it is set to 1; The normalized expression is as follows: Among them, X represents the original data, X * represents the normalized data, μ represents the mean of the original data, and δ represents the standard deviation of the original data; the specific expressions of μ and δ are as follows: n is the number of samples, x i is the i-th sample in the original data; Step 3: Use the data obtained in step 2 to train the deep learning network model; Step 4: Initialize the distributed FPGA service cluster: First, perform self-check and verification on the FPGA hardware resources to ensure the normal operation of the logic unit, storage unit, clock management module and communication interface; then, the system loads the basic configuration file, initializes the dynamic reconstruction controller, presets the partition of the reconstruction area and the reconfigurable logic module, and establishes the dynamic reconstruction management table; Step 5: The system starts the monitoring module to complete the real-time monitoring of the cluster system status and collect the CPU usage, memory usage, and hard disk usage of each server node in the cluster in real time; Step 6: After normalizing and standardizing the data collected in real time, input it into the deep learning network model trained in step 3 for prediction, obtain the future system status results and perform abnormal status judgment; when there is no abnormality in the system, return to step 5 for real-time monitoring; when there is an abnormality in the system, find the abnormal server node and generate a dynamic reconstruction plan; Step 7: Execute the FPGA dynamic reconstruction solution, use the FPGA dynamic partial reconstruction technology, load the new Bitstream to the specified area or suspend the FPGA component service, and perform self-check on the reconstructed FPGA server to ensure the normal operation of each component.
Citation Information
Patent Citations
Network-based remote FPGA reconstruction method, FPGA and system
CN118264696A
Distributed FPGA (Field Programmable Gate Array) reconstruction management method and system
CN119166583A
Signal processing detection system based on FPGA architecture
CN119167065A
KR20240077313A