FPGA real-time simulation method and device based on adaptive base value and variable step size, and readable storage medium thereof
Through the synergy of machine learning models and PID controllers, adaptive adjustment of base values and step sizes in FPGA real-time simulation is achieved, solving the problems of low efficiency and insufficient precision in existing technologies, improving simulation efficiency and accuracy, adapting to complex systems, and reducing hardware resource consumption.
Patent Information
- Application Number
- CN202510923143.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-04
AI Technical Summary
In existing FPGA real-time simulation technology, the fixed base value and fixed step size methods are difficult to adapt to dynamic system changes, resulting in low simulation efficiency and insufficient accuracy, especially poor performance in strongly nonlinear and multi-scale systems. In addition, the floating-point operation hardware resource consumption is large, making it difficult to meet the microsecond-level real-time simulation requirements.
The machine learning model is used to train the base value prediction capability, combined with PID real-time error correction and variable step-size-base value collaborative control. Through the method of adaptive base value and variable step size, the base value selection is automated and the step size adjustment is intelligent. Combined with fixed-point operations, efficient simulation is performed on FPGA.
The efficiency of base value adjustment has been improved by more than 90%, the accuracy has been improved by 60%, the hardware resource utilization has been improved by 40%, and the computing efficiency has been increased by 3 times. It can adapt to strong nonlinear and multi-scale systems, break through the limitations of traditional solutions, and meet the needs of microsecond real-time simulation.
Smart Images

Figure CN120430257B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of real-time simulation technology, and in particular to a high-precision real-time simulation method on a field programmable gate array (FPGA), and more particularly to a dynamic system simulation optimization technology based on adaptive control and high-order numerical algorithms. Background Art
[0002] In the field of FPGA real-time simulation, existing technologies (such as CN118260026A) combine per-unit scaling with fixed-point operations. While this can reduce hardware resource consumption, it has two core limitations:
[0003] 1. Insufficient adaptability of fixed-base per-unit scaling: CN118260026A uses fixed-base per-unit scaling to scale differential equations. When system parameters change dynamically (such as sudden changes in motor load or power system failures), the per-unit value may exceed the fixed-point representation range, resulting in overflow or loss of precision. This makes it particularly difficult to adapt to strongly nonlinear or multi-scale systems.
[0004] 2. Efficiency drawbacks of fixed-step discretization: CN118260026A uses low-order fixed-step discretization methods such as the Euler method. Extremely small step sizes are required to maintain stability in strong transient links, resulting in a surge in computational complexity. Small step sizes are still used for smooth links, wasting hardware resources and failing to achieve a dynamic balance between accuracy and efficiency.
[0005] In addition, although traditional floating-point arithmetic solutions can handle a wide range of values, they consume a lot of hardware resources on FPGAs (for example, CLBLUT consumes more than four times the fixed-point solution) and have high calculation delays (500ns), making it difficult to meet microsecond-level real-time simulation requirements.
[0006] Therefore, there is an urgent need for an FPGA real-time simulation method, device and readable storage medium based on adaptive base value and variable step size to solve the problems existing in the prior art. Summary of the Invention
[0007] The embodiments of the present invention provide an FPGA real-time simulation method, device and readable storage medium based on adaptive base value and variable step size. These methods address the problems of current technologies that use fixed base values or simple feedback to adjust base values, rely on manual parameter adjustment for variable step size strategies, lack the ability to intelligently predict the dynamic characteristics of complex systems, and result in low simulation efficiency and insufficient accuracy, making it particularly difficult to adapt to highly nonlinear and multi-scale industrial scenarios.
[0008] The core technology of this invention is to build a triple innovative architecture of "intelligent prediction + dynamic optimization + hardware acceleration" on FPGA through offline training of base value prediction capability of machine learning model, combined with PID real-time error correction and variable step-size-base value collaborative control, to achieve automation of base value selection, intelligent step-size adjustment and efficient reuse of computing resources.
[0009] In a first aspect, the present invention provides an FPGA real-time simulation method based on adaptive base value and variable step size, the method comprising the following steps:
[0010] By training historical data through machine learning models, a mapping relationship between system states and per-unit base values is established to generate a base value prediction model. The per-unit base value is the benchmark value for converting physical quantities in differential equations into dimensionless per-unit values.
[0011] The base value prediction model is used to output the initial base value, and the initial base value is dynamically fine-tuned through a controller including proportional, integral, and differential feedback links to obtain the final base value;
[0012] A high-order numerical method with error estimation is used to discretize the per-unit differential equation, and the simulation step size is automatically adjusted according to the dynamic range of the base value.
[0013] The discretized differential equations are converted into hardware description language based on fixed-point operations and deployed on FPGA for execution.
[0014] Furthermore, the initial value of the base value is dynamically fine-tuned through a controller including proportional, integral, and differential feedback links, including:
[0015] Based on the absolute error, cumulative error and error change rate of the current physical quantity, the initial value of the base value is adjusted through a linear combination of proportional term, integral term and differential term.
[0016] Furthermore, the simulation step size is automatically adjusted according to the base value dynamic range and error tolerance, including:
[0017] The local truncation error is estimated using the difference between the calculation results of high-order numerical methods and low-order numerical methods;
[0018] The simulation step size is dynamically expanded or reduced based on the local truncation error, the preset error tolerance, and the base value adjustment range, where the step size adjustment range is limited to 0.1 to 2 times the current step size.
[0019] Furthermore, the machine learning model adopts a lightweight neural network architecture, quantizes parameters into fixed-point numbers through model compression technology, and the inference delay deployed to FPGA is ≤50ns.
[0020] Furthermore, high-order numerical methods include the Runge-Kutta method and the implicit trapezoidal method, which are automatically switched according to the dynamic characteristics of the system:
[0021] When the real part difference of the system eigenvalues satisfies the rigidity condition, the implicit trapezoidal method is used;
[0022] When the system is non-stiff, explicit high-order methods are used.
[0023] In a second aspect, the present invention provides a device equipped with the above-mentioned FPGA real-time simulation method based on adaptive base value and variable step size, comprising:
[0024] Machine learning inference unit: used to run the lightweight base value prediction model and output the initial value of the base value;
[0025] Dual-loop control unit: includes an inner-loop PID controller and an outer-loop variable step-size decision maker. The PID controller is used to fine-tune the base value based on error feedback, and the variable step-size decision maker is used to adjust the step size according to the error estimate.
[0026] Dynamic discretization unit: supports dynamic switching of high-order numerical methods, integrating the Runge-Kutta method pipeline and implicit trapezoidal method iterative solver;
[0027] Fixed-point arithmetic unit: used to convert differential equations into hardware description language and perform parallel calculations.
[0028] Furthermore, the machine learning inference unit uses distributed arithmetic to optimize fixed-point multiplication, including an input buffer, a pipelined multiplication-accumulation unit, and an activation function lookup table, supporting 16-bit fixed-point parallel inference.
[0029] Furthermore, the dual-loop control units share an error calculation module, the PID controller periodically updates the base value, and the variable step size decision maker periodically adjusts the step size.
[0030] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the above-mentioned FPGA real-time simulation method based on adaptive base value and variable step size.
[0031] In a fourth aspect, the present invention provides a readable storage medium, in which a computer program is stored. The computer program includes a program code for controlling a process to execute a process, and the process includes the above-mentioned FPGA real-time simulation method based on adaptive base value and variable step size.
[0032] The main contributions and innovations of the present invention are as follows:
[0033] 1. Intelligent base value selection: Through offline training of machine learning models, a mapping relationship between system status and base values is established, enabling automated prediction of the initial base value. This replaces the traditional trial-and-error method, improving base value adjustment efficiency by over 90% and avoiding errors caused by human experience (e.g., the error of traditional solutions can reach 2.1%, while this method reduces it to 0.18%).
[0034] 2. Dual-loop collaborative optimization accuracy and efficiency:
[0035] Inner loop: The PID controller dynamically fine-tunes the base value based on real-time error to resolve overflow or precision loss caused by a fixed base value and adapt to parameter mutation scenarios (for example, reducing the response time to a sudden motor load change from 200μs to 30μs).
[0036] Outer loop: The variable step size algorithm automatically adjusts the step size in combination with the base value dynamic range. A high-order method (such as RK4) is used for strong transient links, reducing the amount of calculation by more than 50%. At the same time, the error tolerance adaptive mechanism improves the accuracy by 60%.
[0037] 3. Efficient reuse of FPGA hardware resources:
[0038] The lightweight machine learning model achieves inference latency of ≤50ns through fixed-point quantization and pipeline design. Although resource consumption increases from 14,458 to 28,300 (still far less than the 67,586 required for traditional floating-point operations), computing efficiency is more than tripled (14,500 of these are used for the distributed arithmetic lookup table of the machine learning inference core, 8,800 for the PID control state machine, and 5,000 for high-order discretization logic). A three-core pipeline (200MHz clock) achieves 2 million simulation steps per second, a three-fold improvement in efficiency compared to the traditional solution (667,000 steps per second). Using traditional floating-point operations to achieve the same accuracy would require 67,586 CLBLUTs, increasing hardware costs by 138.8% ((67,586 - 28,300) / 28,300 ≈ 138.8%)).
[0039] Dynamically reconstructed discretization units support method switching, increasing hardware utilization by 40%, and are particularly suitable for multi-physics coupling systems (such as variable operating condition simulation of lithium battery packs).
[0040] 4. Significantly improved generalization capabilities: Breaking through the limitations of traditional solutions for linear systems, it can handle highly nonlinear and multi-scale scenarios (such as chaotic oscillation circuits and multi-modal simulation of aircraft engines), broadening the industrial application boundaries of real-time simulation.
[0041] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below so that other features, objects, and advantages of the invention are more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0043] Figure 1 Flowchart of an FPGA real-time simulation method based on adaptive base value and variable step size according to an embodiment of the present invention;
[0044] Figure 2FIG. 4 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0045] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.
[0046] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0047] Our existing technology, CN118260026A, can convert simulated systems of varying numerical scales to a unified numerical scale through per-unit normalization, enabling simulation calculations on an FPGA based on fixed-point arithmetic. This can improve simulation scale, frequency, and accuracy at a limited cost. However, it utilizes a variable step-size strategy that uses a fixed base value or simple feedback to adjust the base value, relying on manual parameter adjustment. This strategy lacks the ability to intelligently predict the dynamic characteristics of complex systems, resulting in low simulation efficiency and insufficient accuracy. It is particularly difficult to adapt to highly nonlinear, multi-scale industrial scenarios.
[0048] Based on this, the present invention solves the problems existing in the prior art by combining the offline training base value prediction capability of the machine learning model with PID real-time error correction and variable step-size-base value collaborative control.
[0049] Example 1
[0050] The present invention aims to propose an FPGA real-time simulation method based on adaptive base value and variable step size, specifically, referring to Figure 1 , the method comprises the following steps:
[0051] Step 1: Use a machine learning model to train historical data, establish a mapping relationship between system status and normalized base values, and generate a base value prediction model;
[0052] In this embodiment, in the offline stage, multi-operating condition data of the system (such as the voltage / current waveform of the motor at different speeds and temperatures) is collected, and a base value prediction model is trained through a fully connected neural network (FCN) or a Transformer model (the model is actually selected according to the specific application object, for example, FCN is used for static features (such as nonlinear mapping of SOC and voltage base value), and LSTM is used for time series features (such as time series scenarios with temperature change rate > 1°C / s). It is not limited to these two models. Different objects can be applied to different models, which will not be repeated here). The output is the optimal base value K of each physical quantity. opt .
[0053] In the online stage, the machine learning model is used to output K init = K opt (x input ), as the initial value of the PID controller.
[0054] Among them, the model structure of the machine learning model includes:
[0055] Input layer: system state variables (such as voltage u, current i, temperature T) and operating parameters (such as load L, speed ω);
[0056] Hidden layer: Uses the attention mechanism to capture the coupling relationship between variables, such as the joint impact of voltage and temperature on the motor's internal resistance;
[0057] Output layer: the base value of each physical quantity (K u ,K i ,K T ,...), which is obtained by mapping through linear layers.
[0058] Preferably, in order to achieve lightweight deployment of the model, knowledge distillation is used to compress the complex model into a lightweight network represented by 16-bit fixed-point numbers (for example, the parameter size is reduced by 80%), and at the same time, the floating-point weights are converted into fixed-point numbers (such as Q15 format), and the FPGA's fixed-point multiplier (DSP48E) is used to accelerate inference.
[0059] In this embodiment, the data preprocessing process is as follows:
[0060] Perform sliding window feature extraction on the collected historical data (such as lithium battery voltage / current waveforms). The window size is set to 100 sampling points (corresponding to 10ms duration). Statistical features such as mean, variance, and kurtosis are extracted as supplementary dimensions for model input.
[0061] Use the normalization method: convert the physical quantity into the [-1,1] interval.
[0062] In this embodiment, the training parameter configuration may be:
[0063] Optimizer: Use Adam optimizer, initial learning rate , decays by 50% every 50 epochs;
[0064] Loss function: for example, combined loss ,in:
[0065] ,
[0066] ;
[0067] in, is the predicted base value of the i-th sample; is the true base value, and N is the total number of samples. RMSE measures the absolute deviation between the predicted value and the true value and is more sensitive to larger errors (because the squared error amplifies the difference). MAPE measures the relative deviation between the predicted value and the true value, expressed as a percentage, and is suitable for unified evaluation of base values of different scales (for example, the error between a voltage base value of 3.7V and a current base value of 10A can be uniformly compared using percentages). The traditional solution (CN118260026A) uses only a single loss function (such as MSE) and cannot simultaneously address both the "absolute error of large-scale base values" and the "relative error of small-scale base values." By combining the losses, the present invention reduces the base value prediction error from 2.1% of the traditional solution to 0.18%.
[0068] Among them, the weight ratio is determined by cross-validation:
[0069] 0.5RMSE+0.5MAPE: average error 0.35%, overflow rate 5% for large base value scenarios;
[0070] 0.7RMSE+0.3MAPE: average error 0.18%, overflow rate 0.3%;
[0071] 0.8RMSE+0.2MAPE: The average error is 0.21%. The relative error increases in small base value scenarios.
[0072] Therefore, 0.7 / 0.3 is selected to balance absolute and relative accuracy.
[0073] Data augmentation: Inject ±1%FS Gaussian noise into the temperature parameter (FS is the full scale range of the temperature sensor, 100°C, i.e. ±1°C) and add ±5% random offset to the voltage / current waveform to simulate industrial-grade sensor errors (accuracy level ±0.5%FS). (This is just an example; the actual adjustment range is set according to requirements.)
[0074] The base value for per-unit system (BV) is the reference value used when converting physical quantities (such as voltage, current, and resistance) to dimensionless per-unit values. Its core function is to normalize variables of different numerical scales to the range representable by fixed-point numbers (such as the [-1, 1] interval for 16-bit fixed-point numbers) through the "physical quantity / base value" normalization operation. This adapts to the hardware computing characteristics of FPGAs and avoids the high resource consumption of floating-point operations. The calculation method for the base value is existing technology, such as CN118260026A, and is therefore not further described.
[0075] In summary, the purpose of step one is to enable the computer to discover the hidden patterns between system states (such as voltage, current, and temperature) and optimal baseline values from historical data, thereby enabling the computer to automatically recommend baseline values based on the "observe state" principle. This trained model can then be directly fed with existing data to generate optimal baseline values.
[0076] Step 2: Use the base value prediction model to output the initial base value, and dynamically fine-tune the initial base value through a controller including proportional, integral, and differential feedback links to obtain the final base value;
[0077] In this embodiment, the base value prediction model outputs the base value initial value, which means outputting independent base values for different physical quantities, such as the voltage base value K u , Current base value K i , internal resistance base value K R Etc., forming a basis value vector K init = [K u , K i , K R , ...].
[0078] In this embodiment, the PID controller is based on the real-time error Err k Fine-tuning the base value yields:
[0079]
[0080] Among them, K final is the final base value after dynamic adjustment; K init is the initial value of the base value; K p , K i , K d are the proportional, integral, and differential coefficients; is the cumulative sum of errors from the initial moment to the current moment n; is the absolute error of the physical quantity at the previous moment. The absolute error is the absolute value of the difference between the measured value and the initial value. For example, if the measured voltage is 3.75V and the initial value is 3.8V, the per-unit value is 3.75 / 3.8 ≈ 0.987. The ideal range is [-1, 1], so the error is 0.013 (very small, but requires fine-tuning).
[0081] In the proportional link, if the current error If the voltage per unit value exceeds the upper limit, increase K u to expand the representation range); in the integral phase, historical errors are accumulated to ensure that the error approaches zero in the steady state (for example, when the error is continuously small, the base value is gradually adjusted through the integral term); in the differential phase, the trend is predicted based on the error change rate, and the base value is adjusted in advance (for example, when a rapid increase in error is detected, the differential effect is increased to suppress overshoot).
[0082] For example, the machine learning model outputs the initial value of the voltage base value K based on SOC=20%, temperature=0℃, and charging current 2A. init =3.8V, the normalized measured voltage is 3.6V, which is 3.6 / 3.8≈0.947 (safety range).
[0083] Error feedback during charging: After 10ms, the voltage rises to 3.75V, the per-unit value is 3.75 / 3.8≈0.987, the error =0.013; proportional term: K p =0.5, the contribution is 0.5×0.013=0.0065; Integral term: the sum of the first five errors is 0.05, K i =0.1, the contribution is 0.1×0.05=0.005; differential term: the last error =0.01, rate of change 0.013-0.01=0.003, K d = 0.2, the contribution is 0.2 × 0.003 = 0.0006; the final base value is: 3.8 + 0.0065 + 0.005 + 0.0006 = 3.8121V, and the per-unit value is 3.75 / 3.8121 ≈ 0.984, which is closer to the center of the ideal range. If the charging current suddenly increases at this time (similar to driving on a steep slope), the differential link will detect the rapid increase in error and increase the base value in advance to prevent the per-unit value from overflowing, just like downshifting and accelerating in advance to prevent difficulty climbing a hill.
[0084] Preferably, the Ziegler-Nichols method can also be used to adjust the parameters: first set K i =K d =0, gradually increase K p Until the system critical oscillation, record the oscillation period T u With critical gain K u , then K p=0.6Ku, K i =1.2K u / T u , K d =0.075K u T u For example, in the case of lithium batteries, K u =2.3, T u =15ms, calculated K p =1.38, K i =0.184, K d =0.026, and the actual empirical value is 0.5 / 0.01 / 0.005 to balance stability.
[0085] In summary, the core of step two is to "add an intelligent fine-tuner to the base value," much like the automatic steering wheel correction system when driving, allowing the base value to be dynamically optimized based on real-time conditions. This dual mechanism of "prediction + feedback" allows the base value to flexibly respond to various complex working conditions like a seasoned driver, reducing the simulation error from 2.1% with traditional solutions to 0.18%, just as autonomous driving is smoother and more accurate than manual driving.
[0086] Step 3: Discretize the per-unit differential equation using a high-order numerical method with error estimation, and automatically adjust the simulation step size according to the dynamic range of the base value;
[0087] The traditional Euler method (low-order) is like driving in first gear, maintaining the same speed regardless of whether going up or downhill. This method can cause the engine to "stall" (simulation divergence) on steep slopes (strong transients), or it can be forced to use extremely small step sizes (slowing to a crawl) to avoid stalling, wasting time. To address this issue, a high-order numerical method (RK4 / trapezoidal method) is employed in this embodiment. This is similar to an automatic transmission car, where the system shifts to a lower gear on a steep slope (the rigid subsystem switches to the implicit trapezoidal method) and shifts to a higher gear on a flat road (the non-rigid subsystem switches to the explicit RK4 method). For example, the "slope" (a stiffness indicator, such as "rigidity is determined when the ratio of the real parts of the eigenvalues (which can be determined by the Jacobian matrix) is greater than 100") of the FPGA real-time calculation system. For example, when a lithium battery is charged with a high current, the internal resistance changes significantly (steep slopes), automatically switching to the implicit trapezoidal method; during the constant voltage phase, where the change is gentle (flat roads), the system switches back to the RK4 method.
[0088] Among them, RK4 (4th-order Runge-Kutta method): used for flat roads / gentle slopes, calculates four "trial speeds" (slopes) at a time, accurately predicts the next position, similar to looking at four points ahead to predict the route when driving; Implicit Trapezoidal Method: used for steep slopes / sharp turns, considers the next position and reversely infers the current operation, similar to driving downhill and then stepping on the brakes to ensure no slipping (simulation stability).
[0089] In this embodiment, according to the base value dynamic range (such as K finalMultiple relationship) automatically adjust the upper limit of step size h max , if the base value is magnified 10 times, the maximum step size is allowed to increase accordingly to improve efficiency. For example, the simulation step size is automatically adjusted to:
[0090] Computing local truncation error using embedded error estimation ,in The state value of the system at the next moment calculated by high-order numerical methods; The system state value at the next moment is calculated by the low-order numerical method; it is like "the GPS shows the middle of the road, and the road signs show the side of the road". The bigger the difference, the more direction needs to be adjusted.
[0091] Adjust the step size according to the error tolerance TOL (the larger the deviation, the smaller the step size needs to be, and the smaller the deviation, the larger the step size can be):
[0092] , where h new is the adjusted simulation step size; h old is the simulation step before adjustment; TOL is the preset error tolerance, similar to the error range; f(K) is the base value adjustment factor, f(K)=K final / K init , reflecting the prediction confidence of the machine learning model or the dynamic adjustment range of the base value, which is used to balance accuracy and efficiency.
[0093] Among them, if the base value K increases to α times the initial value (such as α=5), it means that the system enters the large-scale working condition, and the upper limit of the allowed step size (h max ) from h max Expand to , in order to reduce the amount of calculation. The error tolerance TOL of the variable step size algorithm is dynamically adjusted according to the base value prediction accuracy. For example, when the model prediction error is <0.5%, TOL can be relaxed to 10 -2 , to improve calculation efficiency. It is worth noting that the coefficients 0.1 and 0.2 in the above formula can also be 0.05 and 1.5, which can be set according to actual needs. For example, the coefficients are preferably 0.05 and 1.5. Verification through motor starting conditions: the minimum step size can be reduced to 0.05h old (1μs→50ns), maximum amplification to 1.5h old (10μs→15μs), avoiding missed detection and computational redundancy.
[0094] In summary, the purpose of step three is to use this intelligent collaboration similar to "gear position + vehicle speed" to enable the simulation to accurately capture sudden changes (such as the impact of motor startup) and efficiently calculate during stable phases. This is more like driving by an experienced driver than the traditional fixed-step solution, and is both safer and more fuel-efficient (computing power).
[0095] The "per-unit differential equations" in step three are essentially derived from the physical modeling of real systems and are mathematical abstractions of real systems (such as motors, lithium batteries, and power grids). These equations are first established based on physical laws (such as Kirchhoff's laws and the laws of thermodynamics). Per-unit processing is then used to convert physical quantities into dimensionless values to adapt to the fixed-point operations of FPGAs.
[0096] For ease of understanding, a popular analogy is from a "Chinese recipe" to an "English operating manual":
[0097] Primitive differential equations: These are like recipes written in Chinese (describing the physical laws of the system), but the units (V, A, Ω) of different ingredients (physical quantities) are mixed up, and a foreign chef (FPGA) might not be able to understand them directly.
[0098] Per-unit processing: Translate the recipe into English and convert the ingredient amounts into "cups, spoons" (per-unit values), for example "100 grams of sugar" → "0.5 cups of sugar" (the base value is 200 grams);
[0099] The base value in step 1 / 2 is equivalent to determining the conversion ratio of "1 cup = 200 grams". This ratio needs to be dynamically adjusted based on the freshness of the ingredients (system status) (for example, if the sugar becomes moist, 1 cup may = 220 grams);
[0100] Discretization of step three: The foreign chef follows the English recipe (normalized equation) and the dynamically adjusted conversion ratio (base value) step by step (iterative calculation) to make a dish that suits the taste (high-precision simulation result).
[0101] Therefore, the differential equation in step three is the normalized differential equation. The specific operation is the existing technology, and can be found in CN118260026A.
[0102] Step 4: Convert the discretized differential equation into hardware description language based on fixed-point operations and deploy it on the FPGA.
[0103] Because traditional floating-point operations are like hand-made custom parts, the decimal point position of each number is not fixed. FPGA processing requires complex circuits (such as floating-point adders), which is time-consuming and space-consuming (resource consumption is more than four times that of fixed-point). Therefore, using fixed-point operations, decimals can be converted into integers for calculation. For example, a voltage of 3.8V is recorded as 38 (one decimal place is the default), and it is unified into a "16-bit integer" standard part (such as Q15 format: 1 sign bit + 15 decimal places). The FPGA multiplier (DSP48E) can directly and quickly process it. For example:
[0104] / / Define fixed-point register (16 bits, Q15 format)
[0105] reg [15:0] yfix, Rfix, ufix, kfix;
[0106] / / Pipeline first stage: multiplication operation
[0107] wire [31:0] mulresult = $signed(Rfix) * $signed(yfix);
[0108] wire [15:0] mulscaled = mulresult[31:16]; / / right shift 15 bits to remove the decimal point
[0109] / / Pipeline second stage: negation + addition
[0110] always @(posedge clk) begin
[0111] kfix<= $signed(-mulscaled) + $signed(ufix);
[0112] End
[0113] This is equivalent to telling the FPGA: "Do the multiplication first, shift the result right 15 bits, then take the inverse and add u, and process a batch of data in each clock cycle."
[0114] Preferably, FPGA uses a parallel architecture to optimize:
[0115] 1. Tri-core pipeline design:
[0116] Machine Learning Inference Core: handles base value predictions and uses distributed arithmetic (DA) to optimize fixed-point multiplication;
[0117] PID control core: real-time error feedback adjustment, running in parallel with the inference core;
[0118] High-order discretization core: Supports dynamic switching between RK4 (4th-order Runge-Kutta method) and trapezoidal method, with a pipeline depth of 6 stages (coefficient calculation → 4th-order slope → accumulation → error estimation → step size decision). The 6-stage pipeline latency = 6 × 5ns (clock cycle) = 30ns, BRAM weight read latency is 15ns, AXI bus transmission latency is 5ns, and the total inference latency is ≤ 50ns. The relationship between pipeline stage number and latency is as follows:
[0119] Input buffer: 1 cycle (5ns);
[0120] Fully connected computation: 3 cycles (15ns);
[0121] Activation function: 1 cycle (5ns);
[0122] Base value output: 1 cycle (5ns).
[0123] 2. Resource reuse technology:
[0124] Shared Multiply Accumulator (MAC): The inference core and PID core share the FPGA's DSP48E unit, improving resource utilization by 40%;
[0125] Dynamic reconstruction logic: Switch the hardware configuration of the discretization kernel in real time according to the current working condition (rigid / non-rigid), such as switching the 4th-order slope calculation unit of RK4 to the iterative solver of the trapezoidal method.
[0126] In summary, the core of Step 4 is to "translate mathematical calculations into hardware circuits, enabling FPGAs to operate at high speeds like an assembly line," much like converting a recipe into a control program for automated kitchen equipment. This "mathematical formula → hardware circuit" transformation enables FPGAs to achieve microsecond-level real-time simulation through parallel computing, over three times faster than traditional CPU solutions, much like how automated factories are far more efficient than manual production.
[0127] Example 2
[0128] Based on the same concept, the present invention also proposes an FPGA real-time simulation device based on adaptive base value and variable step size, comprising:
[0129] Machine learning inference unit: used to run the lightweight base value prediction model and output the initial value of the base value;
[0130] Dual-loop control unit: the inner loop is a PID controller, and the outer loop is a variable step size decision maker;
[0131] Dynamic discretization unit: supports dynamic switching of high-order numerical methods, integrating the 4th-order Runge-Kutta method pipeline and the trapezoidal method iterative solver;
[0132] Fixed-point arithmetic unit: implements parallel computation of differential equations based on hardware description language.
[0133] In this embodiment, the matrix multiplication of the neural network of the machine learning inference unit is decomposed into a multi-stage pipeline (input buffer → multiply-accumulate → activation function → output), and each clock cycle processes one layer of neurons; for example:
[0134] Input cache level: reads 8-channel status data (such as voltage, current, SOC) in parallel and stores it in BRAM cache;
[0135] Multiply-accumulate stage: uses 8 DSP48E units to calculate the dot product of input and weight in parallel, and stores the result in register;
[0136] Activation level: ReLU activation is achieved through LUT, and the base value initial value K is output init, single-cycle processing delay is 20ns.
[0137] The base value table and neural network weights share BRAM memory, reducing resource usage. The PID controller updates the base value every 10 simulation steps, and the variable step size decision maker adjusts the step size every 2 simulation steps. The machine learning inference unit and the dual-loop control unit transmit the base value initialization and error data via the AXI bus.
[0138] Preferably, the AXI bus configuration timing is:
[0139] t0: Send method switching instruction (1 cycle);
[0140] t1-t5: BRAM reloads the trapezoidal coefficients (5 cycles);
[0141] t6: Pipeline restart, total delay = 6 × 5ns = 30ns ≤ 200ns.
[0142] State continuity during switching: Transition is performed through the forward Euler method to avoid jumps.
[0143] In order to verify the technical effect of the present invention, the present invention is compared with the existing technology and the solution of the present invention after removing machine learning. The results are shown in the following table:
[0144]
[0145] The charge and discharge cut-off voltage is set to 3.0V~4.2V, and the temperature control accuracy is ±0.5℃. Under -20℃ and 3C charging conditions, the base value adjustment trajectory is as follows:
[0146] When the traditional solution has a fixed base value of 3.7V, the per-unit value overflows at 1.05, resulting in an error of 2.1%;
[0147] The base value of the present invention is dynamically adjusted to 3.9V, the per-unit value is 0.97, and the error is 0.18%.
[0148] The reason why the hardware resource consumption of the traditional solution is 18200 CLB LUTs instead of 14458 CLB LUTs is that after actual simulation, the per-unit / de-per-unit module: approximately 4000 CLB LUTs (including base value storage and fixed-point implementation of division operations); Euler method calculation unit: approximately 10000 CLB LUTs (including state storage and addition iteration logic); input / output interface: approximately 4200 CLB LUTs (including data cache and format conversion), totaling approximately 18200 CLB LUTs (the remaining data has also been simulated and is different from the situation of CN118260026A).
[0149] As can be seen, although the present invention adds a machine learning inference unit (approximately 5000 CLB LUTs, including fixed-point multiplication and activation functions for lightweight neural networks); a dual-loop control module (approximately 3000 CLB LUTs, including a PID controller and step-size decision state machine); and high-order discretization switching logic (approximately 2100 CLB LUTs, including dynamic switching between RK4 and trapezoidal methods), the total resource consumption reaches 28,300 (mainstream industrial-grade FPGAs (such as the Xilinx UltraScale+ ZU19EG) contain over 200,000 CLB LUTs, and 28,300 CLB LUTs only account for 14.1%, far below the resource limit). However, the present invention significantly reduces benchmark adjustment time, calculation latency, and operating mode switching response time, achieving a significant performance boost with limited resources. Furthermore, the present invention does not consume all FPGA DSP48E units (only approximately 300 of the total resources exceed 2000), leaving room for expansion.
[0150] Example 3
[0151] Based on the same concept, this example, based on Example 1 or Example 2, uses a new energy vehicle lithium battery pack as the simulation object, demonstrating how to achieve high-precision real-time simulation through machine learning-assisted base value optimization and variable-step-size high-order discretization. Lithium batteries exhibit strong nonlinear characteristics under different states of charge (SOC), temperatures (T), and charge and discharge currents (I). Traditional fixed base value solutions have an error of over 2%, while this invention reduces the error to within 0.18% through three innovations. The details are as follows:
[0152] 1. Application Scenarios
[0153] The simulation object is a battery pack consisting of four lithium-ion batteries connected in series, including the following differential equations:
[0154] 1. Terminal voltage equation: ;
[0155] (E is electromotive force, R is internal resistance, and C is polarization capacitance);
[0156] 2.SOC dynamic equation: ;
[0157] (Q is the battery capacity).
[0158] Key variables are normalized:
[0159] Voltage base value K U = 3.7V (nominal voltage);
[0160] Current base value K I = 10A (rated current);
[0161] Internal resistance base value K R = 5mΩ (initial internal resistance).
[0162] 2. Machine Learning Base Value Prediction Model Training
[0163] 1. Data collection:
[0164] 100,000 sets of data were collected at a temperature range of -20°C to 60°C and a charge and discharge rate of 0.1C to 3C, including:
[0165] Input: [SOC, T, I, AmbientPressure]; AmbientPressure is the ambient pressure.
[0166] Output: Optimal basis value [K U ,K R ] (determined by high-precision floating-point simulation) Model construction:
[0167] An LSTM neural network (3 hidden layers, 128 neurons per layer) is used, and the loss function is the root mean square error (RMSE), which is:
[0168]
[0169] 2. Model lightweight:
[0170] Knowledge distillation: Compressing LSTM into a 2-layer fully connected network (FCN), reducing parameters by 80%;
[0171] Fixed-point quantization: The weights are converted to Q15 format (16-bit fixed-point numbers) and deployed to the FPGA's BRAM memory.
[0172] 3. Dynamic Adjustment of Base Value
[0173] 1. Online reasoning:
[0174] When the battery is at a low temperature (0°C) and a 2C charging condition, the machine learning inference unit outputs the initial base value: , (Prediction error <0.5%).
[0175] 2.PID real-time fine-tuning:
[0176] Real-time error calculation: (reference values are derived from measured data);
[0177] 3.PID formula:
[0178] ;
[0179] 4. Adjustment result: The voltage per unit value at the initial charging stage is close to the upper limit. Adjust from 3.6V to 3.8V to avoid overflow.
[0180] 4. Variable Step Size High-Order Discretization
[0181] 1. Selection of discretization method:
[0182] Initial charging (strong transient, dI / dt>100A / s): Automatically switches to the implicit trapezoidal method, and the differential equation is:
[0183]
[0184] Solve I through FPGA parallel iteration n+1 , step size h=0.5us.
[0185] Constant voltage stage (non-rigid, dI / dt<10A / s): Use the RK4 method, the formula is:
[0186]
[0187] The step size is automatically expanded to h=2us to improve calculation efficiency.
[0188] 2. Step size adjustment logic:
[0189] Adjustment factor based on base value , the step length formula is:
[0190] ;
[0191] When the local error Err h =0.0015, h new =1.5us, dynamically balancing accuracy and efficiency.
[0192] 5. FPGA Hardware Deployment
[0193] Hardware architecture:
[0194] 1. Machine Learning Inference Unit:
[0195] It adopts a 6-stage pipeline (input cache → 3-layer FCN calculation → activation function → base value output), uses 8 DSP48E units to parallelly calculate multiplication and accumulation, processes 16-bit fixed-point numbers in a single cycle, and has an inference delay of 20ns.
[0196] 2. Dual-loop control unit:
[0197] The PID core and the variable step size core run in parallel and share the error calculation module (consuming 222 DSP units);
[0198] The state machine (FSM) triggers a base value update every 10 steps and adjusts the step size every 2 steps.
[0199] 3. Dynamic discretization unit:
[0200] The hardware configuration is dynamically reconfigured through the AXI bus, and the switching between the RK4 pipeline and the ladder method iterator is completed within 200ns.
[0201] For ease of understanding, the following is a systematic explanation of the key professional terms of the present invention:
[0202] 1. FPGA (Field Programmable Gate Array)
[0203] Definition: An integrated circuit chip that consists of modules such as configurable logic blocks (CLBs), memory (BRAM), and digital signal processing units (DSPs). It can dynamically reconfigure circuit functions through hardware description languages (such as Verilog / VHDL) to achieve parallel computing and is suitable for scenarios with high real-time requirements.
[0204] Application of this invention: Utilize the parallel architecture of FPGA to accelerate machine learning reasoning, PID control and high-order numerical calculations, achieve microsecond-level real-time simulation, and reduce computing latency by more than 50% compared to CPU / GPU.
[0205] 2. StiffSubsystem
[0206] Definition: In numerical analysis, a system of differential equations containing both rapidly varying and slowly varying components, where the real parts of their eigenvalues differ significantly (usually by a ratio > 100). Explicit numerical methods would require extremely small step sizes to stabilize, leading to a surge in computational complexity.
[0207] Application of the present invention: Scenarios such as high-current charging of lithium batteries and motor starting belong to rigid systems. The present invention avoids the step size limitation of the explicit method by automatically switching to the implicit trapezoidal method, thereby improving the calculation efficiency by more than 3 times.
[0208] 3.CLBLUT (Configurable Logic Block Look-Up Table)
[0209] Definition: The core component of the configurable logic block (CLB) in an FPGA. Essentially, it is an SRAM-based lookup table (LUT). It implements combinational logic operations by storing a truth table of logic functions. Each LUT typically supports 4-6 input variables.
[0210] Significance of the invention: CLBLUT consumption is a key indicator of FPGA resource usage. The CLBLUT consumption of traditional floating-point solutions is more than four times that of fixed-point solutions. This invention increases the CLBLUT from 14458 to 28300 through fixed-point operations and lightweight models, while improving computing efficiency by three times, demonstrating optimized resource utilization.
[0211] 4.DSP48E (Digital Signal Processing Block)
[0212] Definition: A dedicated digital signal processing unit built into the FPGA that efficiently performs operations such as multiplication, accumulation, and multiply-accumulate (MAC) and supports 18×18-bit fixed-point multiplication. It is the core hardware resource for high-performance digital signal processing and numerical computation.
[0213] Applications of this invention: It is used to accelerate fixed-point multiplication of machine learning models, error calculation of PID controllers, and slope iteration of high-order numerical methods, increasing the computing speed by 5 times compared to general logic resources.
[0214] 5. Knowledge Distillation
[0215] Definition: A model compression technique in machine learning that reduces parameter size while maintaining accuracy by allowing a lightweight "student model" to learn the output distribution (such as soft labels) of a complex "teacher model". It is often used for embedded device deployment.
[0216] The invention's innovation is to distill complex models such as LSTM into a two-layer fully connected network, reducing parameters by 80%. Combined with fixed-point quantization, the FPGA inference delay is reduced to ≤50ns, solving the high latency problem of traditional floating-point models on FPGAs.
[0217] 6. Fixed-Point Quantization
[0218] Definition: Converts floating-point values into integer representations with a fixed decimal point position (e.g., Q15 format: 1 sign bit + 15 decimal bits in 16 bits), replacing floating-point operations with shift operations, significantly reducing hardware resource consumption and computational latency.
[0219] The value of this invention: Quantizing the neural network weights and activation values into 16-bit fixed-point numbers reduces FPGA multiplier resource consumption by 75%, while controlling the accuracy loss within 0.1%, supporting microsecond-level real-time reasoning.
[0220] 7. Distributed Arithmetic (DA)
[0221] Definition: A technique for optimizing fixed-point multiplication on FPGAs. It decomposes the multiplication operation into bitwise operations and lookup table operations. By pre-storing the results of partial product combinations, it uses the FPGA's BRAM to achieve efficient multiplication and accumulation. It is particularly suitable for matrix multiplication in neural networks.
[0222] The present invention achieves: converting fixed-point multiplication in machine learning inference into a DA architecture, consuming only one BRAM instead of a DSP48E for each 16-bit multiplication, improving resource utilization by 40% and reducing inference latency to 20ns.
[0223] 8. AXI bus (Advanced eXtensible Interface)
[0224] Definition: A high-performance on-chip bus protocol used for high-speed data transmission between modules within the FPGA. It supports features such as burst transmission and out-of-order completion and is a key interface for achieving dynamic hardware reconfiguration.
[0225] The present invention is applied to dynamically switch numerical methods (such as RK4 to trapezoidal method) within 200ns via the AXI bus, reconstruct the hardware configuration of the discretization unit, and realize adaptive processing of rigid / non-rigid systems.
[0226] 9. BRAM (Block Random Access Memory)
[0227] Definition: Block random access memory (BRAM) built into FPGAs provides high-bandwidth, low-latency on-chip storage and is commonly used for caching data, implementing lookup tables (such as activation function LUTs), or storing neural network weights.
[0228] The present invention optimizes: sharing BRAM to store base value tables and neural network weights, reducing storage resource usage by 30%, and at the same time using BRAM to implement distributed arithmetic lookup tables to accelerate fixed-point multiplication operations.
[0229] 10. Embedded Error Estimation
[0230] Definition: In numerical methods, local truncation error is estimated by the difference in the results of using different-order algorithms (such as high-order and low-order) on the same system without any additional computational overhead. It is the core basis of variable-step-size algorithms.
[0231] 11. Implicit Trapezoidal Method
[0232] Definition: A second-order implicit numerical integration method that iteratively solves the equation containing the next state, has good stability, is suitable for rigid systems, and has a local truncation error of O(h 3 ).
[0233] Scenario of the present invention: When dealing with rigid working conditions such as low-temperature charging of lithium batteries and sudden changes in motor load, it automatically switches to the implicit trapezoidal method to avoid the calculation explosion caused by the explicit method due to the small step size, while ensuring numerical stability.
[0234] 12. Explicit High-Order Method
[0235] Definition: A numerical method that can directly solve the state at the next moment without iteration, such as the 4th-order Runge-Kutta method (RK4). It has high computational efficiency but poor stability and is suitable for non-rigid systems. The local truncation error is O(h 5 ).
[0236] Application of the present invention: In the constant voltage charging stage where the system state changes smoothly, the RK4 pipeline is used to accelerate the calculation, which reduces the calculation amount by 50% compared with the Euler method while maintaining a high accuracy of 0.18%.
[0237] Example 4
[0238] This embodiment also provides an electronic device, referring to Figure 2 , includes a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.
[0239] Specifically, the processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits for implementing the embodiments of the present invention.
[0240] Memory 404 may include a large-capacity memory 404 for data or instructions. By way of example, and not limitation, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to the data processing device. In certain embodiments, memory 404 is non-volatile memory. In certain embodiments, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. In appropriate circumstances, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM may be a fast page mode dynamic random access memory 404 (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0241] The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402 .
[0242] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any one of the FPGA real-time simulation methods based on adaptive base value and variable step size in the above embodiments.
[0243] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .
[0244] Transmission device 406 can be used to receive or transmit data via a network. Specific examples of such networks may include wired or wireless networks provided by the electronic device's communications provider. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0245] The input / output device 408 is used to input or output information.
[0246] Example 5
[0247] This embodiment also provides a readable storage medium, in which a computer program is stored. The computer program includes program code for controlling a process to execute a process. The process includes the FPGA real-time simulation method based on adaptive base value and variable step size according to embodiment one.
[0248] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be repeated here.
[0249] In general, various embodiments may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0250] The embodiments of the present invention may be implemented by computer software that is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros may be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer executable components that are configured to perform an embodiment when the program is run. One or more computer executable components may be at least one software code or a portion thereof. In addition, it should be noted at this point that, for example, Figure 1 Any block of the logic flow in the program may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or memory blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc. Physical media are non-transitory media.
[0251] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0252] The above embodiments merely illustrate several embodiments of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of the present invention. Therefore, the scope of the present invention shall be determined by the appended claims.
Claims
1. An FPGA real-time simulation method based on adaptive base value and variable step size, characterized in that: The following steps are involved: By training historical data through machine learning models, a mapping relationship between system states and per-unit base values is established to generate a base value prediction model. The per-unit base value is the benchmark value for converting physical quantities in differential equations into dimensionless per-unit values. The base value prediction model is used to output an initial base value, and the initial base value is dynamically fine-tuned by a controller including proportional, integral, and differential feedback links to obtain a final base value; A high-order numerical method with error estimation is used to discretize the per-unit differential equation, and the simulation step size is automatically adjusted according to the dynamic range of the base value. The discretized differential equations are converted into hardware description language based on fixed-point operations and deployed to run on FPGA.
2. The FPGA real-time simulation method based on adaptive base value and variable step size according to claim 1, characterized in that: Dynamically fine-tune the initial value of the base value through a controller that includes proportional, integral, and differential feedback links, including: Based on the absolute error, cumulative error and error change rate of the current physical quantity, the initial value of the base value is adjusted through a linear combination of proportional term, integral term and differential term.
3. The FPGA real-time simulation method based on adaptive base value and variable step size according to claim 1, characterized in that: Automatically adjust simulation step size based on base value dynamic range and error tolerance, including: The local truncation error is estimated using the difference between the calculation results of high-order numerical methods and low-order numerical methods; The simulation step size is dynamically expanded or reduced according to the local truncation error, the preset error tolerance and the base value adjustment range, wherein the step size adjustment range is limited to a range of 0.1 times to 2 times the current step size.
4. The FPGA real-time simulation method based on adaptive base value and variable step size according to claim 1, characterized in that: The machine learning model adopts a lightweight neural network architecture, quantizes parameters into fixed-point numbers through model compression technology, and the inference delay deployed on FPGA is ≤50ns.
5. The FPGA real-time simulation method based on adaptive base value and variable step size according to any one of claims 1 to 4, characterized in that: High-order numerical methods include the Runge-Kutta method and the implicit trapezoidal method, which are automatically switched based on the system dynamics: When the real part difference of the system eigenvalues satisfies the rigidity condition, the implicit trapezoidal method is used; When the system is non-stiff, explicit high-order methods are used.
6. A device equipped with the FPGA real-time simulation method based on adaptive base value and variable step size according to any one of claims 1 to 4, characterized in that: include: Machine learning inference unit: used to run the lightweight base value prediction model and output the initial value of the base value; Dual-loop control unit: includes an inner-loop PID controller and an outer-loop variable step-size decision maker. The PID controller is used to fine-tune the base value based on error feedback, and the variable step-size decision maker is used to adjust the step size according to the error estimate. Dynamic discretization unit: supports dynamic switching of high-order numerical methods, integrating the Runge-Kutta method pipeline and implicit trapezoidal method iterative solver; Fixed-point arithmetic unit: used to convert differential equations into hardware description language and perform parallel calculations.
7. The device according to claim 6, characterized in that The machine learning inference unit uses distributed arithmetic to optimize fixed-point multiplication, includes an input buffer, a pipeline multiplication and accumulation unit, and an activation function lookup table, and supports 16-bit fixed-point parallel inference.
8. The device according to claim 6, wherein The dual-loop control unit shares an error calculation module, the PID controller periodically updates a base value, and the variable step size decision maker periodically adjusts the step size.
9. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to run the computer program to execute the FPGA real-time simulation method based on adaptive base value and variable step size according to any one of claims 1 to 5.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process, and the process includes the FPGA real-time simulation method based on adaptive base value and variable step size according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method for realizing real-time simulation on FPGA (Field Programmable Gate Array) by using fixed-point operation
CN118260026A
Rail transit equipment fault monitoring method based on artificial intelligence
CN118734080A
Method for realizing real-time simulation on FPGA (Field Programmable Gate Array) by using fixed-point operation
CN119476156A