A Fast Solving Method for 5-Degree-of-Freedom Space Motion Equation

Through dynamic configuration of data paths and frequency multiplication of the operator frequency and time-division multiplexing, combined with fourth-order Runge-Kutta and fourth-order Adams prediction-correction operations, the problem of slow solution speed of spatial motion equations is solved, and the rapid solution effect of hardware acceleration is achieved.

CN115079998BActive Publication Date: 2025-07-22HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210687363.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-24
Filing Date
2022-06-16
Publication Date
2025-07-22
Estimated Expiration
2042-06-16

AI Technical Summary

Technical Problem

In the prior art, the solution speed of spatial motion equations is slow and cannot meet the real-time requirements, and the hardware acceleration method cannot be directly applied to this field.

Method used

The dynamic configuration of data paths and the frequency multiplication method of the operator frequency multiplication method, combined with the fourth-order Runge-Kutta and fourth-order Adams prediction-correction operation, the rapid solution of spatial motion equations is achieved through hardware acceleration.

Benefits of technology

Faster solution speed is achieved on limited resources, improving the system's resource utilization and computing performance, reducing the complexity of storage control logic, and improving the system's working efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115079998B_ABST
    Figure CN115079998B_ABST
Patent Text Reader

Abstract

The present application discloses a method for quickly solving a 5-degree-of-freedom space motion equation. Among them, an iteration manager is used to trigger an algorithm calculation unit to perform iterative calculations; a data path dynamic configuration module in the algorithm calculation unit is provided with a plurality of arithmetic units. The data path dynamic configuration module is used to configure the paths between the plurality of arithmetic units according to the state output by the state manager, and input the operation results into the space motion equation calculation unit; the space motion equation calculation unit is used to respectively input the calculated derivative values into the algorithm calculation unit and input the calculated eigenvalue into the iteration manager to enable the state manager to perform state conversion; the iteration manager is further used to send an iteration termination instruction to the algorithm calculation unit to output the operation results when it is determined that the iteration termination condition is satisfied. Through the technical solution in the present application, the acceleration effect of solving the space motion equation is improved by means of hardware acceleration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of digital signal processing. Specifically, it relates to a method for quickly solving a 5-degree-of-freedom space motion equation. Background Art

[0002] As a kind of high-order ordinary differential equation, the space motion equation is one of the key means to study the orbital mechanics of traditional aircraft. Calculating relatively accurate numerical solutions of the space motion equation (SME) by using a suitable method plays an irreplaceable role in the orbital control of launch vehicles. Due to the characteristics of multiple variables, high complexity, large number of iterations, and high-precision requirements of the SME, in the equation solving process, the huge amount of computation leads to a slow calculation speed. Therefore, how to improve the solving speed of the space motion equation has become an important research topic.

[0003] In the prior art, although the ordinary differential equation can be solved by using the hardware acceleration method, which has been applied in many fields and obtained obvious acceleration effects, due to the different application backgrounds and equation structure characteristics of its equations from those of the space motion equation, such hardware solving methods cannot be applied to the solution of the space motion equation.

[0004] Therefore, the existing solutions of the space motion equation are still based on the optimization algorithms and equation simplification methods of software calculation, and the acceleration effects brought by these methods in large-scale operations are not optimistic, and still cannot meet the real-time requirements in practical applications. Summary of the Invention

[0005] The purpose of this application is to: how to improve the acceleration effect of solving the space motion equation by using the hardware acceleration method and improve the real-time performance of the solution.

[0006] The technical solution of this application is: a method for quickly solving a 5-degree-of-freedom space motion equation is provided. The method includes: Step 1, after all the data to be operated on are received, perform state initialization; Step 2, read the data to be operated on, perform state iterative conversion, and according to the operation rules and the converted state, configure paths for multiple arithmetic units to calculate the operation result, where the operation rule is one of the fourth-order Runge-Kutta operation and the fourth-order Adams predictor-corrector operation; Step 3, determine whether the iteration termination condition is satisfied. If so, output the operation result. If not, use the operation result as the intermediate operation data for the next iteration, perform state iterative conversion, and execute Step 2.

[0007] In any of the above technical solutions, further, when the operation rule is the fourth-order Adams predictor-corrector operation, step 3 includes: Step 31, after receiving the intermediate operation data, convert the state iteration to state S_A1. According to state S_A1 and the fourth-order Adams predictor-corrector operation, configure the paths of multiple arithmetic units and calculate the first intermediate operation data. Specifically, record the calculation result of the fourth-order Runge-Kutta operation corresponding to state S_M4 as the intermediate operation data, convert state S_M4 to state S_A1, and at the same time refresh the link relationship between its internal arithmetic units, reconfigure multiple arithmetic units, calculate the algorithm operation task corresponding to the current state S_A1, calculate the first intermediate operation data, and perform the spatial motion equation calculation. Then reconfigure multiple arithmetic units again, calculate the algorithm operation task corresponding to the current state S_A2, calculate the second intermediate operation data, and judge the iteration termination condition again; Step 32, perform the spatial motion equation calculation according to the first intermediate operation data to generate intermediate eigenvalues, and convert the state iteration to state S_A2; Step 33, configure the paths of multiple arithmetic units according to state S_A2 and the fourth-order Adams predictor-corrector operation; Step 34, use the configured multiple arithmetic units to operate on the first intermediate operation data and the intermediate eigenvalues to generate the second intermediate operation data; Step 35, determine that the iteration termination condition is satisfied, and output the second intermediate operation data as the operation result.

[0008] The beneficial effects of this application are as follows:

[0009] The technical solution in this application adopts the folding technology. Through two methods of dynamic configuration of the data path and frequency doubling and time-sharing multiplexing of arithmetic units, it saves arithmetic unit resources as much as possible and realizes complex designs on limited resources. At the same time, it adopts a pipeline design structure to calculate as many equations as possible under a fixed operation cycle, fully improving the resource utilization rate and operation performance of the system and enhancing the actual working efficiency of the system. Specifically, its beneficial effects are as follows:

[0010] 1. This application adopts the fourth-order Adams predictor-corrector method with a faster solution speed and uses the fourth-order Runge-Kutta method as an auxiliary algorithm. By adopting the folding technology, it completes the task mapping and dynamic switching of the two algorithms on limited hardware resources, solves the spatial motion equation faster, and realizes the full reuse of algorithm arithmetic unit resources;

[0011] 2. This application adopts the method of dynamic configuration of the data path. Combined with the folded structure of the algorithm, it dynamically configures different operation links of the two algorithms in multiple operation states, reducing the demand of the algorithm for arithmetic unit resources;

[0012] 3. The present application uses a highly efficient and reusable memory. When the fourth-order Adams predictor-corrector method runs, by polling and configuring 4 groups of memory access requests and memory physical ports, the memory access frequency of Memory data and the complexity of the storage control logic are reduced, reducing the pressure on FPGA placement and routing;

[0013] 4. The present application uses a folding technique to achieve frequency doubling and time-sharing multiplexing of the arithmetic unit, saving the arithmetic unit resources for square root, division, trigonometric functions, inverse trigonometric functions, and matrix multiplication. At the same time, it effectively reduces the calculation delay of the space motion equation, effectively increases the working main frequency of the system and the system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and / or additional advantages of the present application will become obvious and easy to understand in the description of the embodiments in conjunction with the following drawings, where:

[0015] Figure 1 is a schematic block diagram of a 5-degree-of-freedom space motion equation fast solving device according to an embodiment of the present application;

[0016] FIG. 2(a) is a schematic diagram of a data path configuration scheme in the S_M1 state according to an embodiment of the present application;

[0017] FIG. 2(b) is a schematic diagram of a data path configuration scheme in the S_M2 state according to an embodiment of the present application;

[0018] FIG. 2(c) is a schematic diagram of a data path configuration scheme in the S_M3 state according to an embodiment of the present application;

[0019] FIG. 2(d) is a schematic diagram of a data path configuration scheme in the S_M4 state according to an embodiment of the present application;

[0020] FIG. 2(e) is a schematic diagram of a data path configuration scheme in the S_A1 state according to an embodiment of the present application;

[0021] FIG. 2(f) is a schematic diagram of a data path configuration scheme in the S_A2 state according to an embodiment of the present application;

[0022] Figure 3 is a task state jump relationship diagram of a state manager according to an embodiment of the present application;

[0023] Figure 4 is a task mapping scheme of a fourth-order Runge-Kutta method according to an embodiment of the present application;

[0024] Figure 5 is a task mapping scheme of a fourth-order Adams predictor-corrector method according to an embodiment of the present application;

[0025] Figure 6 is a scatter plot of the relative error of the software and hardware calculation results of variable x according to an embodiment of the present application;

[0026] Figure 7 is a scatter plot of the relative error of the software and hardware calculation results of variable y according to an embodiment of the present application;

[0027] Figure 8 is a scatter plot of the relative error of the software and hardware calculation results of variable v according to an embodiment of the present application. Detailed implementation manners

[0028] In order to more clearly understand the above objects, features and advantages of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.

[0029] In the following description, many specific details are set forth in order to fully understand the present application. However, the present application may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present application is not limited by the specific embodiments disclosed below.

[0030] Embodiment 1:

[0031] As Figure 1 shown, this embodiment provides a device for quickly solving a 5-degree-of-freedom space motion equation, aiming to improve the calculation speed of a large number of space motion equations. On limited FPGA resources, a hardware solver capable of accelerating the solution of the space motion equation is designed, while ensuring that the arithmetic unit has a faster operation speed and acceleration effect.

[0032] The device includes: an algorithm calculation unit 1, a space motion equation calculation unit 3, and an iteration manager 4; wherein, the device further includes an AXI data bus and a PCI-E interface for data transmission, and a storage module 2 for storing data. The storage module 2 includes: an initialization data memory 21, a DDR SDRAM module 22, and a highly multiplexed memory 23.

[0033] In this embodiment, the algorithm calculation unit 1 includes a data path dynamic configuration module 6 and a state manager 5. A plurality of arithmetic units are provided in the data path dynamic configuration module 6. The data path dynamic configuration module 6 is used to configure the paths between the plurality of arithmetic units according to the state output by the state manager 5, and input the operation results into the space motion equation calculation unit 3. Among them, the arithmetic unit at least includes an adder and a multiplier. The algorithm calculation unit 1 is configured to perform four fourth-order Runge-Kutta operations and two fourth-order Adams predictor-corrector operations in sequence by an iterative method.

[0034] Specifically, the algorithm calculation unit 1 sequentially reads the initial values from the initialization data memory 21 or the highly multiplexed memory 23, completes the operation tasks of two algorithms, namely the fourth-order Runge-Kutta method and the fourth-order Adams predictor-corrector method, and transmits the results to the spatial motion equation calculation unit 3, which accelerates the calculation of the 5-degree-of-freedom spatial motion equation.

[0035] Specifically, after completing one iteration, the algorithm calculation unit 1 writes the iteration result into the DDR SDRAM module 22, and at the same time, according to the termination instruction of the iteration manager 4, determines whether the iteration result corresponding to the initial value is written into the initialization data memory 21; among them, the data written into the DDR SDRAM module 22 is stored as the operation result, and the data written into the initialization data memory 21 is used as the input for the next iteration.

[0036] In this embodiment, by reconstructing multiple arithmetic units in the algorithm calculation unit 1 and adopting the folded calculation structure, only one algorithm operation unit and one spatial motion equation calculation unit 3 are required to complete the operation task mapping of two solution methods, namely the fourth-order Runge-Kutta operation and the fourth-order Adams predictor-corrector operation. The two mapping schemes are respectively as Figure 4 and Figure 5 shown.

[0037] Furthermore, the state manager 5 is used to define multiple algorithm running states according to the folded sets after folding the two algorithms, dynamically and precisely switch the states by monitoring the algorithm running process in real time, and output the states to the data path dynamic configuration module 6 in the algorithm calculation unit 1.

[0038] Among them, as Figure 3 shown, the state manager 5 includes 7 states: an initialization state, four states of the fourth-order Runge-Kutta method (S_M1 to S_M4), and two states of the fourth-order Adams predictor-corrector method (S_A1 and S_A2).

[0039] The data path dynamic configuration module 6 designs the connection relationships of multiple arithmetic units in each state according to the multiple states in the state manager 5, and configures the connection relationships of multiple arithmetic units according to the states provided by the state manager 5 during the algorithm running process to establish the operation links in different states.

[0040] In this embodiment, the spatial motion equation calculation unit 3 is used to calculate the derivative value and the eigenvalue, and inputs the calculated derivative value into the algorithm calculation unit 1 and the eigenvalue into the iteration manager 4 to enable the state manager 5 to perform state conversion.

[0041] Specifically, the spatial motion equation calculation unit 3 is responsible for calculating the 5-degree-of-freedom spatial motion equation. For several arithmetic units such as square root operation, trigonometric function operation, inverse trigonometric function operation, division operation, and matrix multiplication operation, folding technology is adopted, and time-sharing multiplexing is realized through ping-pong operation after the frequency doubling of the arithmetic unit. The remaining operations and controls adopt the base frequency design. By the method of frequency doubling and time-sharing multiplexing, the resource consumption of this module is reduced. The spatial motion equation calculation unit 3 adopts pipeline design technology, and the interior is composed of multiple pipelines to form a complete pipeline structure. The pipeline depth is 260, that is, the spatial motion equation calculation unit 3 can calculate up to 260 equations with different initial values in batches.

[0042] In this embodiment, the data path dynamic configuration module 6 is configured to, in the manner of a folding calculation structure, according to the state output by the state manager 5, refer to Figure 4 and Figure 5 the operation rules in the operation task sub-state mapping scheme shown, configure the paths between multiple arithmetic units, where the operation rules are fourth-order Runge-Kutta operation and fourth-order Adams predictor-corrector operation.

[0043] Furthermore, the multiple arithmetic units in the data path dynamic configuration module 6 include adders and multipliers. When the algorithm calculation unit 1 performs the fourth-order Runge-Kutta operation, the state manager 5 is configured to sequentially output four different states of S_M1, S_M2, S_M3, and S_M4. The data path dynamic configuration module 6 is configured to configure the paths between multiple arithmetic units according to the state output by the state manager 5, specifically including:

[0044] As shown in FIG. 2(a), in the S_M1 state: Using three groups of multipliers and four groups of adders, respectively complete the operation tasks and a total of 5 operations of M1 + 0;

[0045] Among them, the output end of the fifth group of multipliers 205 is connected to the input end of the fifth group of adders 215 to complete the calculation;

[0046] The output end of the seventh group of multipliers 207 is connected to the input end of the sixth group of adders 216 to complete the calculation;

[0047] The eighth group of multipliers 208 completes the operation, the fourth group of adders 214 completes the calculation of M1 + 0, and the ninth group of adders 219 completes

[0048] It should be noted that during the path configuration process, the input ends of the adders and multipliers that are not otherwise occupied are used to input source data, such as h / 2, which will not be elaborated here.

[0049] As shown in Fig. 2(b), in the S_M2 state: Using four sets of multipliers and six sets of adders, respectively complete a total of 6 operations of M_sum_y + M2 and M_sum_y + 2M2;

[0050] Among them, the output end of the sixth set of multiplier 206 is connected to the input end of the fifth set of adder 215, and the output end of the fifth set of adder 215 is connected to the input end of the eighth set of adder 218 to complete the calculation of;

[0051] The output end of the seventh set of multiplier 207 is connected to the input end of the sixth set of adder 216 to complete the calculation of;

[0052] The output end of the eighth set of multiplier 208 is connected to the input end of the seventh set of adder 217 to complete the calculation of M_sum_y + 2M2;

[0053] The fifth set of multiplier 205 completes the calculation of, the fourth set of adder 214 completes the calculation of M_sum_y + M2, and the ninth set of adder 219 completes the calculation of;

[0054] As shown in Fig. 2(c), in the S_M3 state: Using four sets of multipliers and six sets of adders to complete t n + h, y n + hz n + Δy_M, z n + hM3, z0 * h, M_sum_y + M3 and M_sum_y + 2M3 6 operations;

[0055] Among them, the output end of the fifth set of adder 215 and the output end of the sixth set of multiplier 206 are respectively connected to the two input ends of the eighth set of adder 218 to complete y n + hz n + Δy_M's calculation;

[0056] The output end of the seventh set of multiplier 207 is connected to the input end of the sixth set of adder 216 to complete z n + hM3's calculation;

[0057] The output end of the eighth set of multiplier 208 is connected to the input end of the seventh set of adder 217 to complete the calculation of M_sum_y + 2M3;

[0058] The fifth set of multiplier 205 completes the calculation of z0 * h, the fourth set of adder 214 completes the calculation of M_sum_y + M3. The ninth set of adder 219 completes t nCalculation of +h;

[0059] As shown in Fig. 2(d), in the S_M4 state: Two sets of multipliers and four sets of adders are used to complete M_sum_z = M_sum_z + M4, and y n +Δy_z + M_sum_y for four operations;

[0060] Among them, the output terminal of the fourth set of adder 214 is connected to the input terminal of the seventh set of multiplier 207, and the output terminal of the seventh set of multiplier 207 is connected to the input terminal of the sixth set of adder 216 to complete M_sum_z = M_sum_z + M4 and the calculation;

[0061] The output terminal of the fifth set of multiplier 205 is connected to the input terminal of the fifth set of adder 215, and the output terminal of the fifth set of adder 215 is connected to the input terminal of the eighth set of adder 218 to complete and y n +Δy_z + M_sum_y calculation.

[0062] Among the inputs required by the adders and multipliers in the four states of S_M1 to S_M4, the input data M1, M2, M3, and M4 come from the spatial motion equation calculation unit 3, and the rest of the inputs come from the memory; the outputs of the multipliers and adders vary according to the state. Specifically, in the three states of S_M1 to S_M3, the variables y, z, and t are output to the spatial motion equation calculation unit 3, and the rest of the outputs are written into the memory 23; in the S_M4 state, the variables y, z, and t are output to the spatial motion equation calculation unit 3 and also written into the memory 23, and the rest of the outputs are written into the memory 23.

[0063] Further, the arithmetic unit includes an adder and a multiplier, and the state manager 5 is configured to output the S_A1 state and the S_A2 state in sequence. Among them, in the S_A1 state and the S_A2 state, the arithmetic unit in the data path dynamic configuration module 6 is configured with the same path (link relationship). In the S_A1 state, the data path dynamic configuration module 6 is configured to: form a first intermediate unit with three adders and three multipliers. Among them, the output ends of the first group of multipliers 201 and the second group of multipliers 202 are respectively connected to two input ends of the first group of adders 211, the output end of the third group of multipliers 203 is connected to one input end of the second group of adders 212, the output ends of the first group of adders 211 and the second group of adders 212 are respectively connected to two input ends of the third group of adders 213, and the output end of the third group of adders 213 is the output end of the first intermediate unit; the output end of the first intermediate unit and the output end of the fourth group of multipliers 204 are respectively connected to the input end of the fourth group of adders 214; form a second intermediate unit with three adders and three multipliers, where the second intermediate unit has the same structure as the first intermediate unit; the output end of the eighth group of multipliers 208 is connected to one input end of the seventh group of adders 217.

[0064] Specifically, as shown in FIGS. 2(e) and 2(f), the output ends of the first group of multipliers 201 and the second group of multipliers 202 are connected to the input ends of the first group of adders 211, the output end of the third group of multipliers 203 is connected to the input end of the second group of adders 212, the output ends of the first group of adders and the second group of adders 212 are connected to the input ends of the third group of adders 213, and the output end of the third group of adders 213 and the output end of the fourth group of multipliers 204 are connected to the input end of the fourth group of adders 214 to complete or the operation;

[0065] The output ends of the fifth group of multipliers 205 and the sixth group of multipliers 206 are connected to the input ends of the fifth group of adders 215, the output end of the seventh group of multipliers 207 is connected to the input end of the sixth group of adders 216, and the output ends of the fifth group of adders 215 and the sixth group of adders 216 are connected to the input end of the eighth group of adders 218 to complete or the calculation;

[0066] The output end of the eighth group of multipliers 208 is connected to the input end of the seventh group of adders 217 to complete or the calculation.

[0067] In the two states of S_A1 and S_A2, among the inputs required by the adder and the multiplier, f(n) and The rest of the inputs come from the memory, while the input related to the spatial motion equation calculation unit 3 comes from the said unit 3. The outputs of the multiplier and adder are determined by the state.

[0068] Specifically, in the S_A1 state, the variables y, z, and t are output to the spatial motion equation calculation unit 3, and the rest of the outputs are written into the memory 23. In the S_A2 state, the variables y, z, and t are output to the spatial motion equation calculation unit 3 and also written into the memory 23, while the rest of the outputs are written into the memory 23.

[0069] In this embodiment, the initialization data memory 21 is responsible for storing data such as the batch operation quantity received from the PCI-E interface, the initial values of the equations required for the batch operation of the spatial motion equation, all the parameters required for the equation operation, the reference values for judging the termination of the equation iteration, and the transmission completion instruction; and storing the results of each iteration during the operation of the algorithm calculation unit 1.

[0070] The DDR SDRAM module 22 is responsible for saving the results of each iteration during the equation calculation process, and after all the batch operations are completed, reading the data and outputting it through the PCI-E interface.

[0071] The highly efficient and multiplexed memory 23 is used to poll and configure the four groups of memory access commands and the four groups of memory ports of the fourth-order Adams predictor-corrector method by means of polling arbitration, dynamically update the corresponding relationship between the instructions and the physical ports according to the iteration process, and realize the data refresh on the premise of avoiding frequent data transfer.

[0072] In this embodiment, the iteration manager 4 is used to trigger the algorithm calculation unit 1 to perform iterative calculation; the iteration manager 4 is also used to send an iteration termination instruction to the algorithm calculation unit 1 when it is determined that the iteration termination condition is satisfied, so that the algorithm calculation unit 1 outputs the operation result.

[0073] Specifically, the iteration manager 4 compares the operation result of each iteration of the algorithm calculation unit 1 and the eigenvalue calculated by the spatial motion equation calculation unit 3 with multiple iteration termination conditions in the initialization data memory 21, and judges in real time whether the equation solution satisfies the iteration termination condition and generates an iteration termination instruction.

[0074] The iteration termination conditions in this embodiment can be set as one or more of the following three conditions: the acceleration a reaches a certain value, the height h reaches a certain value, and the speed v reaches a certain value.

[0075] According to the monitoring instruction transmitted by the algorithm calculation unit 1 after the iteration is completed, the iteration manager 4 performs a multi-parameter mixed judgment on the iteration operation result of the algorithm calculation unit 1 and the eigenvalue calculated by the spatial motion equation calculation unit 3. Among them, the iteration result includes velocity, and the eigenvalue includes height and acceleration, to determine whether the iteration process of each set of initial values meets the termination condition. If it meets, the termination instruction of this set of initial values is set to 1 and transmitted to the algorithm calculation unit 1; at the same time, the termination instruction of this set of initial values is recorded to facilitate the judgment of whether all the batch initial values have been iterated; if it does not meet, the termination instruction remains unchanged.

[0076] The iteration manager 4 also needs to continuously judge in real time whether the termination instructions of all initial values are all 1 when the monitoring instruction is valid. If all are 1, the calculation is completed, a finish signal is transmitted to the algorithm calculation unit 1, the DDR SDRAM 22 read operation is started, and the data is output through the PCI-E interface as the solution result of the five-degree-of-freedom spatial motion equation for this time; if not, the calculation continues.

[0077] Furthermore, the device also includes: a PCI-E interface and a register; the PCI-E interface is used to receive the data to be operated; the register is used to store the data to be operated; the iteration manager 4 is also used to judge whether the value in the register is within the preset interval [1, 260]. If so, it triggers the algorithm calculation unit 1 to read the data to be operated from the register; if not, it controls the PCI-E interface to terminate receiving the data to be operated.

[0078] Embodiment 2:

[0079] This embodiment provides a method for quickly solving a five-degree-of-freedom spatial motion equation, and the method includes:

[0080] Step 1, when all the data to be operated is received, perform state initialization;

[0081] Specifically, data is received through the PCI-E interface, and the received data includes the quantity of batch operations, the equation initial values required for batch operations of the spatial motion equation, all parameters required for equation operations, the reference value for equation iteration termination judgment, the iteration step size, and the transmission termination instruction, etc.; after the PCI-E receives the correct data, the quantity of batch operations is stored in the register pipe_deepth, and the other received data is stored in the initialization data memory.

[0082] After receiving the quantity of batch operations, first determine whether the batch value is legal according to a preset reference range, that is, determine whether the value of the register pipe_deepth is within the interval [1, 260]. If it is within the interval, the batch number meets the requirements and other data can be continuously received; if it is not within the interval, send an error request through the PCI-E interface, terminate receiving data, reset the relevant control logic, and wait for the PCI-E interface to resend the correct data.

[0083] Step 2, read the data to be operated, perform state iterative conversion, and configure paths for multiple arithmetic units according to the operation rules and the converted state, and calculate the operation result.

[0084] Among them, the operation rule is one of the fourth-order Runge-Kutta operation and the fourth-order Adams predictor-corrector operation.

[0085] Specifically, when all the data to be operated is received and a transmission termination instruction is received, trigger the calculation of the 5-degree-of-freedom space motion equation and start the iterative calculation of the equation; initialize the iteration counter n = 0, set the state to INIT, and set the termination signals of all initial values to 0.

[0086] As Figure 4 shown, during the process of state iterative conversion, enter the first state S_M1 from the INIT state, read the data required for the current iteration from the initialization data memory 21 according to the S_M1 state, complete the relevant operations using the configured multiple arithmetic units, and perform space motion equation calculation according to the relevant calculation results to obtain eigenvalues.

[0087] When configuring multiple arithmetic units, use three groups of multipliers and 4 groups of adders to complete the operation tasks respectively and a total of 5 operations including M1 + 0; among them, the output end of the fifth group of multipliers 205 is connected to the input end of the fifth group of adders 215 to complete the calculation; the output end of the eighth group of multipliers 208 is connected to the input end of the sixth group of adders 216 to complete the calculation; the eighth group of multipliers 208 completes the operation, the fourth group of adders 214 completes the calculation of M1 + 0, and the ninth group of adders 219 completes

[0088] During the calculation of the spatial motion equation, corresponding parameters are read from the initialization data memory 21, and the iteration termination is judged according to the calculated eigenvalue. If the iteration termination condition is not satisfied, the state S_M1 is switched to the state S_M2, the link relationship between its internal arithmetic units is refreshed, the configuration of multiple arithmetic units is re-performed, the algorithm operation task in the current state S_M2 is calculated, the spatial motion equation is calculated, and the iteration termination condition is judged again.

[0089] In the state S_M2, multiple arithmetic units are configured as follows: the output end of the sixth group multiplier 206 is connected to the input end of the fifth group adder 215, and the output end of the fifth group adder 215 is connected to the input end of the eighth group adder 218 to complete the calculation of; the output end of the seventh group multiplier 207 is connected to the input end of the sixth group adder 216 to complete the calculation of; the output end of the eighth group multiplier 208 is connected to the input end of the seventh group adder 217 to complete the calculation of M_sum_y + 2M2; the fifth group multiplier 205 completes the calculation of, the fourth group adder 214 completes the calculation of M_sum_y + M2, and the ninth group adder 219 completes the calculation of.

[0090] If the iteration termination condition is not satisfied, the state S_M2 is switched to the state S_M3, the link relationship between its internal arithmetic units is refreshed, the configuration of multiple arithmetic units is re-performed, the algorithm operation task in the current state S_M3 is calculated, the spatial motion equation is calculated, and the iteration termination condition is judged again.

[0091] In the state S_M3, multiple arithmetic units are configured such that the output end of the fifth group adder 215 is connected to the input end of the eighth group adder 218, and the output end of the sixth group multiplier 206 is connected to the input end of the eighth group adder 218 to complete y n + hz n + Δy_M calculation; the output end of the seventh group multiplier 207 is connected to the input end of the sixth group adder 216 to complete z n + hM3 calculation; the output end of the eighth group multiplier 208 is connected to the input end of the seventh group adder 217 to complete the calculation of M_sum_z + 2M3; the fifth group multiplier 205 completes the calculation of z0 * h; the fourth group adder 214 completes the calculation of M_sum_y + M3; the ninth group adder 219 completes t n + h calculation.

[0092] If the iteration termination condition is not met, the state S_M3 is switched to the state S_M4, the link relationship between its internal arithmetic units is refreshed, the configuration of multiple arithmetic units is re-performed, the algorithm operation task in the current state S_M4 is calculated, the spatial motion equation is calculated, and the iteration termination condition is judged again.

[0093] In the state S_M4, multiple arithmetic units are configured as follows: the output end of the fourth group of adders 214 is connected to the input end of the seventh group of multipliers 207, and the output end of the seventh group of multipliers 207 is connected to the input end of the sixth group of adders 216, completing the calculation of M_sum_z = M_sum_z + M4 and ; the output end of the fifth group of multipliers 205 is connected to the input end of the fifth group of adders 215, and the output end of the fifth group of adders 215 is connected to the input end of the eighth group of adders 218, completing the calculation of and y n +Δy_z + M_sum_y.

[0094] At this time, 1 time of fourth-order Runge-Kutta operation is completed, a monitoring instruction is output to the iteration manager 4, and according to the termination instruction fed back by the iteration manager 4, it is judged whether each group of results needs to be written into the initialization data memory 21; and all operation results are stored in the DDR SDRAM 22;

[0095] When the monitoring instruction is valid, the iteration manager 4 judges in real time whether the iteration of each group of initial values terminates, outputs the termination instruction to the algorithm calculation unit 1, and records the termination instruction; the iteration manager 4 also needs to judge according to the monitoring instruction whether all termination instructions are 1. If they are all 1, it means that the equation operations of the current batch are all completed, and a finish signal is output to the algorithm calculation unit 1.

[0096] In this embodiment, four times of fourth-order Runge-Kutta operations and two times of fourth-order Adams predictor-corrector operations are sequentially performed to complete the solution process of the 5-degree-of-freedom spatial motion equation. Therefore, the initial value of the iteration counter n is set to 0. Each time a fourth-order Runge-Kutta operation is executed, the iteration counter n is incremented by one; the size of the iteration counter n is judged. If n is less than 3, the state is switched to S_M1; if n is equal to 3, the fourth-order Adams predictor-corrector operation is executed, and the state is switched to S_A1.

[0097] Step 3, judge whether the iteration termination condition is met. If so, output the operation result. If not, use the operation result as the intermediate operation data for the next iteration, perform state iteration conversion, and execute step 2.

[0098] Further, when the operation rule is the fourth-order Adams predictor-corrector operation, step 3 includes:

[0099] Step 31: After receiving the intermediate operation data, convert the state iteration to state S_A1. According to state S_A1 and the fourth-order Adams predictor-corrector operation, configure the paths of multiple arithmetic units and calculate the first intermediate operation data. Specifically, record the calculation result of the fourth-order Runge-Kutta operation corresponding to state S_M4 as the intermediate operation data, convert state S_M4 to state S_A1, simultaneously refresh the connection relationships between its internal arithmetic units, reconfigure multiple arithmetic units, calculate the algorithm operation tasks corresponding to the current state S_A1, calculate the first intermediate operation data, perform spatial motion equation calculations, reconfigure multiple arithmetic units again, calculate the algorithm operation tasks corresponding to the current state S_A2, calculate the second intermediate operation data, and then determine the iteration termination condition again.

[0100] In state S_A1, multiple arithmetic units are configured as follows: the output ends of the first group of multipliers 201 and the second group of multipliers 202 are connected to the input end of the first group of adders 211, the output end of the third group of multipliers 203 is connected to the input end of the second group of adders 212, the output ends of the first group of adders and the second group of adders 212 are connected to the input end of the third group of adders 213, the output end of the third group of adders 213 and the output end of the fourth group of multipliers 204 are connected to the input end of the fourth group of adders 214 to complete the calculation; the output ends of the fifth group of multipliers 205 and the sixth group of multipliers 206 are connected to the input end of the fifth group of adders 215, the output end of the seventh group of multipliers 207 is connected to the input end of the sixth group of adders 216, the output ends of the fifth group of adders 215 and the sixth group of adders 216 are connected to the input end of the eighth group of adders 218 to complete the calculation; the output single of the eighth group of multipliers 208 is connected to the input end of the seventh group of adders 217 to complete the calculation;

[0101] Step 32: Perform spatial motion equation calculations based on the first intermediate operation data to generate intermediate eigenvalues, and convert the state iteration to state S_A2;

[0102] Step 33: According to state S_A2 and the fourth-order Adams predictor-corrector operation, configure the paths of multiple arithmetic units;

[0103] Specifically, in state S_A2, multiple arithmetic units are configured as follows: the output ends of the first group of multipliers 201 and the second group of multipliers 202 are connected to the input ends of the first group of adders 211, the output end of the third group of multipliers 203 is connected to the input end of the second group of adders 212, the output ends of the first group of adders and the second group of adders 212 are connected to the input end of the third group of adders 213, the output end of the third group of adders 213 and the output end of the fourth group of multipliers 204 are connected to the input end of the fourth group of adders 214 to complete the calculation of; the output ends of the fifth group of multipliers 205 and the sixth group of multipliers 206 are connected to the input ends of the fifth group of adders 215, the output end of the seventh group of multipliers 207 is connected to the input end of the sixth group of adders 216, the output ends of the fifth group of adders 215 and the sixth group of adders 216 are connected to the input end of the eighth group of adders 218 to complete the calculation of; the output single of the eighth group of multipliers 208 is connected to the input end of the seventh group of adders 217 to complete the calculation of;

[0104] Step 34: Use the configured multiple arithmetic units to perform operations on the first intermediate operation data and the intermediate eigenvalue to generate second intermediate operation data;

[0105] Specifically, convert state S_A1 to state S_A2, simultaneously refresh the link relationship between its internal arithmetic units, reconfigure the multiple arithmetic units, calculate the algorithm operation task corresponding to the current state S_A2, and calculate the second intermediate operation data.

[0106] At this time, 1 iteration of the fourth-order Adams predictor-corrector method has been completed, an output monitoring instruction is sent to the iteration manager 4, and according to the termination instruction fed back by the iteration manager 4, it is judged whether each group of results needs to be written into the initialization data storage module 2; and all operation results are stored in the DDR SDRAM 22.

[0107] Step 35: Determine that the iteration termination condition is satisfied, and output the second intermediate operation data as the operation result.

[0108] Specifically, if a finish signal transmitted by the iteration manager 4 is received, after writing the iteration result into the DDR SDRAM module 22, start the read operation of the DDR SDRAM 22, output the read data through PCI-E, and perform a reset; otherwise, according to the second intermediate operation data, re-execute step 31.

[0109] To verify the effectiveness of the technical solution in this embodiment, it is set that in the actual application of this embodiment, there are 260 groups of initial values for a single operation task in the operation, and the number of iteration times varies from 300,000 to 500,000. Select 10 operation tasks, each operation task contains 260 groups of initial values, the range of the number of iteration times of the 260 initial values is less than 10,000, and the number of iteration times of the 2,600 groups of initial values of the 10 tasks is all between 300,000 and 500,000. The experiment proves the acceleration effect of the device proposed by the present invention by comparing the computing time consumption of software and hardware; at the same time, by analyzing the complex error of the final solution results of software and hardware, the reliability of the solution accuracy of the device is proved.

[0110] Perform software and hardware calculations on the selected test sets respectively, and the calculated computing time consumption is statistically shown in Table 1. By averaging the acceleration ratios of the 10 groups of tasks in Table 1, the acceleration ratio of the device proposed by the present invention relative to the traditional software solution is 12.765.

[0111] Table 1

[0112]

[0113] Perform error analysis on the software and hardware calculation results of 2,600 groups of initial values. The calculation results are the displacements x and y and the velocity v respectively, and their relative error analysis is as follows Figures 6 to 8 shown. Through relative error analysis, it is found that the relative error of the hardware solution is within the allowable range, and the reliability of the solution accuracy is relatively high.

[0114] The technical solution of the present application has been described in detail above in conjunction with the accompanying drawings. The present application proposes a device and method for quickly solving a 5-degree-of-freedom space motion equation. Among them, the iteration manager in the device is used to trigger the algorithm calculation unit to perform iterative calculation; the algorithm calculation unit includes a data path dynamic configuration module and a state manager. Multiple arithmetic units are set in the data path dynamic configuration module. The data path dynamic configuration module is used to configure the paths between multiple arithmetic units according to the state output by the state manager, and input the operation results into the space motion equation calculation unit; the space motion equation calculation unit is used to calculate the derivative value and eigenvalue, and input the calculated derivative value into the algorithm calculation unit and input the calculated eigenvalue into the iteration manager to enable the state manager to perform state conversion; the iteration manager is further used to send an iteration termination instruction to the algorithm calculation unit when it is determined that the iteration termination condition is met, so that the algorithm calculation unit outputs the operation result. Through the technical solution in the present application, the acceleration effect of solving the space motion equation is improved by means of hardware acceleration.

[0115] The steps in the present application can be adjusted, combined and deleted according to actual needs.

[0116] The units in the device of the present application can be combined, divided, and deleted according to actual needs.

[0117] Although the present application has been disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and are not intended to limit the application of the present application. The protection scope of the present application is defined by the appended claims and may include various variations, modifications, and equivalent solutions made to the invention without departing from the protection scope and spirit of the present application.

Claims

1. A fast solution method for a 5-degree-of-freedom spatial motion equation, characterized in that, The method includes: Step 1: After all the data to be operated on is received, perform state initialization. Step 2: Read the data to be operated on, perform state iterative conversion, and configure paths for multiple arithmetic units according to the operation rule and the converted state, and calculate the operation result. Wherein, the operation rule is one of the fourth-order Runge-Kutta operation and the fourth-order Adams predictor-corrector operation. Step 3: Determine whether the iteration termination condition is satisfied. If so, output the operation result. If not, use the operation result as the intermediate operation data for the next iteration, perform state iterative conversion, and execute Step 2. When the operation rule is the fourth-order Adams predictor-corrector operation, Step 3 includes: Step 31: After receiving the intermediate operation data, convert the state iteratively to state S_A1, configure paths for multiple arithmetic units according to state S_A1 and the fourth-order Adams predictor-corrector operation, and calculate the first intermediate operation data. Specifically, record the calculation result of the fourth-order Runge-Kutta operation corresponding to state S_M4 as the intermediate operation data, convert state S_M4 to state S_A1, simultaneously refresh the link relationship between its internal arithmetic units, reconfigure multiple arithmetic units, calculate the algorithm operation task corresponding to the current state S_A1, calculate the first intermediate operation data, perform spatial motion equation calculation, reconfigure multiple arithmetic units again, calculate the algorithm operation task corresponding to the current state S_A2, calculate the second intermediate operation data, and determine the iteration termination condition again. Step 32: Perform spatial motion equation calculation based on the first intermediate operation data to generate intermediate eigenvalues, and convert the state iteratively to state S_A2. Step 33: Configure paths for multiple arithmetic units according to state S_A2 and the fourth-order Adams predictor-corrector operation. Step 34: Use the configured multiple arithmetic units to operate on the first intermediate operation data and the intermediate eigenvalues to generate the second intermediate operation data. Step 35: Determine that the iteration termination condition is satisfied, and output the second intermediate operation data as the operation result.

Citation Information

Patent Citations

  • Montgomery modular multiplication device and embedded security chip with same

    CN104793919A

  • NURBS interpolation position parameter optimization and parallel computation hardening method

    CN110780642A