Accelerator device for solving linear programming problem and execution method
By designing an accelerator device including controller module, parallel computing module, data handling module, data distribution module and data receiving module, the problem that existing hardware cannot efficiently run the PDLP algorithm, and efficient solution to large-scale linear planning problems are achieved.
Patent Information
- Application Number
- CN202411877621.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
AI Technical Summary
Existing hardware cannot efficiently run the PDLP algorithm with SpMV operation as the core to solve the general large-scale linear programming problem.
An accelerator device for solving linear programming problems is designed, including a controller module, a parallel computing module, a data transfer module, a data distribution module and a data receiving module. Through the coordinated work of these modules, the parallel processing of SpMV operations and vector operations is realized, meeting the parallelism and randomness requirements of sparse matrices.
It improves the computing efficiency of linear programming problems, can effectively deal with large-scale linear programming problems, and meets the requirements of SpMV operations for parallelism and irregularity.
Smart Images

Figure CN119937983A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hardware accelerators, and in particular relates to an accelerator device and an execution method for solving linear programming problems. Background Art
[0002] Linear programming (LP) is a mathematical method in operations research that is used to optimize (maximize or minimize) a linear objective function under a set of linear inequality or equality constraints. Linear programming problems can usually be expressed in the following standard form:
[0003] Objective function: The linear function to be optimized, which can be cost minimization or profit maximization.
[0004] Constraints: A set of linear inequalities or equalities that define the set of feasible solutions.
[0005] Variables: The decision variables in the problem, which need to take values under constraints.
[0006] The general form of a linear programming problem can be expressed as:
[0007] Maximize f=c T x
[0008] stAx≤b,
[0009] x≥0
[0010] Where f is the objective function, c = {c1, c2, ..., c n} is the coefficient of the objective function, x={x1,x2,…,x n} is the decision variable, A is the coefficient matrix of the constraint condition, b={b1,b2,…,b m} is the constant term of the constraint.
[0011] The application areas of linear programming are very wide, including but not limited to: production planning, transportation problems, inventory management, financial investment, energy distribution, agricultural planning, project management, supply chain management, environmental management, etc.
[0012] Linear programming uses a systematic approach to analyze and solve resource allocation problems and has become an indispensable tool in modern management science and engineering.
[0013] Currently, the main algorithms for solving linear programming problems include the simplex method, the interior point method (IPM), and the first-order method (FOM) based on gradient descent.
[0014] In a linear programming problem, each constraint divides the solution space into two parts; multiple constraints together form a polyhedron in the solution space. A feasible solution to a linear programming problem that satisfies the constraints is located on the surface and inside of the polyhedron, and the optimal solution is at the vertex of the polyhedron.
[0015] The simplex method traverses the vertices of the polyhedron through multiple iterations. In each iteration, the linear equations are solved to obtain the coordinates of one of the vertices of the polyhedron, and the point is judged to be the optimal solution. If it is not the optimal solution, the linear equations are updated for the next iteration, and if it is, the iteration is terminated. The simplex algorithm is simple, but the number of iterations increases exponentially with the size of the problem.
[0016] Unlike the simplex method that traverses the vertices of a polyhedron, the feasible solution generated by each iteration of the interior point method is located inside the polyhedron, and the number of iterations is less than that of the simplex method. However, the iteration of the interior point method requires solving the inverse matrix, and the amount of calculation for a single iteration is large.
[0017] The first-order method uses the gradient descent method to solve the optimal solution, which has the problem of slow convergence speed. However, its core operation is sparse matrix-vector multiplication (SpMV), which has high parallelism and can improve the convergence speed through some improvements. It has great potential in solving large-scale linear programming problems.
[0018] Currently, the algorithms that use the first-order method to solve linear programming problems mainly include SCS, ECLIPSE, PDLP, etc. Among them, the PDLP algorithm has advantages in accuracy and convergence speed.
[0019] The solution of the linear equations required for the simplex method iteration and the inverse matrix calculation required for the interior point method iteration are based on the matrix decomposition algorithm. In addition to storing the original constraint matrix, additional memory is required to store the decomposition matrix. As the scale of the linear programming problem increases, the additional storage overhead of storing the decomposition matrix may exceed the memory limit of the computing platform and make it impossible to solve.
[0020] Since the algorithm complexity of matrix decomposition algorithm is generally O(n 3 ), which takes a lot of time to solve large-scale linear programming problems. At the same time, due to the low parallelism of the matrix decomposition algorithm, its acceleration mainly comes from the improvement of CPU main frequency and IPC, and the former is difficult to improve due to the gradual failure of Moore's Law, and the future development prospects are limited.
[0021] The core calculation of the PDLP algorithm is sparse matrix-vector multiplication (SpMV), which has a computational complexity of O(n) and has no data dependency during the calculation process, resulting in high parallelism. Compared with CPUs with low parallelism, GPUs that are good at large-scale parallel computing are more suitable for performing SpMV operations. However, GPUs have weak control capabilities and cannot adapt to the randomness of sparse matrices, resulting in the inability to efficiently perform SpMV operations. In order to adapt to the parallelism and randomness of SpMV operations, it is necessary to design dedicated hardware accelerators.
[0022] Currently, there is no dedicated hardware accelerator for the PDLP algorithm. Mainstream hardware accelerators mainly optimize the single SpMV operation. In order to execute the PDLP algorithm, the main processor and vector accelerator are required to cooperate to control the algorithm and calculate the vector respectively.
[0023] The accelerator RSQP for solving quadratic programming problems is optimized for SpMV operations with specific sparse structure matrices and is compatible with vector operations. It has a control unit to control the above operations and can run the PDLP algorithm to accelerate specific linear programming problems. However, its optimization effect on SpMV operations with specific sparse structure matrices is poor compared with that of a dedicated SpMV accelerator, and it is not optimized for SpMV operations with general sparse structure matrices, so the acceleration effect is poor when running the PDLP algorithm to solve general linear programming problems.
[0024] In summary, existing hardware cannot efficiently run the PDLP algorithm with SpMV operation as the core to solve general large-scale linear programming problems. Summary of the invention
[0025] The present invention aims to provide an accelerator device and an execution method for solving linear programming problems, so as to solve the technical problem that the existing hardware cannot efficiently run the PDLP algorithm with SpMV operation as the core to solve general large-scale linear programming problems.
[0026] In order to solve the above technical problems, the specific technical solutions of an accelerator device and an execution method for solving linear programming problems of the present invention are as follows:
[0027] An accelerator device for solving linear programming problems includes a controller module, a parallel computing module, a data handling module, a data distribution module and a data receiving module. When the PDLP algorithm starts to be executed, the main processor sends a pre-prepared accelerator control code to the accelerator device to start algorithm acceleration: the controller module is responsible for the overall control of the accelerator, including controlling the parallel computing module to complete SpMV operations and vector operations; controlling the data handling module to carry data from an external memory; and controlling the data distribution module to send data to a computing unit inside the parallel computing module.
[0028] The control data receiving module accumulates the results of the parallel computing module into the final result and writes the data back to the external memory. Furthermore, the controller module finally completes the PDLP algorithm by initiating multiple different parallel operations and notifies the main processor to take away the algorithm results.
[0029] Furthermore, the parallel computing module is used to calculate SpMV operations and vector operations. The interior of the parallel computing module is composed of multiple basic operation units, each basic operation unit includes a floating-point multiplication and accumulator for calculation, a register for temporarily storing operation results, and multiple control logics. Each basic operation unit performs SpMV operations or calculations of an element in vector operations according to the control of the controller module. The data of the basic operation unit comes from the data distribution module, and the calculation results are sent to the data receiving module.
[0030] Furthermore, the data handling module is responsible for carrying the matrix and vector data required for SpMV operations and vector operations from the external memory, and storing the carried data in the internal memory. The internal memory includes SRAM and multiple registers. The SRAM is responsible for storing streaming data such as matrix data, and supports reading and writing multiple floating-point values at the same time. The registers are responsible for storing random data and can read multiple floating-point values at the same time.
[0031] Furthermore, the data distribution module is responsible for sending the data in the data handling module to the parallel computing module, including two types of data distribution: 1. Distribution of streaming data, sending multiple floating-point numbers in the SRAM to the basic computing units of the parallel computing module in sequence; 2. Distribution of random data, sending data at multiple arbitrary positions in the register of the data handling module to different basic computing units of the parallel computing module at the same time according to the control of the controller module. The distribution of random data is realized by the improved Benes network, and the routing algorithm is used to realize arbitrary routing from N inputs to M outputs to meet the parallel random access requirements of the SpMV operation for vectors, where N≤M. Furthermore, the data receiving module is responsible for merging the calculation results of each basic computing unit in the parallel computing module according to the control of the controller module and writing them back to the external memory.
[0032] The present invention also discloses an execution method of an accelerator device for solving linear programming problems, comprising the following steps:
[0033] Step 1: Generate the code of the controller module for a specific problem;
[0034] Step 2: Use the main processor to load the code into the controller module;
[0035] Step 3: Use the host processor to start the controller module to execute the PDLP algorithm through the controller module register interface;
[0036] Step 4: The controller module starts one iteration of the PDLP algorithm;
[0037] Step 5: The controller module controls the data transport module to transport data;
[0038] Step 6: The controller module controls the data distribution module to distribute data;
[0039] Step 7: The controller module controls the parallel computing module to perform SpMV and vector calculations;
[0040] Step 8: The controller module controls the data receiving module to recycle data;
[0041] Step 9: The controller module determines whether the termination condition is met. If not, it returns to step 4; if it is met, it goes to step 10;
[0042] Step 10: The controller module notifies the main processor that the calculation is completed;
[0043] Step 11: The main processor reads the result and ends.
[0044] An accelerator device and execution method for solving linear programming problems of the present invention have the following advantages: the present invention proposes using a dedicated accelerator to run the PDLP algorithm to solve large-scale linear programming problems, which meets the requirements of SpMV operations for parallelism and irregularity and improves operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a diagram of the accelerator device architecture of the present invention;
[0046] Figure 2 This is a diagram of the random data distribution part of the data distribution module of the present invention;
[0047] Figure 3 It is a schematic diagram of the working process of the accelerator device of the present invention. DETAILED DESCRIPTION
[0048] In order to better understand the purpose, structure and function of the present invention, an accelerator device and an execution method for solving linear programming problems of the present invention are further described in detail below in conjunction with the accompanying drawings.
[0049] like Figure 1As shown, an accelerator device for solving linear programming problems of the present invention mainly includes a controller module, a parallel computing module, a data handling module, a data distribution module and a data receiving module. When the PDLP algorithm starts to be executed, the main processor sends the accelerator control code prepared in advance to the accelerator device to start algorithm acceleration: the controller module is responsible for the overall control of the accelerator, including 1. Controlling the parallel computing module to complete SpMV operations and vector operations (multiplication and addition, inner product, etc.); 2. Controlling the data handling module to carry data from the external memory; 3. Controlling the data distribution module to send data to the computing unit inside the parallel computing module; 4. Controlling the data receiving module to accumulate the results of the parallel computing module into the final result and write the data back to the external memory.
[0050] The controller module completes the PDLP algorithm by initiating a variety of parallel operations and notifies the main processor to take away the algorithm results.
[0051] The parallel computing module is mainly used to calculate SpMV operations and vector operations. The parallel computing module includes one or more basic operation units, each of which includes a floating-point multiplication and accumulation unit for calculation, a register for temporarily storing the operation results, and one or more control logics. Each basic operation unit calculates an element in the SpMV operation or vector operation according to the control of the controller module. The data of the basic operation unit comes from the data distribution module, and the calculation results are sent to the data receiving module.
[0052] The data handling module is responsible for handling the matrix and vector data required for SpMV operations and vector operations from the external memory, and storing the handled data in the internal memory. The internal memory includes SRAM and one or more registers. SRAM is responsible for storing streaming data such as matrix data, and supports reading and writing multiple floating-point values at the same time. Registers are responsible for storing random data such as vectors, and can read multiple floating-point values at the same time.
[0053] The data distribution module is responsible for sending the data in the data handling module to the parallel computing module, which mainly includes two types of data distribution: (1) streaming data distribution, sending multiple floating-point numbers in the SRAM to the basic computing units of the parallel computing module in sequence; (2) random data distribution, under the control of the controller module, sending data at multiple arbitrary positions in the register of the data handling module to different basic computing units of the parallel computing module at the same time. The distribution of random data is implemented by the improved Benes network, which can realize arbitrary routing from N inputs to M outputs (N≤M) with a specific routing algorithm to meet the parallel random access requirements of the SpMV operation for vectors. The specific network structure is shown in the attached figure. Figure 2 shown.
[0054] The data receiving module is responsible for merging the calculation results of each basic operation unit in the parallel computing module according to the control of the controller module and writing them back to the external memory.
[0055] like Figure 3 As shown, the steps of executing a PDLP algorithm by the linear programming algorithm accelerator device of the present invention are as follows:
[0056] Step 1: Generate the code for the controller module for your specific problem.
[0057] Step 2: Use the host processor to load the code into the controller module.
[0058] Step 3: Use the host processor to start the controller module to execute the PDLP algorithm through the controller module register interface.
[0059] Step 4: The controller module starts one iteration of the PDLP algorithm.
[0060] Step 5: The controller module controls the data transport module to transport data.
[0061] Step 6: The controller module controls the data distribution module to distribute data.
[0062] Step 7: The controller module controls the parallel computing module to perform SpMV and vector calculations.
[0063] Step 8: The controller module controls the data receiving module to recycle data.
[0064] Step 9: The controller module determines whether the termination condition is met. If not, it returns to step 4; if it is met, it goes to step 10.
[0065] Step 10: The controller module notifies the main processor that the calculation is complete.
[0066] Step 11: The main processor reads the result and ends.
[0067] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.
Claims
1. An accelerator device for solving linear programming problems, characterized in that: It includes a controller module, a parallel computing module, a data handling module, a data distribution module and a data receiving module. When the PDLP algorithm starts to be executed, the main processor sends the pre-prepared accelerator control code to the accelerator device to start algorithm acceleration: the controller module is responsible for the overall control of the accelerator, including controlling the parallel computing module to complete SpMV operations and vector operations; controlling the data handling module to carry data from the external memory; controlling the data distribution module to send data to the operation unit inside the parallel computing module; and controlling the data receiving module to accumulate the results of the parallel computing module into the final result and write the data back to the external memory.
2. The accelerator device for solving linear programming problems according to claim 1, characterized in that: The controller module finally completes the PDLP algorithm by initiating a variety of different parallel operations and notifies the main processor to take away the algorithm results.
3. The accelerator device for solving linear programming problems according to claim 1, characterized in that: The parallel computing module is used to calculate SpMV operations and vector operations. The interior of the parallel computing module is composed of multiple basic computing units. Each basic computing unit includes a floating-point multiplication and accumulation unit for calculation, a register for temporarily storing the calculation results, and multiple control logics. Each basic computing unit performs the calculation of an element in the SpMV operation or the vector operation according to the control of the controller module. The data of the basic computing unit comes from the data distribution module, and the calculation result is sent to the data receiving module.
4. The accelerator device for solving linear programming problems according to claim 1, characterized in that: The data transport module is responsible for transporting the matrix and vector data required for SpMV operations and vector operations from the external memory, and storing the transported data in the internal memory. The internal memory includes SRAM and multiple registers. The SRAM is responsible for storing streaming data such as matrix data and supports reading and writing multiple floating-point values at the same time. The registers are responsible for storing random data and can read multiple floating-point values at the same time.
5. The accelerator device for solving linear programming problems according to claim 1, characterized in that: The data distribution module is responsible for sending the data in the data handling module to the parallel computing module, including two types of data distribution:
1. Distribution of streaming data, sending multiple floating-point numbers in the SRAM to the basic computing unit of the parallel computing module in sequence; 2. Distribution of random data, sending data at multiple arbitrary positions in the register of the data handling module to different basic computing units of the parallel computing module at the same time according to the control of the controller module. The distribution of random data is realized by an improved Benes network, and the routing algorithm is used to realize arbitrary routing from N inputs to M outputs to meet the parallel random access requirements of the SpMV operation for the vector, wherein, .
6. The accelerator device for solving linear programming problems according to claim 1, characterized in that: The data receiving module is responsible for merging the calculation results of each basic operation unit in the parallel calculation module according to the control of the controller module and writing them back to the external memory.
7. A method for executing the accelerator device for solving linear programming problems according to any one of claims 1 to 6, characterized in that: The steps include: Step 1: Generate the code of the controller module for a specific problem; Step 2: Use the main processor to load the code into the controller module; Step 3: Use the host processor to start the controller module to execute the PDLP algorithm through the controller module register interface; Step 4: The controller module starts one iteration of the PDLP algorithm; Step 5: The controller module controls the data transport module to transport data; Step 6: The controller module controls the data distribution module to distribute data; Step 7: The controller module controls the parallel computing module to perform SpMV and vector calculations; Step 8: The controller module controls the data receiving module to recycle data; Step 9: The controller module determines whether the termination condition is met. If not, it returns to step 4; if it is met, it goes to step 10; Step 10: The controller module notifies the main processor that the calculation is completed; Step 11: The main processor reads the result and ends.