Processing device, method and equipment for computing acceleration in business application

By using SIMD architecture and first-in-first-out queue technology to encapsulate hardware instructions in the processing device, the problem of sparse matrix multiplication calculation bottleneck and low efficiency is solved, and efficient calculation acceleration of sparse matrix is ​​achieved.

CN120179291APending Publication Date: 2025-06-20SHANGHAI SMARTLOGIC TECHNOLOGY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256323.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When dealing with large-scale sparse matrix multiplication, the prior art faces the problems of computing bottlenecks and low computing efficiency, especially in data-intensive business scenarios.

Method used

A processing device and method are designed. Based on the SIMD architecture, by determining the adapted number of computing components, and encapsulating multiple time-consuming and repeated computing instructions into hardware instructions through the first-in-first-out queue technology, reducing the condition judgment and the number of instructions, thereby accelerating the computing of the sparse matrix.

Benefits of technology

Through this method, the calculation efficiency of sparse matrix multiplication is significantly improved, processing time is reduced, and it is suitable for data-intensive business application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179291A_ABST
    Figure CN120179291A_ABST
Patent Text Reader

Abstract

The invention discloses a processing device, method and equipment for computing acceleration in business application. The processing device for computing acceleration in the business application provided by the embodiment of the invention comprises an application scene and operand acquisition module used for acquiring a current business application scene and an operation matrix A and an operation matrix B participating in matrix operation in the business application scene; the hardware instruction generation module is used for determining operation instructions involved in matrix operation participated by the operation matrix A and the operation matrix B according to the service application scene, and packaging a plurality of operation instructions into hardware instructions; the component selection module is used for determining an adaptive number of operation components according to the data types of the operation matrix A and the operation matrix B; and the operation processing module is used for carrying out matrix operation processing on the operation matrix A and the operation matrix B according to the hardware instruction through each operation component. The device accelerates the operation of the sparse matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing devices, and in particular, to a processing device, method, and equipment for accelerating calculations in business applications. Background Art

[0002] With the acceleration of the social informatization process, artificial intelligence technology has been widely applied in business scenarios such as text classification, autonomous driving, and image recognition. The intelligent computing chips of artificial intelligence need to quickly process the acquired data through a hardware architecture and built-in software (such as deep learning algorithms, filtering operations), and then perform various operations.

[0003] However, there are a large number of sparse matrix (a matrix with most elements being zero) multiplication operations in the built-in software. For example, in deep learning algorithms, the weight matrix of a neural network and input data (such as a feature map processed by a ReLU activation function) often have sparsity.

[0004] To efficiently store and calculate sparse matrices, special storage formats are usually adopted, such as row compression format (CSR), column compression format (CSC), etc. Although it is possible to calculate sparse matrix multiplication relatively quickly in the CSR or CSC format compared to the dense case, when dealing with large-scale matrices, since existing sparse matrix multiplication calculation methods mostly run on a single instruction single data (SISD) architecture such as a traditional CPU, it will reach a calculation bottleneck when facing data-intensive operations, slowing down the overall calculation speed; at the same time, because the application scenarios faced by artificial intelligence are diverse, it is impossible to determine the number of non-zero elements in the sparse matrix, and the traditional processor instruction set only provides basic instructions such as addition, subtraction, multiplication, division, shift, data load / storage, and jump, resulting in sparse matrix multiplication operations often requiring a large number of instructions and conditional judgments, thereby reducing the operation efficiency of sparse matrix multiplication and the processing ability of artificial intelligence. Summary of the Invention

[0005] The present invention provides a processing device, method, and equipment for accelerating calculations in business applications to solve the problems that due to the inability to determine the number of non-zero elements in a sparse matrix stored in a special storage format, a large number of instructions and conditional judgments are required during the multiplication calculation process, resulting in a large amount of calculation, low efficiency, and long duration of sparse matrix multiplication.

[0006] In a first aspect, an embodiment of the present invention provides a processing device for accelerating calculations in business applications, and the device includes:

[0007] An application scenario and operand acquisition module, configured to acquire the current business application scenario and the operand matrices A and B participating in matrix operations in the business application scenario;

[0008] A hardware instruction generation module, configured to determine operation instructions involved in matrix operations of the operation matrix A and the operation matrix B according to the service application scenario, and encapsulate the multiple operation instructions into hardware instructions;

[0009] A component selection module, configured to determine an appropriate number of operation components according to the data types of the operation matrix A and the operation matrix B;

[0010] An operation processing module, configured to perform matrix operation processing on the operation matrix A and the operation matrix B according to the hardware instructions through each of the operation components.

[0011] In a second aspect, an embodiment of the present invention provides a processing method for computing acceleration in a service application. The method includes:

[0012] Obtain the current service application scenario and the operation matrix A and the operation matrix B participating in matrix operations in the service application scenario;

[0013] Determine operation instructions involved in matrix operations of the operation matrix A and the operation matrix B according to the service application scenario, and encapsulate the multiple operation instructions into hardware instructions;

[0014] Determine an appropriate number of operation components according to the data types of the operation matrix A and the operation matrix B;

[0015] Perform matrix operation processing on the operation matrix A and the operation matrix B according to the hardware instructions through each of the operation components.

[0016] In a third aspect, an embodiment of the present invention provides an electronic device. The electronic device includes the processing device for computing acceleration in a service application according to any embodiment of the present invention, and further includes:

[0017] A memory communicatively connected to the processing device;

[0018] Wherein, the memory stores a computer program executable by the processing device, and when the computer program is executed by the processing device, it can implement the processing method for computing acceleration in a service application according to any embodiment of the present invention.

[0019] The technical solution of the embodiment of the present invention. A processing device for computing acceleration in a service application provided by an embodiment of the present invention includes: an application scenario and operand acquisition module, configured to acquire the current service application scenario and the operation matrices A and B participating in matrix operations in the service application scenario; a hardware instruction generation module, configured to determine, according to the service application scenario, the operation instructions involved in the matrix operations of the operation matrices A and B, and encapsulate the multiple operation instructions into hardware instructions; a component selection module, configured to determine an appropriate number of operation components according to the data types of the operation matrices A and B; and an operation processing module, configured to perform matrix operation processing on the operation matrices A and B according to the hardware instructions through each of the operation components. Based on the design of the SIMD architecture, the device fully utilizes the instruction-level parallelism by determining an appropriate number of operation components, and aggregates and encapsulates multiple operation instructions with long execution times, many repetitions, and that can be hardware-encapsulated involved in matrix operations into hardware instructions through the first-in-first-out queue technology, reducing the explicit conditional judgments in the code and the number of instructions required for operations, thereby accelerating the operation of sparse matrices.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic structural diagram of a processing device for computing acceleration in a service application provided by an embodiment of the present invention;

[0023] Figure 2 It is a flowchart of a processing method for computing acceleration in a service application provided by an embodiment of the present invention;

[0024] Figure 3 It shows a schematic structural diagram of an electronic device that can be used to implement the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] It should be noted that when calculating the sparse matrix multiplication in business applications, a processor with a single instruction single data (SISD) architecture is usually adopted in the intelligent computing chip of artificial intelligence, and the processor instruction set only provides basic instructions such as addition, subtraction, multiplication, division, shift, data loading / storage, and jump. Therefore, when facing data-intensive services, the processor will reach a computing bottleneck, slowing down the overall computing speed. At the same time, due to the need for a large number of instructions and conditional judgments, the operation efficiency of the sparse matrix multiplication and the processing ability of artificial intelligence are further reduced.

[0028] Based on this, the embodiments of the present invention provide a processing device for accelerating calculations in business applications. Figure 1 It is a schematic structural diagram of a processing device for accelerating calculations in business applications provided by the embodiments of the present invention. The embodiments of the present invention are applicable to scenarios involving sparse matrix multiplication operations in business applications. This device can execute a processing method for accelerating calculations in business applications. This device can be implemented in the form of software and / or hardware. Optionally, it is implemented through an electronic device, and the electronic device is preferably a mobile terminal, a desktop computer, a laptop computer, a server, etc.

[0029] Specifically, the processing device for computing acceleration in business applications can be a processor under the Single Instruction Multiple Data (SIMD) architecture. Compared with the traditional Single Instruction Single Data (SISD) architecture, that is, the architecture where one instruction processes a single data, the SIMD architecture uses one instruction to process multiple data, achieving instruction-level parallelism, thereby accelerating operations such as vector operations. For example, a traditional SISD 64-bit wide processor can only calculate one double-precision floating-point number multiplied by one double-precision floating-point number with a single instruction, while a SIMD architecture processor with a bit width of n * 64 bits can calculate n double-precision floating-point numbers multiplied by n double-precision floating-point numbers with a single instruction, and also supports parallel calculation of 2n single-precision floating-point numbers multiplied by 2n single-precision floating-point numbers. When there are m cores and k groups of required basic arithmetic components in the processor core, the parallelism can be extended to k * n * m times the basic parallelism. Therefore, the SIMD architecture is very suitable for compute-intensive services, and the large bit width has a certain degree of flexibility, enabling parallel acceleration throughout the entire process of data loading, computing, and storage. Moreover, the parallelism is only related to the bit width design, the number of cores, and the number of arithmetic components, without the need for modification at the coding level. Additionally, through reasonable data layout and access patterns, the SIMD architecture can optimize memory access and reduce data transfer bottlenecks.

[0030] As Figure 1 shown, the processing device for computing acceleration in business applications provided by the embodiments of the present invention may specifically include: an application scenario and operand acquisition module 10, a hardware instruction generation module 20, a component selection module 30, and an arithmetic processing module 40. Among them,

[0031] The application scenario and operand acquisition module 10 is configured to acquire the current business application scenario and the operand matrices A and B participating in matrix operations in the business application scenario;

[0032] The hardware instruction generation module 20 is configured to determine the arithmetic instructions involved in the matrix operations of the operand matrices A and B according to the business application scenario, and encapsulate the multiple arithmetic instructions into hardware instructions;

[0033] The component selection module 30 is configured to determine an appropriate number of arithmetic components according to the data types of the operand matrices A and B;

[0034] The arithmetic processing module 40 is configured to perform matrix operation processing on the operand matrices A and B according to the hardware instructions through the respective arithmetic components.

[0035] It should be noted that at least one of the operation matrices A and B is a sparse matrix, that is, a matrix in which most of the elements are zero. Compared with a dense matrix (where most elements are non-zero), a sparse matrix is more efficient in storage and calculation. At the same time, the number of columns of the operation matrix A should be the same as the number of rows of the operation matrix B.

[0036] In this embodiment, the application scenario and operand acquisition module 10 can acquire the current business application scenario, such as fluid dynamics simulation, image reconstruction, commodity recommendation, etc., and acquire at least one pair of operation matrices participating in matrix operations (such as matrix multiplication, matrix addition, etc.) in this business application scenario. In this embodiment, a pair of operation matrices is taken as an example, that is, the operation matrices A and B are acquired. The hardware instruction generation module 20 determines the operation instructions involved in the matrix operations of the operation matrices A and B according to the acquired business application scenario. Exemplarily, when the matrix operation is matrix multiplication, the operation instructions involved may include the selection of elements in the matrix, the determination of the operation relationship between the elements in the two matrices, and the determination of the storage address of the calculation result, etc. Then, the operation instructions can be divided based on the functionality of the operation instructions, the business of a specific service, etc. For example, the operation instructions can be divided according to whether the operation instructions participate in element operations or address determination, and multiple operation instructions that are time-consuming, have a large number of repetitions, and can be hardware-packaged and belong to the same function are encapsulated into one hardware instruction through the first-in-first-out queue technology to reduce explicit conditional judgments in the code (for example, when writing data, the data writing and pointer update operations can be combined into one instruction), so as to complete operations such as corresponding selection of elements with operation requirements in the matrix or determination of the storage address of the calculation result through one hardware instruction, and improve the calculation efficiency of sparse matrix operations.

[0037] Continuing with the above description, the component selection module 30 can determine an appropriate number of arithmetic components according to the data types of the operation matrices A and B (such as 32-bit floating-point data type, 64-bit floating-point data type). Exemplarily, considering the bit width of the SIMD architecture, if both operation matrices A and B are of 64-bit floating-point data type, then 2 sets of data readers / writers, 2 sets of data selectors, and 1 set of data arithmetic units are required. The data readers / writers, data selectors, and data arithmetic units are all designed with the same long bit width, such as n * 64-bit design, so that one instruction can calculate n elements simultaneously. It should be noted that only by increasing the number of the above arithmetic components in multiples can the simultaneous calculation of multiple sparse matrix multiplications be achieved. Exemplarily, using 4 times the above arithmetic components, that is, 8 sets of data readers / writers, 8 sets of data selectors, and 4 sets of data arithmetic units can perform the matrix operations of 4 sets of operation matrices simultaneously. The operation processing module 40 performs matrix operation processing on the operation matrices A and B in parallel through the arithmetic components determined by the component selection module 30 according to the hardware instructions encapsulated by the hardware instruction generation module 20. It should be noted that each arithmetic component only needs to repeat its corresponding operation and does not have to wait for other arithmetic components.

[0038] A processing device for computing acceleration in business applications provided by an embodiment of the present invention includes: an application scenario and operand acquisition module, configured to acquire the current business application scenario and the operation matrices A and B participating in matrix operations in the business application scenario; a hardware instruction generation module, configured to determine the operation instructions involved in the matrix operations of the operation matrices A and B according to the business application scenario, and encapsulate multiple said operation instructions into hardware instructions; a component selection module, configured to determine an appropriate number of arithmetic components according to the data types of the operation matrices A and B; and an operation processing module, configured to perform matrix operation processing on the operation matrices A and B according to the hardware instructions through each said arithmetic component. Based on the SIMD architecture design, by determining an appropriate number of arithmetic components, the instruction-level parallelism ability is fully utilized, and multiple time-consuming, frequently repeated, and hardware-encapsulable operation instructions involved in matrix operations are aggregated and encapsulated into hardware instructions through the first-in-first-out queue technology, reducing the explicit conditional judgments in the code and the number of instructions required for operations, thereby accelerating the operation of sparse matrices.

[0039] As a first alternative embodiment of the embodiment of the present invention, on the basis of the above embodiment, the processing device further includes:

[0040] A data processing module 50, configured to convert the operation matrices A and B into sparse matrix forms respectively, and obtain the first index information of the operation matrix A and the second index information of the operation matrix B;

[0041] A storage module 60, configured to read in the first index information and the second index information and store them in a storage component.

[0042] Among them, the first index information and the second index information may respectively include row pointers of rows in the operation matrices A and B where not all elements are zero. The row pointers include the starting address of the column index of the current row (i.e., which non-zero element in the whole matrix is the first non-zero element of the current row), the number of valid elements in the current row, the row number of the current row, and a matrix end flag (used to indicate whether the current row is the last row of the matrix, which can be distinguished by 0 and 1). The first index information may further include an array of column indices (coordinates) of elements in each row in the row, compressed data (i.e., valid elements in the matrix), and row pointer information (including all information of a row pointer and a flag bit indicating whether the row has been processed).

[0043] In this embodiment, the data processing module 50 respectively converts the operation matrices A and B into a certain sparse matrix form, such as converting them into a sparse matrix storage format of CSR (row compression format) or CSC (column compression format), to obtain the first index information of the operation matrix A and the second index information of the operation matrix B. The storage module 60 can read in the first index information and the second index information obtained by the data processing module 50 and store them in the storage component respectively.

[0044] In the above technical solution of this embodiment, the data processing module 50 obtains the first index information of the operation matrix A and the second index information of the operation matrix B, and the storage module 60 stores the first index information and the second index information in the storage component, so as to achieve more efficient storage of the matrix and provide strong data support for subsequent matrix operations, thereby improving the operation efficiency.

[0045] As the second optional embodiment of the present invention, the operation processing module 40 may further specifically include:

[0046] A task division unit 41, configured to split the matrix multiplication operation of the operation matrix A and the operation matrix B into multiple operation subtasks, and each subtask is a multiplication operation of one row in the operation matrix A with the operation matrix B;

[0047] A subtask operation unit 42, configured to, through each of the operation components, respectively execute an operation processing logic for each of the operation subtasks according to the valid data selection instruction in the hardware instruction, and obtain the subtask operation results of each of the operation subtasks;

[0048] A result determination unit 43, configured to generate a result matrix of the multiplication operation of the operation matrix A and the operation matrix B according to each of the subtask operation results.

[0049] Among them, the valid data selection instruction can be understood as follows: when performing the multiplication operation of the operation matrix A and the operation matrix B, when it is determined that the column index corresponding to a single non-zero element in a certain row of the operation matrix A is a, quickly find the row with the row number a in the operation matrix B, send the row pointer information corresponding to the row in the found operation matrix B to the data reader / writer, and send the non-zero element in the operation matrix A to the data arithmetic unit.

[0050] In this embodiment, the operation of multiplying the operation matrix A by the operation matrix B is split into multiple operation subtasks by the task division unit 41. Each operation subtask is the multiplication of one row of the operation matrix A by the operation matrix B. That is, the task division unit 41 can first calculate the multiplication of the first row of the operation matrix A by the operation matrix B, then calculate the multiplication of the second row of the operation matrix A by the operation matrix B, until the multiplication of the last row of the operation matrix A by B is completed, and the entire operation of multiplying the operation matrix A by the operation matrix B is completed. The subtask operation unit 42 then executes the operation processing logic for each of the operation subtasks respectively according to the valid data selection instruction in the hardware instruction through each operation component, and obtains the subtask operation results of each of the operation subtasks. The result determination unit 43 generates the result matrix of the multiplication operation of the operation matrix A and the operation matrix B according to the subtask operation results of each subtask.

[0051] It should be noted that before executing the operation processing logic for each of the operation subtasks respectively according to the valid data selection instruction in the hardware instruction to obtain the subtask operation results of each of the operation subtasks, it is necessary to pre-read all the row pointers of the operation matrix B from the storage module 60 into the cache through the data reader / writer in advance, so as to avoid repeatedly reading the row pointers of the operation matrix B from the outside when using the valid data selection instruction; at the same time, it is necessary to appropriately add the operation of reading the first index information of the next row of the operation matrix A from the storage module 60 into the cache through the data reader / writer, so as to avoid the situation where one row of the operation matrix A has been calculated and the next row has not been loaded yet.

[0052] In the above technical solution of this embodiment, by abstracting the fixed operation sequence in the sparse matrix multiplication using the valid data selection instruction in the hardware instruction, and accelerating the sparse matrix multiplication based on the function of the instruction, the number of instructions required for calculating the sparse matrix multiplication is reduced in principle. At the same time, all the row pointers of the operation matrix B and the first index information of at least one row of the operation matrix A are pre-read into the cache in advance, reducing the number of memory accesses and improving the operation efficiency.

[0053] As one implementation manner, the subtask operation unit 42 may further specifically include:

[0054] The first operator unit 421 is configured to respond to the valid data selection instruction in the hardware instruction, determine the current operation row involved in the current operator task, and determine the operation element column index array corresponding to the current operation row. The valid operation elements in the current operation row and the operation element column index array are respectively read into the data rearrangement unit through the data reader / writer. The current operation row corresponds to a matrix row of the operation matrix A, and the operation element column index array is obtained from the first index information corresponding to the operation matrix A, and the first index information is obtained by accessing the storage component.

[0055] The second operator unit 422 is configured to sequentially select valid operation elements from the current operation row according to the operation element column index array in the data rearrangement unit, perform a multiplication operation with the operation matrix B, and obtain an element operation result relative to the valid operation element until the selection of the valid operation elements in the current operation row is completed, and form a subtask operation result based on each of the element operation results.

[0056] Among them, the operation element column index array can be understood as the column index array corresponding to each valid operation element.

[0057] In this embodiment, the first operator unit 421 responds to the valid data selection instruction in the hardware instruction, determines the current operation row of the operation matrix A involved in the current operator task according to the row pointer of the operation matrix A read into the cache in advance from the first index information of the storage module 60 through the data reader / writer, and determines the operation element column index array corresponding to the valid operation elements in the current operation row according to the first index information. The operation element column index array and the valid operation elements are respectively read into the data rearrangement unit through the data reader / writer. Then, the second operator unit 422 sequentially selects valid operation elements from the current operation row according to the operation element column index array stored in the data rearrangement unit, performs a multiplication operation with the operation matrix B, and obtains an element operation result relative to each of the valid operation elements until the multiplication operation of all the valid operation elements in the current operation row with the operation matrix B is completed, and forms a subtask operation result based on each of the element operation results. Exemplarily, the operation can be performed with the operation matrix B in ascending order of the operation element column index array of the valid operation elements in the current operation row stored in the data rearrangement unit. For example, first calculate the multiplication of the first valid operation element in the current operation row by the operation matrix B, then calculate the multiplication of the second valid operation element in the current operation row by the operation matrix B, and so on until the multiplication of the last valid operation element in the current operation row by the operation matrix B is completed to finish this operator task, and form a subtask operation result based on each of the element operation results, that is, each of the element operation results is commonly regarded as the subtask operation result.

[0058] In the above technical solution of this embodiment, by responding to the valid data selection instruction in the hardware instruction, valid operation elements are sequentially selected from each row of the operation matrix A according to the operation element column index array to perform a multiplication operation with the operation matrix B, and an element operation result relative to the valid operation element is obtained, reducing the calculation of non-zero elements in the matrix and the number of required repeated instructions, and without a large number of conditional judgments, improving the calculation efficiency of sparse matrix multiplication.

[0059] As an alternative embodiment, the second operation sub-unit 422 may specifically perform the following steps:

[0060] a1) Using the valid data selection instruction, determine the operand element row pointer information corresponding to the current operation row from the second index information corresponding to the operation matrix B, and transmit the operand element row pointer information to the data reader / writer. At the same time, transmit the valid operation elements corresponding to the operation element column index array from the data rearranger to the data arithmetic unit, and transmit the operation element row pointer information corresponding to the valid operation elements to the data rearranger. The operation element row pointer information is determined from the first index information corresponding to the operation matrix A, and the second index information is obtained by accessing the storage component.

[0061] Among them, the row pointer information can be understood as all the information including a row pointer and the information of a flag bit recording whether the row has been processed. The operation element row pointer information can be considered as the row pointer information of the current operation row corresponding to the valid operation elements in the operation matrix A, and the operand element row pointer information can be understood as the row pointer information of the current operand row corresponding to the element data in the operation matrix B.

[0062] In this embodiment, the specific manner of determining the operand element row pointer information corresponding to the current operation row from the second index information corresponding to the operation matrix B by using the valid data selection instruction may be as follows: According to the operation element column index array of a certain valid operation element in the current operation row of the operation matrix A, use a data selector to determine the row whose row number is the same as the operation element column index array from the row pointers of the operation matrix B read into the cache in advance from the storage module 60, and determine the operand element row pointer information of this row from the second index information corresponding to the operation matrix B. Then transmit the operand element row pointer information to the data reader / writer. At the same time, transmit the valid operation elements corresponding to the operation element column index array from the data rearranger to the data arithmetic unit, and transmit the operation element row pointer information in the first index information corresponding to the valid operation elements to the data rearranger.

[0063] b1) Through the data reader / writer, use the row pointer information of the operand elements to determine the current operand row of the operation matrix B, and load the element data of the current operand row and the array of column indices of the operand elements corresponding to the element data into the data arithmetic unit and the data re-arranger respectively.

[0064] Among them, the current operand row can be understood as the row that multiplies with a certain valid operation element in the operation matrix A, and the row number of this row is the same as the array of column indices of the operation elements of the valid operation element. The array of column indices of the operand elements can be understood as the array of column indices of the element data in the operation matrix B.

[0065] In this embodiment, use the row pointer information of the operand elements to determine the current operand row in the operation matrix B, and through the data reader / writer, load the element data of the current operand row from the second index information into the data arithmetic unit, and load the array of column indices of the operand elements corresponding to the element data into the data re-arranger.

[0066] c1) Through the data arithmetic unit, perform multiplication operations on the valid operation element and each of the element data corresponding to the valid operation element respectively, and send each multiplication result as the element operation result of the valid operation element to the data reader / writer.

[0067] In this embodiment, through the data arithmetic unit, multiply the valid operation element and each of the element data corresponding to the valid operation element (that is, each element data in the current operand row whose row number is the same as the array of column indices of the operation elements of the valid operation element) respectively, and send each multiplication result as the element operation result of the valid operation element to the data reader / writer through the data selector.

[0068] d1) Through the data selector and the data re-arranger, according to the result address generation instruction in the hardware instruction, determine the result address information of each element operation result, and send each result address information to the data reader / writer.

[0069] Among them, the result address generation instruction can be understood as an instruction used to calculate the position where the multiplication result corresponding to the valid operation element and the element data should be written back to the storage device and send it to the data reader / writer when the row number of the valid operation element in the known operation matrix A and the array of column indices of the operand elements of the element data in the operation matrix B are known.

[0070] In this embodiment, the row pointer information of the operation matrix A corresponding to the operation results of each element and the array of column indexes of the elements to be operated on the operation matrix B are determined from the data rearranger through a data selector. The result address information of each of the element operation results is determined according to the result address generation instruction in the hardware instruction, and each of the result address information is sent to the data reader / writer through the data selector. It should be noted that c1) and d1) can be executed in parallel.

[0071] With the above technical solution of this embodiment, the sparse matrix multiplication can be efficiently executed at the hardware level by using the above four operation components in cooperation, which is especially suitable for scenarios requiring high throughput (such as AI training, image processing). At the same time, combined with the SIMD architecture and the encapsulated hardware instructions, the vectorized operation is realized.

[0072] Further, the execution step of determining the result address information of each of the element operation results according to the result address generation instruction by the data selector and the data rearranger can be specifically implemented as follows:

[0073] a2) In response to the result address generation instruction, determine the matrix row-column information of the product matrix according to the row-column information of the operation matrix A and the operation matrix B.

[0074] In this embodiment, in response to the result address generation instruction, the matrix row-column information of the product matrix is determined according to the row-column information of the operation matrix A and the operation matrix B. Exemplarily, when both the operation matrix A and the operation matrix B are 3*3 matrices, it can be determined that there are 3 elements in one row of the operation matrix and the product matrix is 3 rows and 3 columns, that is, an empty 3*3 matrix can be constructed to store the operation results of each element.

[0075] b2) For each element operation result, determine the row number of the current operation row where the valid operation element corresponding to the element operation result is located, and the array of column indexes of the elements to be operated on the element data corresponding thereto. The row number is obtained by reading the operation element row pointer information.

[0076] In this embodiment, for each element operation result, determine the row number of the current operation row where the valid operation data corresponding to the element operation result is located, and determine the array of column indexes of the elements to be operated on the element data corresponding to the element operation result.

[0077] c2) According to the row number and the array of column indexes of the elements to be operated on, and in combination with the matrix row-column information, determine the position offset value of the element operation result, and use the position offset value as the result address information of the element operation result.

[0078] In this embodiment, the specific method for determining the position offset value of the element operation result according to the row number and the array of column indices of the elements to be operated, in combination with the matrix row-column information, may be to multiply the row number by the number of columns of the operation matrix A and then add the array of column indices of the elements to be operated. Exemplarily, continuing with the above example description, the row number is 1 (starting from 0, that is, the operation matrix A has the 0th row, the 1st row, and the 2nd row), the number of columns of the operation matrix A (3*3) is 3, and the array of column indices of the elements to be operated is 2 (starting from 0, that is, the operation matrix B has the 0th column, the 1st column, and the 2nd column). Then, the position offset value of the element operation result is 5 (1*3 + 2), that is, the result address information of the element operation result is 5.

[0079] In the above technical solution of this embodiment, in response to the result address generation instruction, by determining the result address information corresponding to each element operation result according to each row number and each array of column indices of the elements to be operated, it provides strong support for subsequently generating the result matrix of the multiplication operation of the operation matrix A and the operation matrix B.

[0080] As one implementation manner, the result determination unit 43 may further specifically execute the following steps:

[0081] a3) For each operation subtask, read the element operation results that make up the subtask operation result corresponding to the operation subtask from the data reader / writer, and obtain the result address information of each of the element operation results from the data reader / writer.

[0082] b3) Write the corresponding element operation results back to the storage device according to each of the result address information, and form the result matrix of the multiplication operation of the operation matrix A and the operation matrix B according to the data content formed in the storage device after the write-back.

[0083] In this embodiment, obtain each element operation result and the result address information of each element operation result from the data reader / writer, and write the corresponding element operation results back to the storage device according to the result address information, thereby forming the result matrix of the multiplication operation of the operation matrix A and the operation matrix B. It should be noted that the three rows of the empty matrix (product matrix) are the 0th row to the 2nd row, and the three columns are the 0th column to the 2nd column, and the result address information ranges from 0 to 8. Exemplarily, continuing with the above example description, if the result address information corresponding to the element operation result of 1 is 5, then store 1 in the 1st row and 2nd column of the 3*3 empty matrix, that is, X,X,X (the 0th row), X,X,1 (the 1st row).

[0084] As an alternative embodiment, in the write-back of the element operation result to the storage device, if there is already data content at the address to be written, then accumulate the element operation result and the existing data content to form new data content.

[0085] Among them, the address to be written can be understood as the position in the storage device corresponding to the result of the element operation.

[0086] In this embodiment, when writing the result of the element operation back to the corresponding position of the storage device, if there is already data content at the address to be written, the result of the element operation is accumulated with the existing data content to form new data content and stored at the address to be written.

[0087] Figure 2 The figure is a flowchart of a processing method for computing acceleration in a service application provided by an embodiment of the present invention. This method can be executed by a processing device for computing acceleration in a service application. As Figure 2 shown, the image retrieval method provided by the embodiment of the present invention specifically may include:

[0088] S101. Obtain the current service application scenario and operation matrices A and B participating in matrix operations in the service application scenario;

[0089] S102. Determine the operation instructions involved in the matrix operations of operation matrices A and B according to the service application scenario, and encapsulate multiple operation instructions into hardware instructions;

[0090] S103. Determine the appropriate number of operation components according to the data types of operation matrix A and operation matrix B;

[0091] S104. Perform matrix operation processing on operation matrices A and B according to the hardware instructions through each operation component.

[0092] A processing method for computing acceleration in a service application provided by an embodiment of the present invention includes: obtaining the current service application scenario and operation matrices A and B participating in matrix operations in the service application scenario; determining the operation instructions involved in the matrix operations of operation matrices A and B according to the service application scenario, and encapsulating multiple operation instructions into hardware instructions; determining the appropriate number of operation components according to the data types of operation matrix A and operation matrix B; performing matrix operation processing on operation matrices A and B according to the hardware instructions through each operation component. Based on the design of the SIMD architecture, this device determines the appropriate number of operation components, makes full use of the instruction-level parallelism ability, and aggregates and encapsulates multiple time-consuming, frequently repeated, and hardware-encapsulable operation instructions involved in matrix operations into hardware instructions through the first-in-first-out queue technology, reducing the explicit conditional judgments in the code and the number of instructions required for operations, thereby accelerating the operation of sparse matrices.

[0093] Further, the method further includes:

[0094] Convert the operation matrix A and the operation matrix B into sparse matrix forms respectively to obtain the first index information of the operation matrix A and the second index information of the operation matrix B;

[0095] Read in the first index information and the second index information and store them in the storage component.

[0096] Further, the matrix operation processing of the operation matrix A and the operation matrix B by each of the operation components according to the hardware instruction may specifically include:

[0097] Split the matrix multiplication operation of the operation matrix A and the operation matrix B into multiple operation subtasks, and each subtask is to perform a multiplication operation on one row of the operation matrix A and the operation matrix B;

[0098] Through each of the operation components, perform operation processing logic on each of the operation subtasks respectively according to the valid data selection instruction in the hardware instruction to obtain the subtask operation results of each of the operation subtasks;

[0099] Generate a result matrix for the multiplication operation of the operation matrix A and the operation matrix B according to the subtask operation results of each of the subtasks.

[0100] Further, the performing operation processing logic on each of the operation subtasks respectively according to the valid data selection instruction in the hardware instruction by each of the operation components to obtain the subtask operation results of each of the operation subtasks may specifically include:

[0101] Respond to the valid data selection instruction in the hardware instruction, determine the current operation row involved in the current operation subtask, and determine the operation element column index array corresponding to the current operation row. Read the valid operation elements in the current operation row and the operation element column index array into the data rearrangement unit through a data reader and writer respectively. The current operation row corresponds to a matrix row of the operation matrix A, and the operation element column index array is obtained from the first index information corresponding to the operation matrix A, and the first index information is obtained by accessing the storage component;

[0102] According to the operation element column index array in the data rearrangement unit, sequentially select valid operation elements from the current operation row to perform a product operation with the operation matrix B to obtain the element operation results for the valid operation elements until the selection of the valid operation elements in the current operation row is completed, and form a subtask operation result based on the element operation results of each of the elements.

[0103] Further, according to the operation element column index array in the data rearranger, valid operation elements are sequentially selected from the current operation row to perform a multiplication operation with the operation matrix B to obtain an element operation result relative to the valid operation element. Until the selection of valid operation elements in the current operation row is completed, a sub-task operation result is formed based on each element operation result. Specifically, it may include:

[0104] Using the valid data selection instruction, determine the operand element row pointer information relative to the current operation row from the second index information corresponding to the operation matrix B, and transmit the operand element row pointer information to the data reader / writer. At the same time, send the valid operation elements corresponding to the operation element column index array from the data rearranger to the data arithmetic unit, and send the operation element row pointer information corresponding to the valid operation elements to the data rearranger. The operation element row pointer information is determined from the first index information corresponding to the operation matrix A, and the second index information is obtained by accessing the storage component;

[0105] Through the data reader / writer, use the operand element row pointer information to determine the current operand row of the operation matrix B, and load the element data of the current operand row and the operand element column index array corresponding to the element data into the data arithmetic unit and the data rearranger respectively;

[0106] Through the data arithmetic unit, perform a multiplication operation on the valid operation elements and the respective element data corresponding to the valid operation elements, and send each multiplication operation result as the element operation result of the valid operation element to the data reader / writer;

[0107] Through the data selector and the data rearranger, according to the result address generation instruction in the hardware instruction, determine the result address information of each element operation result, and send each result address information to the data reader / writer.

[0108] Further, the execution steps of determining the result address information of each element operation result through the data selector and the data rearranger according to the result address generation instruction in the hardware instruction include:

[0109] Respond to the result address generation instruction, and determine the matrix row-column information of the product matrix according to the row-column information of the operation matrix A and the operation matrix B;

[0110] For each element operation result, determine the row number of the current operation row where the valid operation element corresponding to the element operation result is located, and the operand element column index array corresponding to the element data. The row number is obtained by reading the operation element row pointer information;

[0111] Determine the position offset value of the element operation result according to the line number and the array of column indexes of the elements to be operated, and combine the matrix row and column information, and use the position offset value as the result address information of the element operation result.

[0112] Further, generating the result matrix of the multiplication operation of the operation matrix A and the operation matrix B according to the operation results of each sub-task specifically includes:

[0113] For each operation sub-task, read the element operation results that make up the sub-task operation result corresponding to the operation sub-task from the data reader, and obtain the result address information of each element operation result from the data reader;

[0114] Write the corresponding element operation results back to the storage device according to each result address information, and form the result matrix of the multiplication operation of the operation matrix A and the operation matrix B according to the data content formed in the storage device after writing back.

[0115] Further, in the write-back of the element operation result to the storage device, if there is already data content at the write address to be written, add the element operation result to the existing data content to form new data content.

[0116] The processing method for computing acceleration in business applications provided by the embodiments of the present invention can be executed by the processing device for computing acceleration in business applications provided by any embodiment of the present invention, and has the corresponding execution methods and beneficial effects of the functional modules.

[0117] Figure 3 FIG. shows a schematic structural diagram of an electronic device 30 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0118] Such as Figure 3As shown, the electronic device 30 includes at least one processing device 31 and a memory communicatively connected to the at least one processing device 31, such as a read-only memory (ROM) 32, a random access memory (RAM) 33, etc. The memory stores a computer program executable by the at least one processing device. The processing device 31 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 32 or the computer program loaded from the storage unit 38 into the random access memory (RAM) 33. In the RAM 33, various programs and data required for the operation of the electronic device 30 can also be stored. The processing device 31, the ROM 32, and the RAM 33 are connected to each other via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.

[0119] Multiple components in the electronic device 30 are connected to the I / O interface 35, including: an input unit 36, such as a keyboard, a mouse, etc.; an output unit 37, such as various types of displays, speakers, etc.; a storage unit 38, such as a disk, an optical disc, etc.; and a communication unit 39, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 39 allows the electronic device 30 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0120] The processing device 31 can be various general and / or special processing components with processing and computing capabilities. Some examples of the processing device 31 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processing device 31 executes the various methods and processes described above, such as the processing method for computing acceleration in business applications.

[0121] In some embodiments, the processing method for computing acceleration in business applications can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 38. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 30 via the ROM 32 and / or the communication unit 39. When the computer program is loaded into the RAM 33 and executed by the processing device 31, one or more steps of the processing method for computing acceleration in business applications described above can be executed. Alternatively, in other embodiments, the processing device 31 can be configured to execute the processing method for computing acceleration in business applications by any other appropriate means (e.g., by means of firmware).

[0122] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processing device, which can be a special-purpose or general-purpose programmable processing device that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0123] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processing devices of a general purpose computer, a special purpose computer, or other programmable data processing devices, such that the computer programs, when executed by the processing devices, cause the functions / operations specified in the flowchart(s) and / or block diagram(s) to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0124] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0125] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0126] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0127] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0128] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0129] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A processing device for computing acceleration in business applications, characterized in that: include: An application scenario and operand acquisition module, used to acquire the current business application scenario and the operation matrix A and operation matrix B involved in the matrix operation in the business application scenario; A hardware instruction generation module, used to determine the operation instructions involved in the matrix operation of the operation matrix A and the operation matrix B according to the business application scenario, and encapsulate a plurality of the operation instructions into hardware instructions; A component selection module, used to determine an adapted number of operation components according to the data types of the operation matrix A and the operation matrix B; The operation processing module is used to perform matrix operation processing on the operation matrix A and the operation matrix B according to the hardware instructions through each of the operation components.

2. The processing device according to claim 1, characterized in that Also includes: A data processing module, used to convert the operation matrix A and the operation matrix B into sparse matrix forms respectively, to obtain first index information of the operation matrix A and second index information of the operation matrix B; The storage module is used to read the first index information and the second index information and store them in a storage component.

3. The processing device according to claim 1, characterized in that The operation processing module includes: A task division unit, used for dividing the matrix multiplication operation of the operation matrix A and the operation matrix B into a plurality of operation subtasks, each subtask is a multiplication operation of a row in the operation matrix A with the operation matrix B; A subtask operation unit, configured to execute operation processing logic for each of the operation subtasks respectively according to valid data selection instructions in the hardware instructions through each of the operation components, so as to obtain a subtask operation result for each of the operation subtasks; The result determination unit is used to generate a result matrix of multiplication operation of the operation matrix A and the operation matrix B according to the operation results of each of the subtasks.

4. The processing device according to claim 3, characterized in that The subtask computing unit specifically includes: A first operator unit, configured to respond to a valid data selection instruction in the hardware instruction, determine a current operation row involved in a current operation subtask, and determine an operation element column index array corresponding to the current operation row, and read the valid operation elements in the current operation row and the operation element column index array into a data rearranger through a data reader / writer, respectively, wherein the current operation row corresponds to a matrix row of the operation matrix A, and the operation element column index array is obtained from first index information corresponding to the operation matrix A, and the first index information is obtained by accessing a storage component; The second operator unit is used to select valid operation elements from the current operation row in turn according to the operation element column index array in the data rearranger, and perform multiplication operation with the operation matrix B to obtain element operation results relative to the valid operation elements until the selection of valid operation elements in the current operation row is completed, and construct subtask operation results based on each element operation result.

5. The processing device according to claim 4, characterized in that The second operator unit is specifically used for: Using the valid data selection instruction, determine the row pointer information of the operated element relative to the current operation row from the second index information corresponding to the operation matrix B, and transmit the row pointer information of the operated element to the data reader / writer, and at the same time, send the valid operation element corresponding to the operation element column index array from the data rearranger to the data operator, and send the operation element row pointer information corresponding to the valid operation element to the data rearranger, the operation element row pointer information is determined from the first index information corresponding to the operation matrix A, and the second index information is obtained by accessing the storage component; The data reader / writer uses the operated element row pointer information to determine the current operated row of the operation matrix B, and loads the element data of the current operated row and the operated element column index array corresponding to the element data into the data operator and the data rearranger respectively; The data operator performs multiplication operations on the effective operation elements and the element data corresponding to the effective operation elements respectively, and sends the multiplication results as the element operation results of the effective operation elements to the data reader / writer; The result address information of each element operation result is determined through the data selector and the data rearranger according to the result address generation instruction in the hardware instruction, and each result address information is sent to the data reader / writer.

6. The processing device according to claim 5, wherein the step of determining the result address information of each element operation result by the data selector and the data rearranger according to the result address generation instruction in the hardware instruction comprises: In response to the result address generation instruction, determining the matrix row and column information of the product matrix according to the row and column information of the operation matrix A and the operation matrix B; For each element operation result, determine the row number of the current operation row where the valid operation element corresponding to the element operation result is located, and the operated element column index array of the corresponding element data, wherein the row number is obtained by reading the operation element row pointer information; According to the row number and the index array of the element column being operated, combined with the row and column information of the matrix, the position offset value of the element operation result is determined, and the position offset value is used as the result address information of the element operation result.

7. The processing device according to claim 3, characterized in that The result determination unit is specifically used for: For each operation subtask, read the element operation result constituting the subtask operation result corresponding to the operation subtask from the data reader / writer, and obtain the result address information of each element operation result from the data reader / writer; According to each of the result address information, the corresponding element operation results are written back to the storage device, and the result matrix of the multiplication operation of the operation matrix A and the operation matrix B is constructed according to the data content formed in the storage device after writing back.

8. The processing device according to claim 7, characterized in that The element operation result is written back to the storage device. If data content already exists at the address to be written, the element operation result is accumulated with the existing data content to form new data content.

9. A processing method for computing acceleration in business applications, characterized in that: include: Obtain the current business application scenario and the operation matrix A and the operation matrix B involved in the matrix operation in the business application scenario; According to the business application scenario, determining the operation instructions involved in the matrix operation of the operation matrix A and the operation matrix B, and encapsulating a plurality of the operation instructions into hardware instructions; Determine an adapted number of computing components according to the data types of the operation matrix A and the operation matrix B; Through each of the computing components, matrix computing is performed on the operation matrix A and the operation matrix B according to the hardware instructions.

10. An electronic device, characterized in that: The processing device for accelerating computing in business applications according to any one of claims 1 to 8 further comprises: a memory in communication with the processing device; The memory stores a computer program that can be executed by the processing device, and the computer program is executed by the processing device to implement the processing method for computing acceleration in business applications described in claim 9.