Processing method and device for triangular matrix inverse operation

By dividing the triangular matrix inverse operation tasks into multiple computing cores on the NPU platform, and using vector instructions and matrix multiplication processing, the problem of inefficient triangular matrix inverse operation on the NPU platform is solved, achieving a more efficient computing process and greater task throughput.

CN119149895BActive Publication Date: 2025-05-06ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411645466.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-05-06
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

When performing triangular matrix inverse operations on the NPU of the artificial intelligence computing platform, the calculation process is less efficient due to the involvement of a large number of scalar operations and a high degree of operation dependence.

Method used

By evenly dividing the target tasks into multiple computing cores of the target processing platform, and gradually solving the inverse matrix of the triangle matrix through vector instructions and matrix multiplication processing.

Benefits of technology

It improves the processing efficiency of inverse operation of triangle matrix, increases task throughput, and solves the problems of slow computing speed and low efficiency in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119149895B_ABST
    Figure CN119149895B_ABST
Patent Text Reader

Abstract

The present invention discloses a processing method and device for triangular matrix inverse operation. The method includes: evenly dividing the target task into multiple computing cores of the target processing platform; obtaining the triangular matrices corresponding to the multiple blocks, and inverting the triangular matrix to obtain the inverse matrix of the triangular matrix; storing the diagonal elements and the unit matrix of the inverse matrix of the triangular matrix into the first storage medium, respectively, and storing the known part of the inverse matrix and the coefficient matrix into the second storage medium; retrieving data from the first storage medium to perform parallel calculations on the data retrieved from the first storage medium, and retrieving data from the second storage medium to perform cube volume calculations on the data in the second storage medium, and obtaining the inverse operation result of the inverse matrix. The present invention solves the technical problem in the related art that the execution of triangular matrix inverse operation on the artificial intelligence computing platform NPU involves a large number of scalar operations and a high degree of operation dependence, making the entire processing process inefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a processing method and device for triangular matrix inverse operation. Background Art

[0002] Compared with the mature software ecosystem of graphics processing unit GPU (Graphics Processing Unit, GPU for short), the ecological environment of neural processing unit NPU (Neural Processing Unit, NPU for short) is still to be improved. Therefore, the performance optimization of basic operators is more important. Performing triangular matrix inverse operations on the artificial intelligence computing platform NPU restricts the parallelism of the operation process because the operation involves a large number of scalar operations and a high degree of operation dependence. If the traditional element-by-element method is adopted, that is, the inverse matrix calculation is realized through scalar operations combined with a loop structure, the calculation efficiency is low.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] An embodiment of the present invention provides a processing method and device for triangular matrix inverse operation, so as to at least solve the technical problem in the related art that performing triangular matrix inverse operation on an artificial intelligence computing platform NPU involves a large number of scalar operations and a high degree of operation dependence, making the entire processing process inefficient.

[0005] According to one aspect of an embodiment of the present invention, a method for processing a triangular matrix inverse operation is provided, comprising: upon receiving a target task, evenly dividing the target task into multiple computing cores of a target processing platform, wherein the target processing platform is used to process the target task, and the target task is a plurality of blocks on the diagonal of a matrix; obtaining triangular matrices corresponding to the plurality of blocks, respectively, and inverting the triangular matrix to obtain an inverse matrix of the triangular matrix; storing the diagonal elements and the unit matrix of the inverse matrix of the triangular matrix into a first storage medium, respectively, and storing the known part and the coefficient matrix of the inverse matrix into a second storage medium, wherein the known part is a result obtained by solving the inverse matrix, the data in the first storage medium is called by parallel instructions, and the data in the second storage medium is called by interactive operation instructions; retrieving data from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and retrieving data from the second storage medium to perform cube volume calculation on the data in the second storage medium, to obtain an inverse operation result of the inverse matrix.

[0006] Optionally, evenly dividing the target task to multiple computing cores of the target processing platform includes: determining the number of tasks of the target task; and evenly dividing the target task to the multiple computing cores of the target processing platform according to the number of tasks.

[0007] Optionally, determining the number of tasks of the target task includes: determining a first dimension of a plurality of the blocks and a second dimension of the matrix; and rounding up a quotient of the first dimension and the second dimension to obtain the number of tasks.

[0008] Optionally, the target tasks are evenly divided to the multiple computing cores of the target processing platform, including: allocating the target tasks of the number of tasks to the multiple computing cores of the target processing platform in a cyclic allocation manner, wherein, in the process of evenly dividing the target tasks, when the target tasks are allocated to the last one of the multiple computing cores and the target tasks are still not allocated completely, returning to the first one of the multiple computing cores to continue to execute the allocation operation of the target tasks until the target tasks are allocated completely.

[0009] Optionally, triangular matrices corresponding to the multiple blocks are obtained, and the triangular matrices are inverted to obtain the inverse matrix of the triangular matrix, including: loading the multiple blocks into the first storage medium, and generating a mask matrix through a vector instruction in the first storage medium, wherein the lower triangle of the mask matrix is ​​all 1, and the remaining elements of the mask matrix are all 0; element-by-element multiplication of the blocks of the matrix with the mask matrix to invert the triangular matrix and obtain the inverse matrix of the triangular matrix.

[0010] Optionally, data is retrieved from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and data is retrieved from the second storage medium to perform cube volume calculation on the data in the second storage medium to obtain the inverse operation result of the inverse matrix, including: generating a mask matrix through a vector instruction, and multiplying the mask matrix by the matrix element by element to obtain the coefficient matrix; performing matrix multiplication processing on the coefficient matrix and the inverse matrix of the triangular matrix to obtain a result matrix; using the 0th row of the result matrix as the polynomial multiplication result of the 0th row of the inverse matrix; using the vector instruction to perform vector subtraction processing on the 0th row of the result matrix and the multiplication result of the polynomial to obtain a vector subtraction processing result; performing vector multiplication processing on the vector subtraction processing result and the reciprocal of the 0th element on the diagonal to obtain the current 0th row element of the inverse matrix of the triangular matrix; using the current 0th row element to update the original 0th row element of the triangular matrix; repeating the above steps until the elements of all rows of the inverse matrix of the triangular matrix are updated.

[0011] Optionally, after retrieving data from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and retrieving data from the second storage medium to perform cube volume calculation on the data in the second storage medium, and obtaining the inverse operation result of the inverse matrix, the processing method of the triangular matrix inverse operation also includes: updating the inverse matrix of the triangular matrix using the inverse operation result to obtain the updated inverse matrix; and storing the updated inverse matrix back into the global memory.

[0012] According to another aspect of an embodiment of the present invention, a processing device for triangular matrix inverse operation is also provided, including: an allocation unit, used to evenly divide the target task into multiple computing cores of a target processing platform when receiving the target task, wherein the target processing platform is used to process the target task, and the target task is a plurality of blocks on the diagonal of the matrix; an acquisition unit, used to acquire the triangular matrices corresponding to the plurality of blocks, respectively, and invert the triangular matrix to obtain the inverse matrix of the triangular matrix; a storage unit, used to store the diagonal elements and the unit matrix of the inverse matrix of the triangular matrix into a first storage medium, respectively, and store the known part and the coefficient matrix of the inverse matrix into a second storage medium, wherein the known part is the result obtained by solving the inverse matrix, the data in the first storage medium is called by parallel instructions, and the data in the second storage medium is called by interactive operation instructions; a processing unit, used to retrieve data from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and retrieve data from the second storage medium to perform cube volume calculation on the data in the second storage medium, to obtain the inverse operation result of the inverse matrix.

[0013] Optionally, the allocation unit includes: a first determination subunit, used to determine the number of tasks of the target task; and a first allocation subunit, used to evenly divide the target task to the multiple computing cores of the target processing platform according to the number of tasks.

[0014] Optionally, the first determination subunit includes: a determination module, used to determine the first dimensions of the plurality of blocks and the second dimension of the matrix; an acquisition module, used to round up the quotient of the first dimension and the second dimension to obtain the number of tasks.

[0015] Optionally, the allocation unit includes: a second allocation sub-unit, used to allocate the target tasks of the number of tasks to the multiple computing cores of the target processing platform in a round-robin manner, wherein, in the process of evenly dividing the target tasks, when the target task is allocated to the last one of the multiple computing cores and the target task is still not allocated completely, it returns to the first one of the multiple computing cores to continue to execute the allocation operation of the target task until the target task is allocated completely.

[0016] Optionally, the acquisition unit includes: a storage subunit, used to load the multiple blocks into the first storage medium, and generate a mask matrix through a vector instruction in the first storage medium, wherein the lower triangle of the mask matrix is ​​all 1, and the remaining elements of the mask matrix are all 0; multiplying the blocks of the matrix by the mask matrix element by element to invert the triangular matrix and obtain an inverse matrix of the triangular matrix.

[0017] Optionally, the processing unit includes: an acquisition subunit, which is used to generate a mask matrix through a vector instruction, and multiply the mask matrix by the matrix element by element to obtain the coefficient matrix; a first processing subunit, which is used to perform matrix multiplication on the coefficient matrix and the inverse matrix of the triangular matrix to obtain a result matrix; a second determination subunit, which is used to use the 0th row of the result matrix as the polynomial multiplication result of the 0th row of the inverse matrix; a second processing subunit, which is used to perform vector subtraction on the 0th row of the result matrix and the multiplication result of the polynomial using the vector instruction to obtain a vector subtraction result; a third processing subunit, which is used to perform vector multiplication on the vector subtraction result and the reciprocal of the 0th element on the diagonal to obtain the current 0th row elements of the inverse matrix of the triangular matrix; an update subunit, which is used to update the original 0th row elements of the triangular matrix using the current 0th row elements; and a fourth processing subunit, which is used to repeat the above steps until all rows of the inverse matrix of the triangular matrix are updated.

[0018] Optionally, the processing device for the inverse operation of the triangular matrix also includes: an updating unit, which is used to retrieve data from the first storage medium to perform parallel calculations on the data retrieved from the first storage medium, and to retrieve data from the second storage medium to perform cube volume calculations on the data in the second storage medium, and after obtaining the inverse operation result of the inverse matrix, update the inverse matrix of the triangular matrix using the inverse operation result to obtain the updated inverse matrix; and a restoring unit, which is used to store the updated inverse matrix back into the global memory.

[0019] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, wherein the computer-readable storage medium includes a stored program, wherein the program executes any one of the above-mentioned processing methods for the inverse operation of a triangular matrix.

[0020] According to another aspect of an embodiment of the present invention, a processor is provided, wherein the processor is used to run a program, wherein the program executes any one of the above-mentioned processing methods for triangular matrix inverse operations when running.

[0021] According to another aspect of an embodiment of the present invention, a computer program product is provided, comprising computer instructions, wherein when the computer instructions are executed by a processor, the processing method for performing any one of the above-mentioned triangular matrix inverse operations is performed.

[0022] In an embodiment of the present invention, when a target task is received, the target task is evenly divided into multiple computing cores of a target processing platform, wherein the target processing platform is used to process the target task, and the target task is a plurality of blocks on the diagonal of a matrix; triangular matrices corresponding to the plurality of blocks are obtained, and the triangular matrix is ​​inverted to obtain an inverse matrix of the triangular matrix; the diagonal elements and the unit matrix of the inverse matrix of the triangular matrix are respectively stored in a first storage medium, and the known part and the coefficient matrix of the inverse matrix are stored in a second storage medium, wherein the known part is a result obtained by solving the inverse matrix, the data in the first storage medium is called by parallel instructions, and the data in the second storage medium is called by interactive operation instructions; data is retrieved from the first storage medium to perform parallel calculations on the data retrieved from the first storage medium, and data is retrieved from the second storage medium to perform cube volume calculations on the data in the second storage medium, to obtain an inverse operation result of the inverse matrix. Through the above-mentioned technical scheme provided by the present invention, the purpose of increasing the computing speed and increasing the task throughput is achieved by allocating computing tasks to multiple computing cores for parallel computing, thereby achieving the technical effect of improving the processing efficiency of triangular matrix inverse operations, thereby solving the technical problem in the related technology that performing triangular matrix inverse operations on the artificial intelligence computing platform NPU involves a large number of scalar operations and a high degree of computational dependence, resulting in low efficiency of the entire processing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0024] Figure 1 It is a hardware structure block diagram of a mobile terminal of a processing method for a triangular matrix inverse operation according to an embodiment of the present invention;

[0025] Figure 2 is a flowchart of a processing method for triangular matrix inverse operation according to an embodiment of the present invention;

[0026] Figure 3 is a flowchart of an optional processing method for triangular matrix inverse operation according to an embodiment of the present invention;

[0027] Figure 4 is a schematic diagram of target task allocation according to an embodiment of the present invention;

[0028] Figure 5is a schematic diagram of a calculation process of an inverse matrix of a triangular matrix of each block according to an embodiment of the present invention;

[0029] Figure 6 is a schematic diagram of a processing method for triangular matrix inverse operation according to an embodiment of the present invention;

[0030] Figure 7 is a schematic diagram of a processing device for triangular matrix inverse operation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] As introduced in the background technology, the related art involves a large number of scalar operations and a high degree of operation dependence in performing triangular matrix inverse operations on the artificial intelligence computing platform NPU, which makes the entire processing process inefficient. In view of the above defects, a processing method and device for triangular matrix inverse operations, a computer-readable storage medium, a processor, and a computer program product are provided in an embodiment of the present invention.

[0034] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0035] The method embodiments provided in the embodiments of the present invention can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a hardware structure block diagram of a mobile terminal of a method for processing a triangular matrix inverse operation according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0036] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the processing method of the triangular matrix inverse operation in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The transmission device 106 is used to receive or send data via a network. The above-mentioned specific examples of the network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0037] According to an embodiment of the present invention, a method embodiment of a processing method for a triangular matrix inverse operation is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0038] Figure 2 is a flowchart of a method for processing a triangular matrix inverse operation according to an embodiment of the present invention. Figure 2 As shown, the processing method of the triangular matrix inverse operation includes the following steps:

[0039] Step S202, when receiving the target task, evenly divide the target task into multiple computing cores of the target processing platform, wherein the target processing platform is used to process the target task, and the target task is a plurality of blocks on the diagonal of the matrix.

[0040] Optionally, the target processing platform may be a platform with data processing capabilities. In this embodiment, the target processing platform has multiple AI Core computing cores, generally 30 or 32.

[0041] Figure 3 is a flowchart of an optional processing method for triangular matrix inverse operation according to an embodiment of the present invention, such as Figure 3 As shown, first, the computing tasks can be divided and evenly distributed to each computing core of the NPU for parallel computing and application of corresponding resources.

[0042] Step S204, obtaining triangular matrices corresponding to the plurality of blocks, and inverting the triangular matrices to obtain an inverse matrix of the triangular matrix.

[0043] Step S206, store the diagonal elements of the inverse matrix of the triangular matrix and the unit matrix into the first storage medium respectively, and store the known part of the inverse matrix and the coefficient matrix into the second storage medium, wherein the known part is the result obtained by solving the inverse matrix, the data in the first storage medium is called by parallel instructions, and the data in the second storage medium is called by interactive operation instructions.

[0044] Step S208, retrieve data from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and retrieve data from the second storage medium to perform cube volume calculation on the data in the second storage medium, to obtain an inverse operation result of the inverse matrix.

[0045] As can be seen from the above, in an embodiment of the present invention, when a target task is received, the target task can be evenly divided into multiple computing cores of a target processing platform, wherein the target processing platform is used to process the target task, and the target task is a plurality of blocks on the diagonal of a matrix; the triangular matrices corresponding to the plurality of blocks are obtained, and the triangular matrix is ​​inverted to obtain the inverse matrix of the triangular matrix; the diagonal elements and the unit matrix of the inverse matrix of the triangular matrix are respectively stored in a first storage medium, and the known part and the coefficient matrix of the inverse matrix are stored in a second storage medium, wherein the known part is the result obtained by solving the inverse matrix, the data in the first storage medium is called by parallel instructions, and the data in the second storage medium is called by interactive operation instructions; data is retrieved from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and data is retrieved from the second storage medium to perform cube volume calculation on the data in the second storage medium, to obtain the inverse operation result of the inverse matrix, thereby achieving the purpose of increasing the calculation speed and increasing the task throughput by distributing the computing tasks to multiple computing cores for parallel computing, and achieving the technical effect of improving the processing efficiency of the inverse operation of the triangular matrix.

[0046] Therefore, the above-mentioned technical solution provided by the embodiment of the present invention solves the technical problem in the related technology that performing triangular matrix inverse operations on the artificial intelligence computing platform NPU involves a large number of scalar operations and a high degree of operation dependence, making the entire processing process inefficient.

[0047] According to the above embodiment of the present invention, evenly dividing the target tasks to multiple computing cores of the target processing platform includes: determining the number of tasks of the target tasks; and evenly dividing the target tasks to multiple computing cores of the target processing platform according to the number of tasks.

[0048] Figure 4 is a schematic diagram of target task allocation according to an embodiment of the present invention, such as Figure 4 As shown, there are multiple blocks on the diagonal of the input matrix. Each block needs to calculate the inverse matrix of the triangular matrix. The calculation process of all blocks is as follows Figure 5 ( Figure 5 is a schematic diagram of the calculation process of the inverse matrix of the triangular matrix of each block according to an embodiment of the present invention. First, the number of calculation tasks can be determined. According to the needs, the size of the block on the diagonal line that needs to be inverted can be determined, that is, M0, and M0 needs to be aligned with 32B.

[0049] In the above embodiment, determining the number of tasks of the target task includes: determining the first dimensions of the plurality of blocks and the second dimension of the matrix; and rounding up the quotient of the first dimension and the second dimension to obtain the number of tasks.

[0050] Here the total number of tasks tilenum=ceil(M / M0), where ceil means rounding up, M is the dimension of the input matrix, and M0 is the size of the diagonal block.

[0051] According to the above embodiment of the present invention, the target tasks are evenly divided into multiple computing cores of the target processing platform, including: the target tasks of the task quantity are allocated to the multiple computing cores of the target processing platform in a cyclic allocation manner, wherein, in the process of evenly dividing the target tasks, when the target tasks are allocated to the last one of the multiple computing cores, if the target tasks are not allocated completely, they are returned to the first one of the multiple computing cores to continue to execute the allocation operation of the target tasks until the target tasks are allocated completely.

[0052] In this embodiment, in the block division, the dimension of the last block may be less than the size of M0. When performing the inverse calculation, the square matrix of M0 is padded with 0 and then the calculation is performed, but when writing back, only valid data is written back.

[0053] like Figure 4 As shown in the figure, each core is responsible for calculating the inverse matrix of a diagonal block. In terms of task allocation, a circular allocation method is adopted to achieve the effect of load balancing. Specifically, task No. 0 is allocated to computing core No. 0, task segment No. 1 is allocated to computing core No. 1, and so on. When there are still task segments not allocated to the last computing core, it will go back and continue to allocate from computing core No. 0. This allocation strategy can improve the L2Cache hit rate and reduce data transfer time.

[0054] According to the above embodiment of the present invention, triangular matrices corresponding to a plurality of blocks are obtained, and the triangular matrices are inverted to obtain an inverse matrix of the triangular matrix, including: loading the plurality of blocks into a first storage medium, and generating a mask matrix through a vector instruction in the first storage medium, wherein the lower triangle of the mask matrix is ​​all 1, and the remaining elements of the mask matrix are all 0; and element-by-element multiplication of the matrix blocks with the mask matrix to invert the triangular matrix to obtain an inverse matrix of the triangular matrix.

[0055] In this embodiment, for the target processing platform NPU, if the vector and cube units need to be called to calculate data, the data needs to be loaded into the corresponding memory. For the vector operation unit, the data needs to be loaded into the UB memory; for the cube operation unit, it needs to be loaded into the L1 Buffer first. First, the inverse process can be expressed as AX=E, where A is the input matrix, E is the unit matrix, and X is the inverse matrix result to be solved. Different parts need to be read into different on-chip memories. That is, Figure 3 As shown, the corresponding diagonal matrix is ​​read from the global memory and copied to the corresponding memory of the device.

[0056] First, since the blocks on the diagonal of the input matrix are all ordinary matrices, it is necessary to obtain their corresponding triangular matrices, so a mask matrix is ​​needed. Specifically, each block needs to be loaded into UB first, and then a mask matrix is ​​generated through vector instructions. Taking the lower triangle as an example, the lower triangle of the mask matrix is ​​all 1, and the remaining elements are all 0. Each block of the input matrix is ​​multiplied element by element with the mask matrix to obtain the corresponding triangular matrix A', which is stored in a space in UB. Secondly, after obtaining the triangular matrix A', the diagonal elements of the triangular matrix and the unit matrix E need to be stored in a space in UB for subsequent vector instructions to use. Finally, when calculating each row of the inverse matrix, it is necessary to calculate the result of a polynomial multiplication. This calculation step is completed by the cube high-speed computing unit. The obtained inverse matrix part (initially all 0) and the coefficient matrix need to be loaded into L1 for subsequent calls to the cube unit instruction. That is, if Figure 3 As shown, the vector operation unit and the cube operation unit can be called to complete the process of solving the inverse matrix.

[0057] According to the above embodiment of the present invention, data is retrieved from a first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and data is retrieved from a second storage medium to perform cube volume calculation on the data in the second storage medium to obtain the inverse operation result of the inverse matrix, including: generating a mask matrix through a vector instruction, and multiplying the mask matrix and the matrix element by element to obtain a coefficient matrix; performing matrix multiplication processing on the coefficient matrix and the inverse matrix of the triangular matrix to obtain a result matrix; using the 0th row of the result matrix as the polynomial multiplication result of the 0th row of the inverse matrix; performing vector subtraction processing on the 0th row of the result matrix and the polynomial multiplication result using the vector instruction to obtain the vector subtraction processing result; performing vector multiplication processing on the vector subtraction processing result and the reciprocal of the 0th element on the diagonal to obtain the current 0th row element of the inverse matrix of the triangular matrix; using the current 0th row element to update the original 0th row element of the triangular matrix; repeating the above steps until the elements of all rows of the inverse matrix of the triangular matrix are updated.

[0058] In this embodiment, each block first needs to be loaded into UB, and then a mask matrix is ​​generated through vector instructions. Taking the lower triangle as an example, the lower triangle of the mask matrix is ​​all 1, and the remaining elements are all 0. Each block of the input matrix is ​​multiplied element by element with the mask matrix to obtain the corresponding triangular matrix A', which is stored in a space in UB. Then, after obtaining the triangular matrix A', the diagonal elements of the triangular matrix and the unit matrix E need to be stored in a space in UB for subsequent vector instructions to use. When calculating each row of the inverse matrix, it is necessary to calculate the result of a polynomial multiplication. This calculation step is completed by the cube high-speed computing unit. The calculated inverse matrix part (initially all 0) and the coefficient matrix need to be loaded into L1 for subsequent calls to the cube unit instruction.

[0059] Each block and the part required for calculation have been loaded into the corresponding memory. Here, the inverse matrix calculation of a block is completed for each core, such as Figure 4 As shown, calculating a block requires iteratively calculating each row of the inverse matrix from top to bottom, and the calculation of subsequent rows will use the calculation results of the previous rows.

[0060] In the embodiment of the present invention, the calculation formula of each row of the inverse matrix can be expressed as: the i-th row of the inverse matrix = (the i-th row of the result matrix - the i-th row of the result matrix of the multiplication of the polynomial) / (diag [i]), i = 0, 1, ..., M_0-1. In the above steps, the parts required for the calculation have been stored in the corresponding memory. The calculation process of this solution is designed to use vector and cube calculation units for joint processing to speed up the calculation speed.

[0061] It should be noted that the elements in the i-th row and the i-th element on the diagonal of the result matrix have been stored in the vector space in UB in the above steps and can be directly used.

[0062] Figure 6 is a schematic diagram of a processing method for triangular matrix inverse operation according to an embodiment of the present invention, such as Figure 6As shown, the polynomial multiplication result matrix of each row needs to be solved, and the polynomial result can be obtained through matrix multiplication. This can be divided into two steps. First, you need to use the vector instruction to generate a mask matrix. The diagonal and above the diagonal of the matrix are 0, and the diagonal is 1; then the triangular matrix A' and the mask matrix are multiplied element by element through the vector instruction to obtain the coefficient matrix A''. Then you need to use the cube operation unit to load the coefficient matrix A'' into the L1 Buffer and then into L0A, load the current inverse matrix result into the L1 Buffer and then into L0B, and call the matrix operation instruction to obtain the result C of the matrix multiplication. Each row of the C matrix is ​​the polynomial multiplication result of the row that needs to be solved. Among them, pay attention to the arrangement of the data when calling the cube operation unit, and the result matrix obtained by the cube at the end has a specific arrangement format, and it is necessary to cooperate with the data arrangement transformation to extract the elements of the same row.

[0063] In this embodiment, the inverse matrix of the row is solved. First, the result of the cube calculation needs to be moved back to UB and the elements of the i-th row are read out. Since multiplication is smaller than division, the division in the formula is converted to multiplication. The result of the row of the inverse matrix is ​​calculated through vector instructions vsub and vmul, and updated to the inverse matrix.

[0064] Therefore, the inverse matrix calculation process of each block is roughly as follows: First, generate a mask matrix through vector, and multiply it element by element with the input matrix to obtain the coefficient matrix. Second, the initial inverse matrix is ​​all 0, and the coefficient matrix is ​​multiplied with the current inverse matrix to obtain the result matrix C, and the 0th row of the C matrix is ​​taken as the polynomial multiplication result for solving the 0th row of the inverse matrix. Third, use the vector instructions vsub and vmul to perform vector subtraction on the 0th row of the result matrix (unit matrix) and the result of the polynomial multiplication obtained in step 2, and then perform vector multiplication on the result and the reciprocal of the 0th element of the diagonal, and the result is the 0th row of the inverse matrix. Fourth, update the 0th row element of the inverse matrix, and repeat steps 1-3 to update the remaining rows of the inverse matrix until the inverse matrix is ​​calculated. This calculation scheme abandons the scalar operation unit and uses all vector and matrix operations, which greatly improves the efficiency of solving the inverse matrix.

[0065] According to the above-mentioned embodiment of the present invention, after retrieving data from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and retrieving data from the second storage medium to perform cube volume calculation on the data in the second storage medium, and obtaining the inverse operation result of the inverse matrix, the processing method of the triangular matrix inverse operation also includes: updating the inverse matrix of the triangular matrix using the inverse operation result to obtain an updated inverse matrix; and storing the updated inverse matrix back into the global memory.

[0066] In this embodiment, after calculating the inverse matrix of the corresponding matrix block, the inverse matrix needs to be stored back in the global memory. Since there are multiple diagonal blocks and only the inverse matrix result needs to be stored back, another global memory is allocated to transfer the calculated inverse matrix result to this global memory. It should be noted that the last block needs to be processed separately and only the valid data is stored back. Specifically, assume that the dimension of the last block is M1 < M0. When storing the result of the last inverse matrix, the matrix with dimension M1 is stored back in the GM.

[0067] Through the technical solution provided by the above embodiment of the present invention, the calculation tasks are first assigned to multiple computing cores for parallel computing, which improves the computing speed and increases the task throughput. This is manifested as different blocks on the diagonal of the input matrix being assigned to different cores for parallel computing, and the calculation of each block is performed iteratively within each core. In addition, high-speed vector and cube computing units are used to replace the traditional slow scalar operations to accelerate the process of finding the inverse of the triangular block matrix. At the same time, since different computing units can be parallel, the speed of the computing process can be further improved and the task throughput can be increased through good pipeline arrangement. That is, in the embodiment of the present invention, high-speed vector and cube computing units are used to replace the traditional slow scalar operations to accelerate the process of finding the inverse of the triangular block matrix. At the same time, since different computing units can be parallel, the speed of the computing process can be further improved and the task throughput can be increased through good pipeline arrangement, which is more obvious when there are many computing tasks.

[0068] Due to the above technical solution provided in the embodiment of the present invention, when finding the inverse matrix of the upper triangular block on the diagonal of the matrix on the target processing platform, there are multiple block matrices on the diagonal of the input matrix. The process of finding the inverse can be expressed as AX = E, where A is an upper triangular matrix or a lower triangular matrix. In the present invention, high-speed computing vectors and matrix computing units are used to parallelly solve the inverse matrices of different blocks, accelerating the computing process and solving the limitations of the traditional method of solving elements one by one through scalars and for loops. The traditional method has a slow computing speed and low efficiency.

[0069] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0070] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0071] According to an embodiment of the present invention, a processing device for a triangular matrix inverse operation for implementing the above-mentioned processing method for a triangular matrix inverse operation is also provided. Figure 7 is a schematic diagram of a processing device for triangular matrix inverse operation according to an embodiment of the present invention, such as Figure 7 As shown, the device includes: an allocating unit 701, an acquiring unit 703, a storage unit 705 and a processing unit 707. The processing device for the inverse operation of the triangular matrix is ​​described below.

[0072] The allocation unit 701 is used to evenly divide the target task into multiple computing cores of the target processing platform when receiving the target task, wherein the target processing platform is used to process the target task, and the target task is a plurality of blocks on the diagonal of the matrix.

[0073] The acquisition unit 703 is used to acquire triangular matrices corresponding to the multiple blocks, and invert the triangular matrices to obtain an inverse matrix of the triangular matrix.

[0074] Storage unit 705 is used to store the diagonal elements of the inverse matrix of the triangular matrix and the unit matrix into the first storage medium respectively, and store the known part of the inverse matrix and the coefficient matrix into the second storage medium, wherein the known part is the result obtained by solving the inverse matrix, the data in the first storage medium is called by parallel instructions, and the data in the second storage medium is called by interactive operation instructions.

[0075] The processing unit 707 is used to retrieve data from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and retrieve data from the second storage medium to perform cube volume calculation on the data in the second storage medium to obtain the inverse operation result of the inverse matrix.

[0076] It should be noted here that the above-mentioned allocation unit 701, acquisition unit 703, storage unit 705 and processing unit 707 correspond to steps S202 to S208 in the above-mentioned embodiment, and the four units and the corresponding steps implement the same instances and application scenarios, but are not limited to the contents disclosed in the above-mentioned embodiment.

[0077] As can be seen from the above, in the scheme recorded in the above embodiment of the present invention, the allocation unit can be used to evenly divide the target task into multiple computing cores of the target processing platform when receiving the target task, wherein the target processing platform is used to process the target task, and the target task is a plurality of blocks on the diagonal of the matrix; the acquisition unit is used to obtain the triangular matrices corresponding to the plurality of blocks, and the triangular matrix is ​​inverted to obtain the inverse matrix of the triangular matrix; the storage unit is used to store the diagonal elements and the unit matrix of the inverse matrix of the triangular matrix into the first storage medium, respectively, and the known part and the coefficient matrix of the inverse matrix are stored into the second storage medium, wherein the known Part is the result of solving the inverse matrix, the data in the first storage medium is called by parallel instructions, and the data in the second storage medium is called by interactive operation instructions; the processing unit is used to call data from the first storage medium to perform parallel calculations on the data called from the first storage medium, and to call data from the second storage medium to perform cube volume calculations on the data in the second storage medium, so as to obtain the inverse operation result of the inverse matrix, thereby achieving the purpose of increasing the calculation speed and increasing the task throughput by distributing the calculation tasks to multiple computing cores for parallel calculations, and achieving the technical effect of improving the processing efficiency of the inverse operation of the triangular matrix.

[0078] Therefore, the above-mentioned technical solution provided by the embodiment of the present invention solves the technical problem in the related technology that performing triangular matrix inverse operations on the artificial intelligence computing platform NPU involves a large number of scalar operations and a high degree of operation dependence, making the entire processing process inefficient.

[0079] Optionally, the allocation unit includes: a first determination subunit, used to determine the number of tasks of the target task; and a first allocation subunit, used to evenly divide the target task into multiple computing cores of the target processing platform according to the number of tasks.

[0080] Optionally, the first determination subunit includes: a determination module, used to determine the first dimension of the plurality of blocks and the second dimension of the matrix; and an acquisition module, used to round up the quotient of the first dimension and the second dimension to obtain the number of tasks.

[0081] Optionally, the allocation unit includes: a second allocation sub-unit, used to allocate a target number of tasks to multiple computing cores of the target processing platform in a round-robin manner, wherein, in the process of evenly dividing the target tasks, when the target tasks are allocated to the last one of the multiple computing cores and are still not allocated, it returns to the first one of the multiple computing cores to continue executing the allocation operation of the target tasks until the target tasks are allocated.

[0082] Optionally, the acquisition unit includes: a storage subunit, used to load multiple blocks into a first storage medium, and generate a mask matrix through a vector instruction in the first storage medium, wherein the lower triangle of the mask matrix is ​​all 1, and the remaining elements of the mask matrix are all 0; multiplying the blocks of the matrix by the mask matrix element by element to invert the triangular matrix and obtain an inverse matrix of the triangular matrix.

[0083] Optionally, the processing unit includes: an acquisition subunit, which is used to generate a mask matrix through a vector instruction, and multiply the mask matrix and the matrix element by element to obtain a coefficient matrix; a first processing subunit, which is used to perform matrix multiplication on the coefficient matrix and the inverse matrix of the triangular matrix to obtain a result matrix; a second determination subunit, which is used to use the 0th row of the result matrix as the polynomial multiplication result of the 0th row of the inverse matrix; a second processing subunit, which is used to perform vector subtraction on the 0th row of the result matrix and the polynomial multiplication result using the vector instruction to obtain a vector subtraction result; a third processing subunit, which is used to perform vector multiplication on the vector subtraction result and the reciprocal of the 0th element on the diagonal to obtain the current 0th row elements of the inverse matrix of the triangular matrix; an update subunit, which is used to update the original 0th row elements of the triangular matrix using the current 0th row elements; and a fourth processing subunit, which is used to repeat the above steps until the elements of all rows of the inverse matrix of the triangular matrix are updated.

[0084] Optionally, the processing device for the inverse operation of the triangular matrix also includes: an updating unit, which is used to retrieve data from the first storage medium to perform parallel calculations on the data retrieved from the first storage medium, and to retrieve data from the second storage medium to perform cube volume calculations on the data in the second storage medium, and after obtaining the inverse operation results of the inverse matrix, update the inverse matrix of the triangular matrix using the inverse operation results to obtain an updated inverse matrix; and a restoring unit, which is used to store the updated inverse matrix back into the global memory.

[0085] According to another aspect of an embodiment of the present invention, a processor is provided, and the processor is used to run a program, wherein when the program is run, any one of the above-mentioned processing methods for inverse operation of a triangular matrix is ​​executed.

[0086] According to another aspect of an embodiment of the present invention, a computer program product is provided, including computer instructions, which, when executed by a processor, perform any one of the above-mentioned processing methods for inverse operation of a triangular matrix.

[0087] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein the program executes any one of the above-mentioned processing methods for the inverse operation of a triangular matrix.

[0088] Optionally, in this embodiment, the computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the communication devices in a communication device group.

[0089] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: when receiving a target task, evenly divide the target task into multiple computing cores of a target processing platform, wherein the target processing platform is used to process the target task, and the target task is a plurality of blocks on the diagonal of a matrix; obtain triangular matrices corresponding to the plurality of blocks, and invert the triangular matrix to obtain an inverse matrix of the triangular matrix; store the diagonal elements and the unit matrix of the inverse matrix of the triangular matrix into a first storage medium, respectively, and store the known part and the coefficient matrix of the inverse matrix into a second storage medium, wherein the known part is the result obtained by solving the inverse matrix, the data in the first storage medium is called by parallel instructions, and the data in the second storage medium is called by interactive operation instructions; retrieve data from the first storage medium to perform parallel calculations on the data retrieved from the first storage medium, and retrieve data from the second storage medium to perform cube volume calculations on the data in the second storage medium to obtain the inverse operation result of the inverse matrix.

[0090] Optionally, in this embodiment, the computer-readable storage medium is configured to store program codes for executing the following steps: determining the number of tasks of the target task; and evenly dividing the target task to multiple computing cores of the target processing platform according to the number of tasks.

[0091] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: determining a first dimension of multiple blocks and a second dimension of the matrix; and rounding up a quotient of the first dimension and the second dimension to obtain the number of tasks.

[0092] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: allocating a target number of tasks to multiple computing cores of the target processing platform in a round-robin manner, wherein, in the process of evenly dividing the target tasks, if the target tasks are not fully allocated when they are allocated to the last one of the multiple computing cores, returning to the first one of the multiple computing cores to continue to execute the target task allocation operation until the target tasks are fully allocated.

[0093] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: loading multiple blocks into a first storage medium, and generating a mask matrix through a vector instruction in the first storage medium, wherein the lower triangle of the mask matrix is ​​all 1, and the remaining elements of the mask matrix are all 0; multiplying the blocks of the matrix by the mask matrix element by element to invert the triangular matrix to obtain an inverse matrix of the triangular matrix.

[0094] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: generating a mask matrix through a vector instruction, and multiplying the mask matrix by the matrix element by element to obtain a coefficient matrix; performing matrix multiplication on the coefficient matrix and the inverse matrix of the triangular matrix to obtain a result matrix; using the 0th row of the result matrix as the polynomial multiplication result of the 0th row of the inverse matrix; performing vector subtraction on the 0th row of the result matrix and the result of the polynomial multiplication using the vector instruction to obtain a vector subtraction result; performing vector multiplication on the vector subtraction result and the reciprocal of the 0th element on the diagonal to obtain the current 0th row elements of the inverse matrix of the triangular matrix; using the current 0th row elements to update the original 0th row elements of the triangular matrix; repeating the above steps until the elements of all rows of the inverse matrix of the triangular matrix are updated.

[0095] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: after retrieving data from the first storage medium to perform parallel calculations on the data retrieved from the first storage medium, and retrieving data from the second storage medium to perform cube volume calculations on the data in the second storage medium, and obtaining the inverse operation results of the inverse matrix, the inverse matrix of the triangular matrix is ​​updated using the inverse operation results to obtain an updated inverse matrix; and the updated inverse matrix is ​​stored back into the global memory.

[0096] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0097] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0098] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0099] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0100] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0101] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.

[0102] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for processing a triangular matrix inverse operation, characterized in that: include: When receiving a target task, evenly divide the target task into multiple computing cores of a target processing platform, wherein the target processing platform is used to process the target task, and the target task is a plurality of blocks on the diagonal of a matrix; Obtaining triangular matrices corresponding to the plurality of blocks respectively, and inverting the triangular matrices to obtain an inverse matrix of the triangular matrix; The diagonal elements and the unit matrix of the inverse matrix of the triangular matrix are stored in a first storage medium respectively, and the known part and the coefficient matrix of the inverse matrix are stored in a second storage medium, wherein the known part is the result obtained by solving the inverse matrix, the data in the first storage medium is called by the parallel instruction, and the data in the second storage medium is called by the interactive operation instruction; Retrieving data from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and retrieving data from the second storage medium to perform cube volume calculation on the data in the second storage medium, to obtain an inverse operation result of the inverse matrix; Wherein, obtaining the triangular matrices corresponding to the plurality of blocks respectively, and inverting the triangular matrices to obtain the inverse matrix of the triangular matrix, comprises: loading the plurality of blocks into the first storage medium, and generating a mask matrix through a vector instruction in the first storage medium, wherein the lower triangle of the mask matrix is ​​all 1, and the remaining elements of the mask matrix are all 0; performing element-by-element multiplication of the blocks of the matrix with the mask matrix to invert the triangular matrix to obtain the inverse matrix of the triangular matrix; Among them, calling data from the first storage medium to perform parallel calculation on the data called from the first storage medium, and calling data from the second storage medium to perform cube volume calculation on the data in the second storage medium to obtain the inverse operation result of the inverse matrix, including: generating a mask matrix through a vector instruction, and multiplying the mask matrix with the matrix element by element to obtain the coefficient matrix; performing matrix multiplication processing on the coefficient matrix and the inverse matrix of the triangular matrix to obtain a result matrix; using the 0th row of the result matrix as the polynomial multiplication result of the 0th row of the inverse matrix; using the vector instruction to perform vector subtraction processing on the 0th row of the result matrix and the multiplication result of the polynomial to obtain a vector subtraction processing result; performing vector multiplication processing on the vector subtraction processing result and the reciprocal of the 0th element on the diagonal to obtain the current 0th row element of the inverse matrix of the triangular matrix; using the current 0th row element to update the original 0th row element of the triangular matrix; repeating the above steps until the elements of all rows of the inverse matrix of the triangular matrix are updated.

2. The method for processing triangular matrix inverse operation according to claim 1, characterized in that: Evenly dividing the target task into multiple computing cores of the target processing platform includes: Determine the number of tasks for the target task; The target tasks are evenly divided into the multiple computing cores of the target processing platform according to the number of tasks.

3. The method for processing the inverse operation of a triangular matrix according to claim 2, characterized in that: Determine the number of tasks for the target task, including: Determining a first dimension of a plurality of the blocks and a second dimension of the matrix; The number of tasks is obtained by rounding up the quotient of the first dimension and the second dimension.

4. The method for processing triangular matrix inverse operation according to claim 2, characterized in that: Evenly dividing the target task into multiple computing cores of the target processing platform includes: The target tasks of the number of tasks are allocated to the multiple computing cores of the target processing platform in a cyclic allocation manner, wherein, in the process of evenly dividing the target tasks, when the target tasks are allocated to the last one of the multiple computing cores and the target tasks are still not allocated completely, the allocation operation of the target tasks is returned to the first one of the multiple computing cores to continue executing until the target tasks are allocated completely.

5. The method for processing triangular matrix inverse operation according to claim 1, characterized in that: After retrieving data from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and retrieving data from the second storage medium to perform cube volume calculation on the data in the second storage medium, and obtaining the inverse operation result of the inverse matrix, the method further includes: Using the inverse operation result, the inverse matrix of the triangular matrix is ​​updated to obtain the updated inverse matrix; The updated inverse matrix is ​​stored back into the global memory.

6. A processing device for triangular matrix inverse operation, characterized in that: include: an allocating unit, configured to evenly divide the target task into a plurality of computing cores of a target processing platform when receiving the target task, wherein the target processing platform is configured to process the target task, and the target task is a plurality of blocks on the diagonal of the matrix; An acquisition unit, used for acquiring triangular matrices corresponding to the plurality of blocks respectively, and inverting the triangular matrices to obtain an inverse matrix of the triangular matrix; A storage unit, used for storing the diagonal elements of the inverse matrix of the triangular matrix and the unit matrix into a first storage medium respectively, and storing the known part and the coefficient matrix of the inverse matrix into a second storage medium, wherein the known part is a result obtained by solving the inverse matrix, the data in the first storage medium is called by parallel instructions, and the data in the second storage medium is called by interactive operation instructions; a processing unit, configured to retrieve data from the first storage medium to perform parallel calculation on the data retrieved from the first storage medium, and retrieve data from the second storage medium to perform cube volume calculation on the data in the second storage medium, to obtain an inverse operation result of the inverse matrix; The acquisition unit includes: a storage subunit, which is used to load the multiple blocks into the first storage medium, and generate a mask matrix in the first storage medium through a vector instruction, wherein the lower triangle of the mask matrix is ​​all 1, and the remaining elements of the mask matrix are all 0; multiply the blocks of the matrix by the mask matrix element by element to invert the triangular matrix and obtain an inverse matrix of the triangular matrix; Wherein, the processing unit includes: an acquisition subunit, which is used to generate a mask matrix through a vector instruction, and multiply the mask matrix with the matrix element by element to obtain the coefficient matrix; a first processing subunit, which is used to perform matrix multiplication processing on the coefficient matrix and the inverse matrix of the triangular matrix to obtain a result matrix; a second determination subunit, which is used to use the 0th row of the result matrix as the polynomial multiplication result of the 0th row of the inverse matrix; a second processing subunit, which is used to perform vector subtraction processing on the 0th row of the result matrix and the multiplication result of the polynomial using the vector instruction to obtain a vector subtraction processing result; a third processing subunit, which is used to perform vector multiplication processing on the vector subtraction processing result and the reciprocal of the 0th element on the diagonal to obtain the current 0th row element of the inverse matrix of the triangular matrix; an update subunit, which is used to update the original 0th row element of the triangular matrix using the current 0th row element; a fourth processing subunit, which is used to repeat the above steps until the elements of all rows of the inverse matrix of the triangular matrix are updated.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the processing method for triangular matrix inverse operation as described in any one of claims 1 to 5.

8. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, a processing method for triangular matrix inverse operation as described in any one of claims 1 to 5 is performed.

Citation Information

Patent Citations

  • Matrix inversion method, device and apparatus and computer readable storage medium

    CN110377875A

  • Method, device and medium for triangular matrix inversion

    CN116089786A