Matrix multiplication and addition operation implementation method and device based on RISC-V and medium

Through the matrix multiplication and addition operation implementation method based on RISC-V, the problem of low efficiency of traditional processors when processing large-scale matrix operations is solved. Through flexible instruction sets and parallel computing units, the computing efficiency and speed are improved.

CN119939097APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510010679.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When traditional processors process large-scale matrix operations, the instruction set is not flexible enough and cannot include all the information in the matrix multiplication and addition operation, resulting in low computing efficiency, frequent instruction transmission and more data movement.

Method used

The matrix multiplication and addition operation implementation method based on RISC-V is adopted. By obtaining the matrix multiplication and addition operation requirement information, corresponding matrix calculation instructions are generated, and the instruction parameter information is determined through the instruction pipeline processing. According to the operand calculation mode parameters, a matrix multiplication calculation unit is constructed, and the unit is used to perform parallel calculations to determine the calculation result.

Benefits of technology

Improve computing efficiency, reduce instruction transmission and data movement, make full use of multi-core and multi-threading technologies, and significantly improve the speed and efficiency of matrix operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939097A_ABST
    Figure CN119939097A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a matrix multiplication and addition operation implementation method and device based on RISC-V. The method comprises the steps that matrix multiplication and addition operation demand information is obtained, a corresponding matrix calculation instruction is generated based on the matrix multiplication and addition operation demand information according to a preset instruction format, and the matrix calculation instruction is sent to a computer; the matrix calculation instruction comprises an operand address and calculation mode information; the matrix calculation instruction is processed through an instruction pipeline, instruction parameter information is determined, and the instruction parameter information comprises multiple pieces of operation data, corresponding operand calculation mode parameters and bias data; according to the operand calculation mode parameters, a corresponding matrix multiplication and addition calculation unit is determined through a pre-constructed matrix multiplication and addition module, and the matrix multiplication and addition calculation unit comprises any one of an integer matrix multiplication and addition calculation unit and a floating-point number matrix multiplication and addition calculation unit; and performing parallel calculation on the plurality of operation data by utilizing a matrix multiplication and addition calculation unit, and determining an operation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present specification relates to the field of computer technology, and in particular to a method, device and medium for implementing matrix multiplication and addition operations based on RISC-V. Background Art

[0002] In today's computing field, matrix multiplication and addition (MMA) has become one of the core operations in key applications such as deep learning, big data analysis, image processing, and high-performance scientific computing. The wide application of this operation stems from its powerful data processing capabilities and support for complex computing patterns.

[0003] With the explosive growth of data volume, the demand for computing efficiency and performance is becoming increasingly urgent. Processors with traditional architectures often seem unable to cope with large-scale matrix operations. First, the instruction set of traditional processors is often not flexible enough. One instruction cannot include all the information required in the matrix multiplication and addition process. It is necessary to transmit multiple instructions during the calculation process, and the data movement during the calculation process is increased. In addition, traditional processors have limitations in parallel processing and it is difficult to fully utilize multi-core or multi-threading technology to accelerate matrix operations. Therefore, the current matrix multiplication operation is limited by instructions and cannot include the required information during the operation process. The number of instruction transmissions and data movement during the calculation process increases, resulting in low efficiency in processing matrix operations. Summary of the invention

[0004] One or more embodiments of the present specification provide a RISC-V-based matrix multiplication and addition operation implementation method, device and medium, which are used to solve the following technical problems: the current matrix multiplication operation is limited by instructions and cannot include the required information during the operation process, which increases the number of instruction transmissions and data movement during the calculation process, resulting in low efficiency in processing matrix operations.

[0005] One or more embodiments of this specification adopt the following technical solutions:

[0006] One or more embodiments of the present specification provide a matrix multiplication and addition operation implementation method based on RISC-V, characterized in that the method includes: obtaining matrix multiplication and addition operation requirement information, and generating corresponding matrix calculation instructions based on the matrix multiplication and addition operation requirement information and in accordance with a preset instruction format, wherein the matrix calculation instruction includes an operand address and calculation mode information; processing the matrix calculation instruction through an instruction pipeline to determine instruction parameter information, wherein the instruction parameter information includes multiple operation data, corresponding operand calculation mode parameters and bias data; determining the corresponding matrix multiplication and addition calculation unit through a pre-constructed matrix multiplication and addition module according to the operand calculation mode parameters, wherein the matrix multiplication and addition calculation unit includes any one of an integer matrix multiplication and addition calculation unit and a floating-point matrix multiplication and addition calculation unit; using the matrix multiplication and addition calculation unit to perform parallel calculations on the multiple operation data to determine the operation result.

[0007] Furthermore, the instruction format includes an instruction opcode identification bit, a plurality of operand address storage identification bits, a biased data storage identification bit, a calculation mode control information identification bit and a rounding mode enable identification bit.

[0008] Furthermore, the matrix calculation instruction is processed through the instruction pipeline to determine the instruction parameter information, specifically including: analyzing the matrix calculation instruction through the instruction pipeline to determine the operand storage address and the bias data storage address corresponding to each operand; and obtaining the multiple operation data and the bias data according to the operand storage address and the bias data storage address.

[0009] Furthermore, according to the operand calculation mode parameters, a corresponding matrix multiplication and addition calculation unit is determined through a pre-constructed matrix multiplication and addition module, specifically including: according to the operand calculation mode parameters, the data types of the multiple operation data are identified to determine the operation matrix data type, wherein the operation matrix data type includes integer type data and floating-point type data; when the operation matrix data type is integer type data, the matrix multiplication and addition calculation unit is determined to be an integer matrix multiplication and addition calculation unit; when the operation matrix data type is a floating-point type, the matrix multiplication and addition calculation unit is determined to be a floating-point matrix multiplication and addition calculation unit, and the rounding mode instruction in the matrix calculation instruction is sent to the floating-point matrix multiplication and addition calculation unit.

[0010] Furthermore, the matrix multiplication and addition calculation unit is used to perform parallel calculation on the multiple operation data to determine the operation result, which specifically includes: performing matrix allocation on the multiple operation data through the matrix multiplication and addition calculation unit to determine the operation elements of each vector inner product calculation unit; performing parallel multiplication and addition operations on the operation elements of each vector inner product calculation unit to determine the output result of each vector inner product calculation unit; and determining the operation result with the output results of multiple vector inner product calculation units.

[0011] Furthermore, the matrix multiplication and addition calculation unit is used to perform matrix allocation on the multiple operation data to determine the operation elements of each vector inner product calculation unit, specifically including: dividing each of the operation data according to rows and columns to determine at least one operation element corresponding to each operation data, wherein the operation element includes row elements and column elements; determining row and column parameters of the bias data according to the row identifier of the row element and the column identifier of the column element corresponding to the operation element, so as to determine the bias parameter in the bias data based on the row and column parameters; determining the element row and column identifier of the operation element corresponding to each operation data, and determining the operation element of each vector inner product calculation unit and the element position of each operation element in the corresponding operation data according to the element row and column identifier and a preset allocation rule.

[0012] Furthermore, parallel multiplication and addition operations are performed on the operation elements of each of the vector inner product calculation units to determine the output result of each of the vector inner product calculation units, specifically including: in each of the vector inner product calculation units, according to the element position of each of the operation elements in the corresponding operation data, multiplication operations are performed on the operation elements at specified positions to obtain multiple multiplication results; and accumulation operations are performed on the multiple multiplication results and the bias parameters to determine the output result of the vector inner product calculation unit.

[0013] Furthermore, after the matrix multiplication and addition calculation unit is used to perform parallel calculation on the multiple operation data and the calculation result is determined, the method also includes: determining the storage location of the calculation result according to the bias data storage identification bit in the matrix calculation instruction, so as to store the calculation result at the storage location.

[0014] One or more embodiments of this specification provide a matrix multiplication and addition operation implementation device based on RISC-V, including:

[0015] at least one processor; and,

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.

[0018] One or more embodiments of the present specification provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.

[0019] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: through the above technical solutions, compared with traditional processors, the instruction set of the RISC-V architecture is more flexible, and can allow the design of more complex and specialized instruction formats to include the complete information required for matrix multiplication and addition operations, reducing the number of instructions that need to be transmitted during the calculation process, reducing the complexity of instruction decoding, and reducing the number of data movements, thereby improving calculation efficiency; when traditional processors process matrix operations, due to the limitations of the instruction set, they often need to move data frequently to perform different calculation steps, while the matrix multiplication and addition operation implementation method based on RISC-V can reduce unnecessary data movement, reduce storage overhead, and improve by optimizing the instruction format and calculation process. Efficiency of data processing; Traditional processors have limitations in parallel processing, and it is difficult to fully utilize the multi-core and multi-threading technology of modern processors to accelerate matrix operations. The matrix multiplication and addition operation implementation method based on RISC-V can efficiently perform parallel calculations by building special matrix multiplication and addition modules and computing units, and fully utilize the advantages of multi-core and multi-threading technology to significantly improve the computing speed and efficiency; The matrix addition operation implementation method based on RISC-V can support complex computing modes and efficiently process large-scale data by providing multiple computing modes and optimized computing units, thereby meeting the needs of these key applications; By optimizing the instruction format, computing process and parallel processing strategy, it can reduce power consumption and improve energy efficiency to meet the application requirements of long-term operation and high efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings:

[0021] Figure 1 A flowchart of a method for implementing matrix multiplication and addition operations based on RISC-V provided in an embodiment of this specification;

[0022] Figure 2A schematic diagram of the instruction format of a matrix calculation instruction provided in an embodiment of this specification;

[0023] Figure 3 A schematic diagram of a calling flow of a matrix multiplication and addition module provided in an embodiment of this specification;

[0024] Figure 4 A schematic diagram of the architecture of a matrix multiplication and addition computing unit provided in an embodiment of this specification;

[0025] Figure 5 A schematic diagram of the structure of a RISC-V-based matrix multiplication and addition operation implementation device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0027] The embodiments of this specification provide a method for implementing matrix multiplication and addition operations based on RISC-V. RISC-V is an open instruction set architecture (ISA) based on the principles of reduced instruction set computing (RISC). It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 A flowchart of a matrix multiplication and addition operation implementation method based on RISC-V is provided in the embodiment of this specification, such as Figure 1 As shown, it mainly includes the following steps:

[0028] Step S101, obtaining matrix multiplication and addition operation requirement information, and generating corresponding matrix calculation instructions based on the matrix multiplication and addition operation requirement information and in accordance with a preset instruction format.

[0029] Wherein, the matrix calculation instruction includes operand address and calculation mode information;

[0030] The instruction format includes an instruction operation code identification bit, a plurality of operand address storage identification bits, a biased data storage identification bit, a calculation mode control information identification bit and a rounding mode enabling identification bit.

[0031] In one embodiment of the present specification, when performing matrix multiplication and addition operations, matrix multiplication and addition operation requirement information is determined according to the operation scenario of the matrix multiplication and addition operation. The matrix multiplication and addition operation requirement information here is used to determine the operation object, operation method, etc. of this operation. The matrix multiplication and addition operation requirement information usually comes from the requirements of upper-level applications or algorithms, including the size of the matrix, data type (such as floating point, integer), specific form of calculation (such as A*B+C), and other possible forms of operation such as bias, rounding mode, etc.

[0032] Based on the R-type instruction format of the RISC-V instruction set, the matrix multiplication and addition instructions are customized. Figure 2 The following is a schematic diagram of the instruction format of a matrix calculation instruction provided in an embodiment of this specification. Figure 2 The instruction shown contains the calculation mode required for matrix multiplication and addition, the rounding mode of floating-point operation, the values ​​of source and destination registers, and other parameters, providing the matrix multiplication and addition module with the choice of calculation mode and rounding mode. This instruction is an R-type instruction format based on the RISC-V instruction set, in which the opcode stores the matrix calculation instruction operation code, which is used to identify and activate the matrix calculation function; the addresses pointed to by rs1 and rs2 store the data of the matrices corresponding to the two cloud operands, and the address pointed to by rd is not only designed to store the bias value in the convolution operation, but more importantly, it is used to store the final result of the matrix multiplication and addition or convolution operation; funct7 carries the control information of the calculation mode, provides the choice of integer data calculation and floating-point data calculation, indicating that the matrix multiplication and addition operation module supports the operation of both integer and floating-point data; funct3 is specific to the floating-point operation scenario, which is used to indicate the rounding mode required in the operation process to ensure the accuracy and controllability of the calculation result. Specifically, this instruction allows the calculation mode of the current calculation requirement to be dynamically selected according to the input data during execution, thereby effectively reducing unnecessary calculation steps and intermediate data storage. It directly optimizes the core links of matrix operations, significantly improves the processing speed of matrix multiplication and addition modules, reduces data movement, and improves the execution efficiency of matrix operations.

[0033] First, according to the matrix multiplication and addition operation requirement information obtained, parse out all necessary parameters, such as matrix size, data type, calculation form, etc. Set the instruction opcode flag according to the operation type. Set the corresponding operand address storage flag according to the addresses of matrices A, B, and C. If the operation contains an offset value, set the offset data storage flag. Set the calculation mode control information flag according to the specific needs of the calculation, such as selecting a specific optimization path or matrix multiplication type. Set the rounding mode enable flag according to whether a specific rounding mode is required. Combine the above-set fields into a complete instruction and output it.

[0034] Through the above technical solution, matrix operations can be directly supported at the hardware level through customized matrix multiplication and addition instructions, reducing the simulation and conversion overhead at the software level; the instructions include the selection of calculation mode and rounding mode, so that the hardware can be dynamically adjusted according to the calculation requirements, reducing unnecessary calculation steps and intermediate data storage, and improving the calculation efficiency; the instruction format supports operations of both integer and floating-point data types, meeting the needs of different application scenarios, and the introduction of bias values ​​enables the instruction to support more complex operation forms, such as biased addition in convolution operations; the configurability of the rounding mode ensures the accuracy and controllability of the calculation results, and is suitable for application scenarios with high precision requirements; by directly optimizing the core links of matrix operations, the overhead of data movement and storage is reduced, and the design of the instruction format fully considers the feasibility of hardware implementation, so that the hardware can perform matrix multiplication and addition operations in a more efficient manner.

[0035] Step S102, processing the matrix calculation instruction through the instruction pipeline to determine instruction parameter information.

[0036] The instruction parameter information includes a plurality of operation data, corresponding operand calculation mode parameters and offset data;

[0037] The matrix calculation instruction is processed through the instruction pipeline to determine the instruction parameter information, specifically including: analyzing the matrix calculation instruction through the instruction pipeline to determine the operand storage address and the bias data storage address corresponding to each operand; and obtaining the multiple operation data and the bias data according to the operand storage address and the bias data storage address.

[0038] In one embodiment of the present specification, the matrix calculation instruction is first analyzed using the instruction pipeline, including decoding the instruction operation code (opcode), identifying the instruction type (such as matrix multiplication, addition or multiplication and addition), and parsing the funct7 and funct3 fields to obtain the calculation mode and rounding mode information. By analyzing the rs1 and rs2 fields in the instruction, the instruction pipeline determines the storage address corresponding to each operand, and these addresses point to the memory location where the matrix data is stored. Similarly, by analyzing the rd field in the instruction, the instruction pipeline determines the storage address of the bias data. After determining the storage address of the operand and the bias data, the corresponding data is read from these addresses, including loading the data of matrices A and B from the memory, and loading the bias value. After obtaining all the necessary data, the instruction pipeline passes these data and the calculation mode parameters to the matrix calculation unit. The matrix calculation unit is configured according to the received calculation mode parameters to ensure that the calculation is performed in the correct way.

[0039] Through the above technical solution, the instruction pipeline efficiently processes the custom matrix calculation instruction to determine the instruction parameter information, and passes these data and calculation mode parameters to the matrix calculation unit; the design of the instruction pipeline enables the analysis, decoding and data acquisition of the instruction to be carried out in parallel, thereby shortening the execution cycle of the instruction, and obtaining the operand storage address and the bias data storage address by directly parsing the fields in the instruction, thereby avoiding additional memory access and calculation overhead; the custom matrix calculation instruction supports multiple calculation modes and rounding modes, and the instruction pipeline can flexibly configure these parameters by parsing the funct7 and funct3 fields to meet the needs of different application scenarios; after determining the storage address of the operand and bias data, the instruction pipeline directly loads these data from the memory, reducing the overhead of data movement and copying, and by optimizing the memory access strategy, such as using cache or prefetch technology, the efficiency of data loading can be further improved; the instruction pipeline directly passes the parsed data and calculation mode parameters to the matrix calculation unit, thereby avoiding the delay and overhead of the intermediate link; the matrix calculation unit is configured according to the received parameters, and can efficiently perform matrix operations, thereby improving the utilization and performance of the calculation unit.

[0040] Step S103, determining the corresponding matrix multiplication-addition calculation unit according to the operand calculation mode parameter through the pre-built matrix multiplication-addition module.

[0041] Wherein, the matrix multiplication and addition calculation unit includes any one of an integer matrix multiplication and addition calculation unit and a floating-point matrix multiplication and addition calculation unit;

[0042] According to the operand calculation mode parameter, a corresponding matrix multiplication and addition calculation unit is determined through a pre-constructed matrix multiplication and addition module, specifically including: according to the operand calculation mode parameter, the data type of the multiple operation data is identified, and the operation matrix data type is determined, wherein the operation matrix data type includes integer type data and floating point type data; when the operation matrix data type is integer type data, the matrix multiplication and addition calculation unit is determined to be an integer matrix multiplication and addition calculation unit; when the operation matrix data type is a floating point type, the matrix multiplication and addition calculation unit is determined to be a floating point matrix multiplication and addition calculation unit, and the rounding mode instruction in the matrix calculation instruction is sent to the floating point matrix multiplication and addition calculation unit.

[0043] In one embodiment of this specification, Figure 3 A schematic diagram of a calling flow of a matrix multiplication and addition module provided in an embodiment of this specification, such as Figure 3As shown in the figure, the architecture core of the matrix multiplication and addition module consists of a selection mechanism for rounding mode and calculation mode, and matrix multiplication and addition calculation units of different data types, aiming to efficiently perform matrix multiplication and addition operations. The selection mechanism for rounding mode and calculation mode allows the system to dynamically select appropriate rounding mode and calculation mode according to the parameters in the instruction, ensuring the accuracy of the operation results and meeting the needs of specific application scenarios. For matrix multiplication and addition calculation units of different data types, two types of matrix multiplication and addition calculation units, integer and floating-point, are equipped to support operations of different data types. After the matrix operation instruction is processed through the instruction pipeline, the system parses the rounding mode, calculation mode and operand data in the instruction. The calculation mode parameter is used to indicate the data type of the operand and the required calculation mode. For example, when the calculation mode parameter is 01, it means that the input operand is 32-bit integer type data; when the calculation mode parameter is 10, it means that the input operand is 32-bit floating point type data. The data type of the operand is identified according to the calculation mode parameter, and the appropriate matrix multiplication and addition calculation unit is selected after the data type is determined. When the operand is integer type data, the integer matrix multiplication and addition calculation unit is selected and activated to perform the corresponding integer matrix multiplication and addition operation. When the operand is a floating-point type data, the floating-point matrix multiplication and addition calculation unit is selected and activated. In addition, the system also sends the rounding mode instruction in the instruction to the module to ensure that the correct rounding rules can be applied when performing floating-point operations.

[0044] Through the above technical solution, by introducing the selection mechanism of calculation mode and rounding mode, it can be dynamically adjusted according to different calculation requirements and data types, thereby improving flexibility and enabling it to adapt to a variety of application scenarios; at the same time, since the design of the architecture core allows the addition of more types of calculation units and rounding modes, the solution has good scalability and can be upgraded and expanded with the development of technology and changes in demand; by directly selecting appropriate calculation units (integer matrix multiplication and addition calculation units or floating-point matrix multiplication and addition calculation units), unnecessary calculation steps and intermediate data storage are avoided, thereby improving calculation efficiency; applying correct rounding rules for floating-point operations can ensure the accuracy of the calculation results while reducing performance losses caused by rounding errors; by assigning different types of operations to specialized calculation units, hardware resources can be fully utilized and the utilization rate of calculation units can be improved.

[0045] Step S104, using a matrix multiplication and addition calculation unit to perform parallel calculations on a plurality of operation data to determine a calculation result.

[0046] Figure 4 A schematic diagram of the architecture of a matrix multiplication and addition computing unit provided in an embodiment of this specification, such as Figure 4As shown, the matrix multiplication and addition calculation unit includes a matrix allocation unit and multiple vector inner product calculation units, wherein the vector inner product calculation unit includes a multiplication unit and an addition unit, which work together to complete complex matrix operations.

[0047] The matrix multiplication and addition calculation unit is used to perform parallel calculation on the multiple operation data to determine the operation result, which specifically includes: performing matrix allocation on the multiple operation data through the matrix multiplication and addition calculation unit to determine the operation elements of each vector inner product calculation unit; performing parallel multiplication and addition operations on the operation elements of each vector inner product calculation unit to determine the output result of each vector inner product calculation unit; and determining the operation result with the output results of multiple vector inner product calculation units.

[0048] In one embodiment of the present specification, it is necessary to distribute multiple operation data in the form of a matrix, and divide the operation data into multiple sub-matrices for subsequent parallel processing. In this process, it is necessary to determine the operation elements of each vector inner product calculation unit, which is achieved by further dividing the sub-matrix into smaller matrix blocks. Each vector inner product calculation unit is responsible for processing one or more such vectors or matrix blocks. After determining the operation elements of each vector inner product calculation unit, parallel multiplication and addition operations are performed. Parallel multiplication and addition operations refer to performing multiplication and addition operations on one or more processors at the same time to speed up the calculation process. Each vector inner product calculation unit will perform parallel multiplication and addition operations on the operation elements it is responsible for. Specifically, for each vector inner product calculation unit, the corresponding elements in its operation elements are multiplied, and all products are added to obtain an output result. After determining the operation elements of each vector inner product calculation unit, parallel multiplication and addition operations are performed. The operation result is determined based on the output results of multiple vector inner product calculation units.

[0049] Through the matrix multiplication and addition calculation unit, the multiple operation data are matrix-allocated to determine the operation elements of each vector inner product calculation unit, specifically including: dividing each operation data according to rows and columns, determining at least one operation element corresponding to each operation data, wherein the operation element includes row elements and column elements; determining row and column parameters of the bias data according to the row identifier of the row element and the column identifier of the column element corresponding to the operation element, so as to determine the bias parameter in the bias data based on the row and column parameters; determining the element row and column identifier of the operation element corresponding to each operation data, and determining the operation element of each vector inner product calculation unit and the element position of each operation element in the corresponding operation data according to the element row and column identifier and a preset allocation rule.

[0050] In one embodiment of the present specification, the operands A, B, and C matrices of the input matrix multiplication and addition calculation unit are all of 4x4 type. First, the data is allocated by rows and columns through the matrix allocation unit. For example, the first row data of the A matrix, the first column data of the B matrix, and the data located at the first row and first column of the C matrix are sent to the vector inner product calculation unit 1, and the second row data of the A matrix, the second column data of the B matrix, and the data located at the second row and second column of the C matrix are sent to the vector inner product calculation unit 2. The remaining data is pushed into the corresponding vector inner product calculation units in this way.

[0051] Through the above technical solution, by dividing the data into rows and columns and allocating them to multiple vector inner product calculation units, the advantages of parallel computing can be fully utilized. Each calculation unit independently processes a part of the data, which significantly improves the calculation efficiency. In particular, when processing large-scale matrix operations, this parallel processing method can greatly shorten the calculation time; according to the row and column identification of the operation data and the preset allocation rules, the data is accurately allocated to each calculation unit, ensuring the effective use of computing resources; it is not only suitable for matrices of a specific size (such as a 4x4 matrix), but also can adapt to matrix operations of different sizes by adjusting the allocation rules; by decomposing complex matrix multiplication and addition calculations into multiple simple vector inner product calculations, the overall calculation complexity is reduced.

[0052] Performing parallel multiplication and addition operations on the operation elements of each of the vector inner product calculation units to determine the output result of each of the vector inner product calculation units, specifically comprising: in each of the vector inner product calculation units, performing multiplication operations on the operation elements at designated positions according to the element positions of each of the operation elements in the corresponding operation data to obtain a plurality of multiplication results; and performing accumulation operations on the plurality of multiplication results and the bias parameter to determine the output result of the vector inner product calculation unit.

[0053] In one embodiment of the present specification, the operation A, B, and C matrices of the input matrix multiplication and addition calculation unit are all 4x4 matrices. The specific implementation process of the vector inner product calculation unit is as shown in the attached figure. Figure 4As shown, first, the multiplication operation at the corresponding position is performed in parallel through its four built-in multiplication units, that is, the first row elements (A11, A12, A13, A14) of the A matrix are multiplied with the first column elements (B11, B21, B31, B41) of the B matrix respectively, to obtain four multiplication results Mul_r1, Mul_r2, Mul_r3, and Mul_r4, and then the above four results Mul_r1, Mul_r2, Mul_r3, and Mul_r4 are input to the accumulation unit for further processing. The implementation of the accumulation unit adopts a multi-level addition strategy to optimize the calculation process and reduce delay. At the same time, the addition of Mul_r1 and Mul_r2, and the addition of Mul_r3 and Mul_r4 are calculated to obtain the calculation results Add_r1 and Add_r2. Then, the addition of Add_r1 and Add_r2 is calculated to obtain the final accumulation result Add_r. Finally, an addition unit is used to add the accumulated result Add_r and the bias value C11 to obtain the final output result Finall_out1 of the vector inner product calculation unit 1. The implementation process of other vector inner product calculation units is the same as that of the vector inner product calculation unit 1, and all vector inner product calculation units use parallel calculation to obtain the calculation result of the matrix multiplication and addition module.

[0054] Through the above technical solution, by building multiple multiplication units into each vector inner product calculation unit, the parallel execution of multiplication operations at corresponding positions is realized, which can greatly improve the calculation speed, especially when processing large-scale matrix operations, the calculation time can be significantly shortened; the accumulation unit adopts a multi-level addition strategy to further optimize the calculation process, and reduces the delay in the calculation process and improves the calculation efficiency by grouping the multiplication results for addition operations and then gradually merging the results; because each vector inner product calculation unit works independently and uses precise multiplication and addition operations, the calculation accuracy is guaranteed.

[0055] After the matrix multiplication and addition calculation unit is used to perform parallel calculation on the multiple operation data and the calculation result is determined, the method further includes: determining the storage location of the calculation result according to the bias data storage identification bit in the matrix calculation instruction, so as to store the calculation result at the storage location.

[0056] In one embodiment of the present specification, after obtaining the result of the matrix multiplication and addition operation, the biased data storage identification bit in the matrix calculation instruction is obtained, the storage location of the operation result is determined, that is, the address pointed to by rd is corresponding, and the operation result of the matrix multiplication and addition operation is stored in this address. By introducing the biased data storage identification bit, the storage location of the operation result can be flexibly specified, and the storage method of the operation result can be dynamically adjusted according to different calculation requirements and data management strategies, thereby improving the efficiency and flexibility of data management; by accurately controlling the storage location of the operation result, memory resources can be used more effectively, for example, the operation result can be directly stored in the location required for subsequent calculations, avoiding unnecessary data movement and copying, thereby reducing memory usage and improving memory access speed; in complex calculation tasks, the operation result may need to be used as the input of subsequent calculations. By specifying the storage location of the operation result, the continuity and consistency of the data flow can be easily achieved, supporting more complex calculation processes and data processing logic.

[0057] Through the above technical solution, compared with traditional processors, the instruction set of the RISC-V architecture is more flexible, which allows the design of more complex and specialized instruction formats to include the complete information required for matrix multiplication and addition operations, reducing the number of instructions that need to be transmitted during the calculation process, reducing the complexity of instruction decoding, and reducing the number of data movements, thereby improving calculation efficiency; when traditional processors process matrix operations, due to the limitations of the instruction set, they often need to move data frequently to perform different calculation steps, while the matrix multiplication and addition operation implementation method based on RISC-V can reduce unnecessary data movement, reduce storage overhead, and improve data processing efficiency by optimizing the instruction format and calculation process; traditional processors are in parallel There are limitations in processing, and it is difficult to fully utilize the multi-core and multi-threading technology of modern processors to accelerate matrix operations. The matrix multiplication and addition operation implementation method based on RISC-V can efficiently perform parallel calculations by building special matrix multiplication and addition modules and computing units, and fully utilize the advantages of multi-core and multi-threading technology to significantly improve the calculation speed and efficiency; the matrix addition operation implementation method based on RISC-V can support complex computing modes and efficiently process large-scale data by providing multiple computing modes and optimized computing units, thereby meeting the needs of these key applications; by optimizing the instruction format, computing process and parallel processing strategy, it can reduce power consumption and improve energy efficiency to meet the application requirements of long-term operation and high efficiency.

[0058] The embodiments of this specification also provide a matrix multiplication and addition operation implementation device based on RISC-V, such as Figure 5As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0059] The embodiments of the present specification also provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.

[0060] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0061] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0062] The devices and media provided in the embodiments of this specification correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0063] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0064] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0065] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0067] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0068] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0069] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0070] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0071] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.

Claims

1. A method for implementing matrix multiplication and addition operations based on RISC-V, characterized in that: The method comprises: Obtaining matrix multiplication and addition operation requirement information, and generating corresponding matrix calculation instructions based on the matrix multiplication and addition operation requirement information and in accordance with a preset instruction format, wherein the matrix calculation instruction includes an operand address and calculation mode information; Processing the matrix calculation instruction through an instruction pipeline to determine instruction parameter information, wherein the instruction parameter information includes a plurality of operation data, corresponding operand calculation mode parameters and bias data; According to the operand calculation mode parameter, a corresponding matrix multiplication-addition calculation unit is determined by a pre-constructed matrix multiplication-addition module, wherein the matrix multiplication-addition calculation unit includes any one of an integer matrix multiplication-addition calculation unit and a floating-point matrix multiplication-addition calculation unit; The matrix multiplication and addition calculation unit is used to perform parallel calculation on the plurality of operation data to determine the calculation result.

2. The method for implementing matrix multiplication and addition operations based on RISC-V according to claim 1, characterized in that: The instruction format includes an instruction operation code identification bit, a plurality of operand address storage identification bits, a biased data storage identification bit, a calculation mode control information identification bit and a rounding mode enabling identification bit.

3. The method for implementing matrix multiplication and addition operations based on RISC-V according to claim 1, characterized in that: The matrix calculation instruction is processed through the instruction pipeline to determine instruction parameter information, specifically including: Analyzing the matrix calculation instruction through the instruction pipeline to determine the operand storage address and the offset data storage address corresponding to each operand; The plurality of operation data and the bias data are acquired according to the operand storage address and the bias data storage address.

4. The method for implementing matrix multiplication and addition operations based on RISC-V according to claim 1, characterized in that: According to the operand calculation mode parameters, the corresponding matrix multiplication and addition calculation unit is determined through the pre-built matrix multiplication and addition module, specifically including: According to the operand calculation mode parameter, the data types of the plurality of operation data are identified to determine the operation matrix data type, wherein the operation matrix data type includes integer type data and floating point type data; When the operation matrix data type is integer type data, determining that the matrix multiplication and addition calculation unit is an integer matrix multiplication and addition calculation unit; When the operation matrix data type is a floating-point type, the matrix multiplication-addition calculation unit is determined to be the floating-point matrix multiplication-addition calculation unit, and the rounding mode instruction in the matrix calculation instruction is sent to the floating-point matrix multiplication-addition calculation unit.

5. The method for implementing matrix multiplication and addition operations based on RISC-V according to claim 1, characterized in that: The matrix multiplication and addition calculation unit is used to perform parallel calculation on the plurality of operation data to determine the calculation result, specifically including: By means of the matrix multiplication and addition calculation unit, the plurality of operation data are matrix-allocated to determine the operation element of each vector inner product calculation unit; Performing parallel multiplication and addition operations on the operation elements of each of the vector inner product calculation units to determine the output result of each of the vector inner product calculation units; The operation result is determined based on the output results of the plurality of vector inner product calculation units.

6. The method for implementing matrix multiplication and addition operations based on RISC-V according to claim 5, characterized in that: By means of the matrix multiplication and addition calculation unit, the plurality of operation data are matrix-allocated to determine the operation element of each vector inner product calculation unit, specifically including: Divide each of the operation data into rows and columns, and determine at least one operation element corresponding to each of the operation data, wherein the operation element includes a row element and a column element; Determine row and column parameters of the bias data according to the row identifier of the row element and the column identifier of the column element corresponding to the operation element, so as to determine the bias parameter in the bias data based on the row and column parameters; Determine the element row and column identifiers of the operation elements corresponding to each operation data, and determine the operation elements of each vector inner product calculation unit and the element position of each operation element in the corresponding operation data according to the element row and column identifiers and a preset allocation rule.

7. The method for implementing matrix multiplication and addition operations based on RISC-V according to claim 6, characterized in that: Performing parallel multiplication and addition operations on the operation elements of each of the vector inner product calculation units to determine the output result of each of the vector inner product calculation units specifically includes: In each of the vector inner product calculation units, according to the element position of each of the operation elements in the corresponding operation data, a multiplication operation is performed on the operation element at the specified position to obtain a plurality of multiplication results; Accumulate the multiple multiplication results and the bias parameter to determine an output result of the vector inner product calculation unit.

8. The method for implementing matrix multiplication and addition operations based on RISC-V according to claim 2, characterized in that: After performing parallel calculations on the plurality of operation data using the matrix multiplication and addition calculation unit and determining the calculation results, the method further includes: According to the offset data storage identification bit in the matrix calculation instruction, the storage location of the operation result is determined to store the operation result at the storage location.

9. A matrix multiplication and addition operation implementation device based on RISC-V, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Matrix vector multiplication optimization method based on RISC-V platform

    CN120406897A

  • Matrix-vector multiplication optimization method based on RISC-V platform

    CN120406897B

  • Matrix allocation and calculation method and system for multiply-add unit and storage medium

    CN121478225A

  • Matrix allocation and computation method, system and storage medium for multiply-add units

    CN121478225B