Chip storage and indexing method and system based on matrix data and electronic equipment

By adopting chip storage and indexing methods based on matrix data in the Transformer model, the problems of low computing efficiency and frequent memory access in the decoding stage are solved, and more efficient computing resource utilization and storage space management are achieved.

CN120179607AActive Publication Date: 2025-06-20SHENZHEN ZHONGWEIDIAN TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510661250.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The Transformer model has low computational efficiency in the decoding stage, resulting in frequent memory access and increasing bus bandwidth pressure. The existing k/v cache mechanism is limited by on-chip storage capacity, which cannot effectively solve this problem.

Method used

The chip storage and indexing method based on matrix data is adopted. By setting the chip storage structure and indexing method, the rows or columns of data are stored first and the shapes and indexing methods of sub-matrixes are set to reduce memory access frequency and storage space requirements.

Benefits of technology

It improves data multiplexing rate, reduces memory access frequency, reduces bus bandwidth pressure, realizes more efficient computing resource utilization, and reduces space waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179607A_ABST
    Figure CN120179607A_ABST
Patent Text Reader

Abstract

The invention discloses a chip storage and indexing method and system based on matrix data and electronic equipment, and the method comprises the following steps: setting a chip storage structure: continuously storing on-chip data according to a row or column priority mode, continuously or discontinuously storing inter-chip data, numbering chip indexes according to a row priority or column priority mode, setting the shape of the sub-matrix to be a starting piece ID, an ending piece ID, a starting row or column ID, an ending row or column ID and the number of pieces contained in each row or column; an access instruction describes the length and width dimensions of a slice, data are accessed through the index modes, the index modes comprise a row index mode and a column index mode, and the row index mode comprises a slice ID, a row start ID and a row end ID; the column index mode comprises a slice ID, a column start ID and a column end ID; and the calculation reference of the piece ID takes the initial piece ID of the mother matrix as an original point for positioning. The method can improve the reuse rate of the data and reduce the access frequency of the memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data storage, and particularly relates to a chip storage and indexing method, system and electronic device based on matrix data. Background Art

[0002] The efficient computation of the Transformer model highly depends on large-scale matrix multiplication operations, which poses strict requirements on real-time inference performance and leads to frequent memory accesses and huge bandwidth pressure. In the existing NVIDIA GPU architecture, the utilization rate of computing units also remains at a low level, making it difficult to fully utilize computing resources.

[0003] In the Ampere architecture, NVIDIA introduced the mma instruction to improve matrix operation efficiency by loading batch data into SRAM for computation. However, this operation mode reduces the computing efficiency due to insufficient reuse rate of a large amount of data during the decoding stage, and may even increase the bus bandwidth pressure.

[0004] Specifically, the computing process of the Transformer mainly includes two stages: prefill and decode. In the prefill stage, multiple input data have significant parallel computing potential, and a large amount of data can be reused, thus effectively reducing the bus pressure. In contrast, in the decode stage, only a single input is processed each time, the parallel computing ability is significantly limited, and frequent memory accesses greatly increase the burden on the bus. In addition, the existing k / v cache mechanism is also limited by the on-chip storage capacity, further exacerbating this problem.

[0005] To address the above challenges, NVIDIA introduced a dedicated Transformer engine in the Hopper architecture, combined with the fourth-generation Tensor Core technology, supporting FP8 precision, which greatly improves the training and inference performance of large language models. Although increasing the on-chip storage capacity and bandwidth and increasing the number of parallel threads can significantly enhance the high-density computing performance, there is still much room for improvement in optimizing data storage and invocation. Summary of the Invention

[0006] The present invention addresses the above problems and provides a chip storage and indexing method, system and electronic device based on matrix data, which improves the data reuse rate and reduces the memory access frequency.

[0007] According to the first aspect of the embodiments of the present disclosure, a chip storage and indexing method based on matrix data is provided. The method includes: Set up a chip - based storage structure: The data within a chip is stored continuously in a row - major or column - major manner, the data between chips can be stored continuously or discontinuously, the numbering method of the chip index is row - first or column - first, and the shape of the sub - matrix is set as the starting chip ID, ending chip ID, starting row or column ID, ending row or column ID, and the number of chips contained in each row or column; Set up the indexing method: The access instruction describes the length and width dimensions of the chip, and accesses the data through the indexing method. The indexing method includes a row - indexing method and a column - indexing method. The row - indexing method includes the chip ID, starting row ID, and ending row ID; the column - indexing method includes the chip ID, starting column ID, and ending column ID; the calculation reference of the chip ID is located with the starting chip ID of the mother matrix as the origin.

[0008] In some embodiments, different chip shapes can be set for different regions of the mother matrix.

[0009] In some embodiments, the number of chips contained in each row or column refers to the number of chips that the sub - matrix can accommodate in a single row or column. When the size of the mother matrix is insufficient, the extra regions are automatically filled with zeros.

[0010] In some embodiments, each sub - matrix and / or the mother matrix storage corresponds to a physical address mapping table of the chip ID. When performing matrix splicing, the physical storage location of the original data is not changed. Based on the physical address mapping table of the mother matrix, the target physical address corresponding to the chip ID and row ID is located. If the target chip ID already exists in the chip index table, the corresponding physical address is directly accessed; if not, a continuous blank physical address is dynamically applied for and registered as a new chip ID, and the data is written into the corresponding storage space to achieve logical matrix splicing.

[0011] In some embodiments, the physical address mapping table also marks whether the matrix data is transposed. If it is transposed, the ID number of the transposed table is used, and the extracted chip attributes also mark the transposed attributes. When performing the transpose marking, the storage data position is not changed. When using, the table of the mother matrix is marked, and the mapping relationship of the table is changed. The transposed chip ID number is updated by the dimension information of the mother matrix.

[0012] In some embodiments, when the matrix data is moved to a higher - level storage structure, through a dedicated thread, the data moved to the shared memory is transposed and rearranged.

[0013] In some embodiments, an index table and a storage mapping table are created based on the matrix data that needs to be stored after calculation. The matrix data is stored in the main memory according to the index table, and the matrix data is registered according to the storage mapping table and the main memory data is transmitted through the PCIe interface.

[0014] According to the second aspect of the embodiments of the present disclosure, a chip - based storage and indexing system based on matrix data is provided. The system includes: Set a chip - type storage structure module, which is used to set the in - chip data to be continuously stored in a row - major or column - major manner. The numbering method of the chip index is row - major or column - major. Set the shape of the sub - matrix as the starting chip ID, ending chip ID, starting row or column ID, ending row or column ID, and the number of chips contained in each row or column. Set an indexing method module, which is used to set the access instruction to describe the length and width dimensions of the chip and access the data through the indexing method. The indexing method includes a row - indexing method and a column - indexing method. The row - indexing method includes the chip ID, row starting ID, and row ending ID; the column - indexing method includes the chip ID, column starting ID, and column ending ID. The calculation reference of the chip ID is positioned with the starting chip ID of the mother matrix as the origin.

[0015] According to the third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above - mentioned chip - type storage and indexing method based on matrix data are implemented.

[0016] According to the fourth aspect of the embodiments of the present disclosure, there is provided a non - temporary computer - readable storage medium, on which computer instructions are stored. When the instructions are executed by the processor, the steps of the above - mentioned chip - type storage and indexing method based on matrix data are implemented.

[0017] A chip - type storage and indexing method, system, electronic device, and storage medium based on matrix data provided by the embodiments of the present disclosure set a chip - type storage structure: the in - chip data is continuously stored in a row - major or column - major manner, the numbering method of the chip index is row - major or column - major, and the shape of the sub - matrix is set as the starting chip ID, ending chip ID, starting row or column ID, ending row or column ID, and the number of chips contained in each row or column. Set the indexing method: the access instruction describes the length and width dimensions of the chip and accesses the data through the indexing method. The indexing method includes a row - indexing method and a column - indexing method. The row - indexing method includes the chip ID, row starting ID, and row ending ID; the column - indexing method includes the chip ID, column starting ID, and column ending ID. The calculation reference of the chip ID is positioned with the starting chip ID of the mother matrix as the origin. Specifically, the beneficial effects of the present invention include: 1. Distributed storage, piecemeal storage and whole - extraction: In terms of storage space, there is no need to store the data completely, reducing the demand for storage space. According to the characteristics of matrix multiplication, it is convenient to extract the entire data to be calculated completely, reducing the number of memory accesses and the probability of loss of calculation data. 2. Flexible operation and reduced space waste: By using the concept of sub-matrices, it is possible to flexibly perform operations on the local data of large matrices. When transposing, there is no need to change the storage order, and the data can be transposed by changing the slice ID number of the table. The operation granularity is a vector with the number of rows in a slice as the size. When splicing, there is no need to apply for a large continuous storage space. Storing by slices effectively reduces space waste.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0019] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0020] Figure 1 It is a schematic diagram of the method for storing data by slices and indexing sub-matrices in the mother matrix in an embodiment of the present invention; Figure 2 It is a schematic diagram of the method for matrix transpose reference and sub-matrix splicing in an embodiment of the present invention; Figure 3 It is a schematic diagram of the method for mapping data storage in hardware in an embodiment of the present invention; Figure 4 It is a schematic diagram of the method for matrix multiplication configuration in the decoding stage of Transformer in an embodiment of the present invention; Figure 5 It is a schematic diagram of the method for storing the key-value matrix back in Transformer in an embodiment of the present invention; Figure 6 It is a schematic diagram of the system structure of the slice-based storage and indexing system for matrix data in an embodiment of the present invention; Figure 7 It is a schematic diagram of an electronic device in an embodiment of the present invention. Detailed Embodiments

[0021] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, rather than all the structures.

[0022] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0023] Embodiments of the present invention are directed to a method, system, electronic device, and storage medium for piecewise storage and indexing based on matrix data, and provide the following embodiments: A method for piecewise storage and indexing based on matrix data, the method comprising: Setting a piecewise storage structure: the data within a piece is continuously stored in a row-major or column-major manner, the data between pieces is stored continuously or discontinuously, the numbering method of the piece index is row-major or column-major, and the shape of the sub-matrix is set as the starting piece ID, ending piece ID, starting row or column ID, ending row or column ID, and the number of pieces contained in each row or column; Setting an indexing method: the access instruction describes the length and width dimensions of the piece, and the data is accessed through the indexing method, where the indexing method includes a row indexing method and a column indexing method. The row indexing method includes the piece ID, row starting ID, and row ending ID; the column indexing method includes the piece ID, column starting ID, and column ending ID; the calculation reference of the piece ID is located with the starting piece ID of the parent matrix as the origin.

[0024] Preferably, different piece shapes can be set for different regions of the parent matrix.

[0025] Preferably, the number of pieces contained in each row or column refers to the number of pieces that can be accommodated by the sub-matrix in a single row or column. When the size of the parent matrix is insufficient, the extra regions are automatically filled with zeros.

[0026] Preferably, each sub-matrix and / or the storage of the parent matrix corresponds to a physical address mapping table of the piece ID.

[0027] Preferably, the physical address mapping table also indicates whether the matrix data is transposed. If it is transposed, the ID number of the transposed table is used, and the extracted piece attributes also indicate the transposed attributes.

[0028] During the specific implementation process, such as Figure 1As shown, a matrix-based chip storage structure. In physical storage, each piece of data needs to be continuously stored in row-major order within a chip, while data between chips does not need to be continuously stored. The numbering method of chip indexes supports two modes: row-major and column-major, and can be flexibly selected according to computing requirements. Each instruction needs to clearly describe the length and width dimensions of the chip, and access data through indexing methods (chip ID, starting row ID, ending row ID). Among them, the calculation benchmark of the chip ID is located with the starting chip ID of the parent matrix as the origin. As Figure 1 shown, the legend demonstrates the row access method. In actual applications, it is also possible to choose block access by column. If not specifically specified, the default starting row ID is the first row, and the ending row ID is the last row. The minimum data access granularity of this scheme is a row vector within a chip; Furthermore, it is allowed to dynamically and flexibly set the shape of the chip, and different chip shapes can be set in different regions of the parent matrix. However, for the convenience of hardware mapping, the dimensions of the chip should match the hardware resources (such as register capacity, bus bandwidth); Furthermore, the shape of each sub-matrix is described by the following parameters: (starting chip ID, ending chip ID, starting row ID, ending row ID, the number of chips contained in each row). Among them, the number of chips per row refers to the number of chips that can be accommodated in a single row of the sub-matrix. If the size of the parent matrix is insufficient, the extra area is automatically filled with zeros to meet subsequent computing requirements; Furthermore, as Figure 1 shown, all matrix storages have a physical address mapping table based on chips. In addition to corresponding the virtual chip IDs, this table also marks whether the table being managed is transposed. If it is transposed, the ID number of the transposed table is used, and the extracted chip attributes will also be marked with transposed attributes.

[0029] Preferably, when matrix data is moved to a higher-level storage structure, through a dedicated thread, the data moved to the shared memory is transposed and rearranged.

[0030] In the specific implementation process, as Figure 2 shown, when performing transpose marking, the stored data position is not changed at all. When using, set the table flag position of the corresponding parent matrix to 1, and change the mapping relationship of the table to column-major (in parentheses). The transposed chip ID number can be updated according to the dimension information of the parent matrix. For example Figure 2 The column-major chip ID number data in the parentheses in the table is , where n is the row-major chip ID number, n* is the column-major chip ID number, and floor() represents rounding down.

[0031] Further, after the data is moved to a higher-level storage structure, efficient transposition support is required at the hardware level. For example, in a CUDA environment, dedicated threads can be used to transpose and rearrange the data moved to the shared memory so that subsequent calculations can efficiently perform transposed access; As Figure 2 shown, the concept of sub-matrices is introduced to achieve matrix splicing. In actual application scenarios, the sizes of the matrices to be spliced usually cannot be determined in advance. Therefore, the shape of the sub-matrices is described based on the mother matrix, specifically defined as (starting slice ID, ending slice ID, number of row slices of the sub-matrix), and is usually indexed and managed in row-major order. In actual storage operations, since the data between slices does not need to be continuously stored, when calculating the spliced matrix, there is no need to apply for a complete space for the entire sub-matrix in advance, and only need to apply for space for each slice separately.

[0032] Preferably, an index table and a storage mapping table are created based on the matrix data that needs to be stored after the calculation is completed. The matrix data is stored in the main memory according to the index table, and the matrix data is registered according to the storage mapping table and the main memory data is transmitted through the PCIe interface.

[0033] In the specific implementation process, as Figure 3 shown, when performing mapping, the index granularity of the data can be less than or equal to one slice, but the granularity size needs to be an integer multiple of the length of a single row within the slice. After the calculation is completed, an index table of the result will be created, and the calculation result will be stored back in the main memory according to this table. It should be noted that the main memory data is usually transmitted through the PCIe interface. Before transmitting the data, a corresponding storage mapping table needs to be registered in advance for the matrix to be transmitted.

[0034] In a specific embodiment, as Figure 4 shown, it is the data index of GPT2.0-345M in the decoding stage. In this table, the data of the left-multiplied matrix is stored continuously by row (1×1024), and the storage dimension of the data of the right-multiplied matrix is 32×32. With this configuration, when the bandwidth is 2 Kbyte (for example, the bandwidth of 32 channels and 512 Bit), this scheme can extract the data of the left- and right-multiplied matrices to the on-chip storage in two requests, so as to provide 64 times of matrix multiplication data for the computing unit (such as a common 4×4 convolution unit), and in the subsequent 1024 / 32 = 32 processes, the data extraction of the left matrix will not be lost, and only the right-multiplied matrix needs to be continuously moved.

[0035] In another specific embodiment, as Figure 5As shown in the figure, the storage matrix configuration of the key value when it is stored. In the general process, the storage space that can store the entire key value is often applied for in the space at one time according to the model information. However, the storage of the key value in the pre-filling stage does not need to fill the entire space, so a large amount of space is wasted and cannot be filled. Due to the existence of the lookup table, that is, the physical address mapping table, the scheme can only apply for the storage space of the next result (minimum is a slice) without applying for the storage space of the mother matrix (that is, the key value matrix). By changing the slice size to the size of the result at one time, the waste of storage space can be minimized. In terms of matrix splicing operations, this method does not change the physical storage location of the original data. Based on the lookup table of the mother matrix, locate the target physical address corresponding to the slice ID and row ID. If the target slice ID already exists in the slice index table, the corresponding physical address is directly accessed; if it does not exist, a continuous blank physical address is dynamically applied for and registered as a new slice ID. The data is written into the corresponding storage space to achieve logical matrix splicing. This method effectively reduces the waste caused by applying for storage space of the entire mother matrix size at one time in multi-head splicing and key-value caching operations in the Transformer model, while avoiding the challenges brought by applying for larger continuous storage space.

[0036] Another embodiment is used to illustrate a slice storage and indexing system based on matrix data, such as Figure 6 As shown, the system 600 includes: A slice storage structure module 610 is set to set the slice data to be stored continuously in a row-first or column-first manner, the slice index numbering method is row-first or column-first, and the shape of the submatrix is ​​set to be a starting slice ID, an ending slice ID, a starting row or column ID, an ending row or column ID, and the number of slices contained in each row or column; An index mode module 620 is set to set the length and width dimensions of the access instruction description slice, and access data through an index mode, wherein the index mode includes a row index mode and a column index mode, the row index mode includes a slice ID, a row start ID, and a row end ID; the column index mode includes a slice ID, a column start ID, and a column end ID; the calculation basis of the slice ID is positioned with the start slice ID of the mother matrix as the origin.

[0037] In addition to the above modules, the system 600 may also include other components. However, since these components are irrelevant to the content of the embodiment of the present disclosure, their illustration and description are omitted here.

[0038] For other specific working processes of the matrix data-based slice storage and indexing system 600, refer to the description of the above-mentioned matrix data-based slice storage and indexing method embodiment, which will not be repeated here.

[0039] Another embodiment is used to illustrate that the system of the present invention can also be used with the help of Figure 7implemented by the architecture of the computing device shown. Figure 7 The architecture of the computing device is shown. As Figure 7 shown, a computer system 710, a system bus 730, one or more CPUs 740, an input / output 720, a memory 750, etc. The memory 750 may store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU, including the program instructions for implementing the method of slice storage and indexing based on matrix data in the embodiments. Figure 7 The architecture shown is only exemplary. When implementing different devices, one or more components in Figure 7 are adjusted according to actual needs. The memory 750, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the method of slice storage and indexing based on matrix data in the embodiments of the present invention (for example, the set slice storage structure module 610 and the set indexing method module 620 in the slice storage and indexing system 600 based on matrix data). One or more CPUs 740 execute various functional applications and data processing of the system of the present invention by running the software programs, instructions, and modules stored in the memory 750, that is, implement the above-mentioned method of slice storage and indexing based on matrix data. The method includes: Setting the slice storage structure: The data within the slice is continuously stored in a row-major or column-major manner, the data between slices is stored continuously or discontinuously, the numbering method of the slice index is row-major or column-major, and the shape of the sub-matrix is set as the starting slice ID, the ending slice ID, the starting row or column ID, the ending row or column ID, and the number of slices contained in each row or column; Setting the indexing method: The access instruction describes the length and width dimensions of the slice and accesses the data through the indexing method. The indexing method includes a row indexing method and a column indexing method. The row indexing method includes the slice ID, the starting row ID, and the ending row ID; the column indexing method includes the slice ID, the starting column ID, and the ending column ID; the calculation benchmark of the slice ID is located with the starting slice ID of the parent matrix as the origin.

[0040] Of course, for the server provided in the embodiments of the present invention, its processor is not limited to executing the method operations as described above, and can also execute the related operations in the method of slice storage and indexing based on matrix data provided in any embodiment of the present invention.

[0041] The memory 750 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 750 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 750 may further include a memory remotely provided with respect to one or more CPUs 740, and these remote memories may be connected to the system through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0042] The input / output 720 may be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the system. The input / output 720 may also include a display device such as a display screen.

[0043] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for chip storage and indexing based on matrix data described in the above embodiment. The computer-readable storage medium of the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.

[0044] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium may send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device.

[0045] The program code contained on the storage medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0046] In addition, for the other specific working processes of a non-temporary computer-readable storage medium, reference may be made to the description of the above-described embodiments of the chip storage and indexing method based on matrix data, and details will not be repeated here.

[0047] According to the above embodiments, a chip storage and indexing method, system, electronic device, and storage medium based on matrix data provided by the present disclosure set up a chip storage structure: the data within the chip is continuously stored in a row-major or column-major manner, the numbering method of the chip index is row-major or column-major, and the shape of the sub-matrix is set as the starting chip ID, ending chip ID, starting row or column ID, ending row or column ID, and the number of chips contained in each row or column; an indexing method is set: the access instruction describes the length and width dimensions of the chip, and data is accessed through the indexing method, where the indexing method includes a row indexing method and a column indexing method, the row indexing method includes the chip ID, row starting ID, and row ending ID; the column indexing method includes the chip ID, column starting ID, and column ending ID; the calculation reference of the chip ID is positioned with the starting chip ID of the mother matrix as the origin. Specifically, the beneficial effects of the present invention include: 1. Distributed storage, piecemeal storage and whole retrieval: There is no need to store data completely in the storage space, reducing the demand for storage space. According to the characteristics of matrix multiplication, it is convenient to extract the entire data that needs to be operated, reducing the number of memory accesses and the probability of loss of calculation data; 2. Flexible operation, reducing space waste: Using the concept of sub-matrices, it is possible to flexibly complete operations on local data of large matrices; when transposing, there is no need to change the storage order, and only by changing the chip ID number of the table, the data can be transposed; the operation granularity is a vector with the number of rows in the chip as the size; when splicing, there is no need to apply for a huge continuous storage space, and storing by chip effectively reduces space waste.

[0048] In this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a step or method including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such step or method.

[0049] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as falling within the protection scope of the present invention.

Claims

1. A chip storage and indexing method based on matrix data, characterized in that, The method includes: Setting up a chip - based storage structure: The data within a chip is continuously stored in a row - major or column - major manner. The data between chips can be stored continuously or discontinuously. The numbering method of the chip index is row - major or column - major. Set the sub - matrix description as the starting chip ID, ending chip ID, starting row or column ID, ending row or column ID, and the number of chips contained in each row or column. Setting up the indexing method: The access instruction describes the length and width dimensions of the chip and accesses the data through the indexing method. The indexing method includes a row - indexing method and a column - indexing method. The row - indexing method includes the chip ID, starting row ID, and ending row ID. The column - indexing method includes the chip ID, starting column ID, and ending column ID. The calculation reference of the chip ID is located with the starting chip ID of the mother matrix as the origin.

2. The chip storage and indexing method based on matrix data according to claim 1, characterized in that, Different regions of the mother matrix can be set with different chip shapes.

3. The chip storage and indexing method based on matrix data according to claim 1, characterized in that, The number of chips contained in each row or column refers to the number of chips that can be accommodated in a single row or column of the sub - matrix. When the size of the mother matrix is insufficient, the extra regions are automatically filled with zeros.

4. The chip storage and indexing method based on matrix data according to claim 1, characterized in that, Each sub - matrix and / or mother matrix storage corresponds to a physical address mapping table of the chip ID. When matrix splicing is performed, the physical storage location of the original data is not changed. Based on the physical address mapping table of the mother matrix, the target physical address corresponding to the chip ID and row ID is located. If the target chip ID already exists in the chip index table, directly access the corresponding physical address. If not, dynamically apply for a continuous blank physical address, register it as a new chip ID, and write the data into the corresponding storage space to achieve logical matrix splicing.

5. The chip storage and indexing method based on matrix data according to claim 4, characterized in that, The physical address mapping table also marks whether the matrix data is transposed. If it is transposed, the ID number of the transposed table is used, and the extracted chip attributes also mark the transposed attributes. When transposition is marked, the storage data location is not changed. When used, the table of the mother matrix is marked, and the mapping relationship of the table is changed. The transposed chip ID number is updated according to the dimension information of the mother matrix.

6. The chip storage and indexing method based on matrix data according to claim 5, characterized in that, When the matrix data is moved to a higher - level storage structure, through a dedicated thread, the data moved to the shared memory is transposed and rearranged.

7. The chip storage and indexing method based on matrix data according to claim 1, characterized in that, Create an index table and a storage mapping table based on the matrix data to be stored after calculation. Store the matrix data in the main memory according to the index table, register the matrix data according to the storage mapping table, and transmit the main - memory data through the PCIe interface.

8. A chip storage and indexing system based on matrix data, characterized in that, The system includes: A module for setting up the chip - based storage structure, which is used to set the data within a chip to be continuously stored in a row - major or column - major manner, the numbering method of the chip index to be row - major or column - major, and the shape of the sub - matrix to be the starting chip ID, ending chip ID, starting row or column ID, ending row or column ID, and the number of chips contained in each row or column. A module for setting up the indexing method, which is used to set the access instruction to describe the length and width dimensions of the chip and access the data through the indexing method. The indexing method includes a row - indexing method and a column - indexing method. The row - indexing method includes the chip ID, starting row ID, and ending row ID. The column - indexing method includes the chip ID, starting column ID, and ending column ID. The calculation reference of the chip ID is located with the starting chip ID of the mother matrix as the origin.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the chip - based storage and indexing method based on matrix data according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, on which computer instructions are stored, characterized in that, When the instruction is executed by a processor, the steps of the chip storage and indexing method based on matrix data according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Graphics processing unit based matrix transpose optimization method

    CN103761215A

  • Systems and methods to load a tile register pair

    CN109992304A

  • Matrix storage method, matrix access method and device and electronic equipment

    CN111176582A

  • Matrix storage method and device

    CN118838537A

  • Access method of matrix data and storage device of the matrix data

    CN1971537A