Data compression and operation method oriented to storage and calculation integrated framework and storage and calculation integrated chip
By aggregating and symmetric processing of the sparse matrix, combined with block coding, the problems of low storage efficiency and poor computing energy efficiency in the integrated storage and computing architecture are solved, and efficient storage resource utilization and computing parallelism are achieved.
Patent Information
- Application Number
- CN202510382866.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has low storage efficiency, poor computing energy efficiency, limited storage density and operation accuracy when oriented towards symmetric sparse matrices. Traditional reordering algorithms do not combine symmetric processing, blocking strategies do not fully consider non-zero element density distribution, and the bit slice mapping of multi-value memory lacks adaptation, making it difficult to achieve high parallel computing.
By compressing the sparse matrix format, non-zero elements are aggregated to the main diagonal of the matrix, combined with symmetry processing, orthogonal symmetric transformation matrix is generated, and the sub-blocks of non-zero elements are binary coded and sliced after blocking, and deployed in a memory integrated memory array. The matrix elements are characterized by conductance values of the resistive variable unit for in-memory simulation multiplication and accumulation operations.
It significantly improves the utilization rate of storage resources, reduces storage space demand by 30%-50%, reduces data migration energy consumption by more than 60%, and improves computing parallelism and computing efficiency.
Smart Images

Figure CN120301435A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of signal processing, and more specifically, relates to a data compression and operation method for a memory-computation integrated architecture, and a memory-computation integrated chip. Background Art
[0002] With the rapid development of artificial intelligence and high-performance computing, the memory-computation integrated architecture has attracted much attention because it can effectively alleviate the "memory wall" problem. This architecture embeds computing units into the memory array and directly performs operations at the data storage location, significantly reducing the data transfer overhead, and is especially suitable for processing data-intensive tasks such as sparse matrices. However, the efficient compression of sparse matrices and in-memory computing still face many challenges and have become a hot research direction.
[0003] In the prior art, there are various methods for sparse matrix compression. For example, the Cuthill-McKee algorithm reduces the matrix bandwidth through reordering, and the block coding technique improves the storage locality. At the same time, the research on the memory-computation integrated architecture focuses on the bit-slice mapping of multi-value memories and the design of analog multiply-accumulate circuits. However, these methods have obvious deficiencies when facing symmetric sparse matrices: First, the traditional reordering algorithm does not combine symmetric processing, resulting in redundant transformation matrices; second, the block strategy does not fully consider the density distribution of non-zero elements, and the storage space utilization rate is low; third, the bit-slice mapping of multi-value memories lacks adaptation to the characteristics of sparse matrices and it is difficult to achieve high-parallel computing. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of this application is to provide a data compression and operation method for a memory-computation integrated architecture, and a memory-computation integrated chip, aiming to solve the problems of low storage efficiency, poor computing energy efficiency, limited storage density, and limited operation accuracy in the prior art.
[0005] The first aspect of this application relates to a data compression method for a memory-computation integrated architecture, including: Compressing the data in the sparse matrix format so that non-zero elements gather towards the main diagonal of the matrix to obtain a first compressed matrix and a transformation matrix; Symmetrically processing the transformation matrix to obtain an orthogonal symmetric transformation matrix; Performing a similarity transformation on the sparse matrix according to the orthogonal symmetric transformation matrix to obtain a second compressed matrix; Based on the density of matrix elements, partitioning the second compressed matrix to obtain sub-blocks with non-zero elements aggregated and sub-blocks with zero elements; Respectively performing binary encoding and slicing on the data in each sub-block with non-zero elements aggregated and the data in the orthogonal symmetric transformation matrix to obtain a plurality of bit-slice matrices; Deploy each bit-slice matrix in the memory-in-memory array, and the resistance state combination of each memory cell represents the multi-bit value corresponding to the matrix bit.
[0006] Preferably, the data in the sparse matrix format is compressed so that non-zero elements gather towards the main diagonal of the matrix. Specifically: regard the input sparse matrix as the adjacency matrix of a graph, and generate a compressed matrix and a transformation matrix through a node reordering algorithm based on the graph structure.
[0007] Preferably, the transformation matrix is symmetrized to obtain an orthogonal symmetric transformation matrix, specifically as follows: Divide the transformation matrix into multiple sub-blocks so that the number of non-zero elements in the non-diagonal sub-block is equal to that in its symmetric sub-block; Perform elementary row and column transformations on each sub-block so that the non-zero elements of the diagonal sub-block are arranged along the main diagonal, and at the same time, the non-diagonal sub-block and its symmetric sub-block form an axisymmetric structure; Reassemble all the block matrices together to form an orthogonal symmetric transformation matrix.
[0008] Preferably, the transformation matrix is divided into multiple sub-blocks so that the number of non-zero elements in the non-diagonal sub-block is equal to that in its symmetric sub-block, specifically as follows: Traverse the points near the envelope line of the first compressed matrix, and use its horizontal and vertical coordinates as the block markers of the transformation matrix; Apply it to the transformation matrix to form a block transformation matrix, so that it forms blocks that are completely symmetric along the diagonal; Verify whether the number of non-zero elements in the non-diagonal sub-block and its symmetric sub-block is equal. If it is equal, retain the block marker; otherwise, do not retain the block marker.
[0009] Preferably, perform elementary row and column transformations on each sub-block so that the non-zero elements of the diagonal sub-block are arranged along the main diagonal, and at the same time, the non-diagonal sub-block and its symmetric sub-block form an axisymmetric structure, specifically as follows: Move the non-zero elements in the diagonal sub-block to the diagonal of the same column, and assign the row coordinate index of the non-zero elements in the non-diagonal sub-block to the column coordinate index of the corresponding non-zero elements in the symmetric sub-block.
[0010] Preferably, the second compressed matrix is divided into blocks based on the density of matrix elements to obtain sub-blocks with concentrated non-zero elements and zero-element sub-blocks, specifically as follows: Screen out the regions in the second compressed matrix where the density of non-zero elements is higher than a preset threshold as high-density regions; Divide the high-density regions into rectangular sub-blocks with sizes adapted to the memory array structure; Retain the relative position relationship information between the sub-blocks.
[0011] Preferably, the deployment of each bit-slice matrix in the memory-in-computation integrated memory array is as follows: The bit slices of each non-zero sub-block in the second compression matrix are each divided into layers in the order from the most significant bit to the least significant bit; Allocate a storage array with a block size equal to that of the non-zero sub-block to store the corresponding layer of bit slices, where the number of states of the storage units in the th block storage array is , so that ; Perform mapping on the bit-slice layers so that each storage array simultaneously satisfies: 1) The storage units with the same coordinates on each block storage array store the corresponding bits of different bit-slice layers within the same group; 2) The different resistance states of the same storage unit represent different encodings of the data at the corresponding coordinates of the mapped bit-slice layers.
[0012] The second aspect of the present application relates to a data operation method for a memory-in-computation integrated architecture, including: S1. Apply the input vector in the form of a driving voltage in parallel to the word lines of the memory array storing the orthogonal symmetric transformation matrix, and implement the multiplication operation of the orthogonal symmetric transformation matrix and the input vector through Ohm's law and Kirchhoff's current law. Obtain the first current vector through the reading circuit as the transformed input vector; S2. Use the elements corresponding to the row indices in the block matrix in the transformed input vector as the input, and apply them in parallel to the word lines of the memory array storing the non-zero sub-blocks of the second compression matrix in the form of a driving voltage. Perform the multiplication operation of the second compression matrix and the transformed input vector through Ohm's law and Kirchhoff's current law, and obtain the second current vector through the reading circuit as the transformed output vector; S3. Use the transformed output vector as the input of the orthogonal symmetric transformation matrix, apply it in parallel to the word lines of the memory array storing the orthogonal symmetric transformation matrix in the form of a driving voltage, and implement the multiplication operation of the inverse matrix of the orthogonal symmetric transformation matrix and the transformed output vector through Ohm's law and Kirchhoff's current law. Obtain the value of the current through the reading circuit as the restored output vector.
[0013] The third aspect of the present application relates to a memory-in-computation integrated chip, including: an in-memory operation module and a digital computing module; where The in-memory operation module includes at least one memory-in-computation integrated core. Each memory-in-computation integrated core includes: a plurality of memory-in-computation operation cores arranged in parallel, each operation core having an independent data input interface and an intermediate result output interface, and being configured to execute the data operation method described in the second aspect; a shared shift-accumulation circuit, whose input end is connected to the output end of the corresponding memory-in-computation operation core; Each memory - computing unit includes: a memory array, a driving circuit, and a reading circuit; wherein, the memory array is composed of non - volatile storage units and is configured to store the matrix conductance values after compression processing; the driving circuit is connected to the instruction output terminal of the processor and is configured to convert digital control signals into driving voltage signals and load them to the word - line nodes of the memory array in a time - sharing manner; the reading circuit is connected to the bit - line nodes of the memory array and is configured to convert the analog current signals generated by the storage units into voltage signals that can be processed; The shift - accumulation circuit is configured to: receive multiplexed voltage signals output by each memory - computing unit in a time - sharing manner, perform a displacement operation with bit - width expansion on each path of signals, accumulate and synthesize the multiplexed signals after displacement in a predetermined time sequence, and transmit the final operation result back to the processor through the data bus; The digital computing module includes a processor and a memory; the processor performs data interaction with the memory - computing core and the memory respectively through a bidirectional bus and is configured to: execute the program stored in the memory. When the program stored in the memory is executed, the processor is configured to execute the data compression method described in the first aspect to generate compressed matrix parameters and distribute matrix operation control instructions to the memory - computing core; the memory stores a computer program.
[0014] The fourth aspect of the present application relates to a computer - readable storage medium. The computer - readable storage medium stores a computer program. When the computer program runs on a processor, it causes the processor to enter the data compression method described in the first aspect, or the data operation method described in the second aspect.
[0015] Generally speaking, compared with the prior art through the above - conceived technical solutions of the present application, the following beneficial effects are obtained: (1) Aiming at the problem of low storage efficiency of sparse matrices, the present application aggregates non - zero elements near the main diagonal and combines block division according to matrix density distribution, and only performs subsequent processing and storage on the sub - blocks where non - zero elements are aggregated, reducing the storage space requirement by 30% - 50% and significantly improving the utilization rate of storage resources.
[0016] (2) Aiming at the problem of large data transfer overhead in the memory - computing separation architecture, the present application maps bit - sliced layers to a multi - value storage array, uses the conductance values of resistive - switching units to directly represent matrix elements, realizes in - memory analog multiplication and accumulation operations, reduces data migration energy consumption by more than 60%, and improves the computing parallelism at the same time.
[0017] (3) Aiming at the problem of insufficient optimization of symmetric matrices, the present application, through the synergistic effect of an orthogonal symmetric transformation matrix and its inverse matrix, transforms the original sparse matrix into a new compressed matrix, and the operation complexity is reduced from to , significantly improving the operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic diagram of a memory - in - computing chip structure provided by an embodiment of the present application.
[0019] Figure 2 is a flowchart of a data compression method for a memory - in - computing architecture provided by an embodiment of the present application.
[0020] Figure 3 is a schematic diagram of the processing process of the data compression method provided by an embodiment of the present application.
[0021] Figure 4 is a flowchart of a data operation method for a memory - in - computing architecture provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] In this embodiment, the term "and / or" is a relationship describing associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In this embodiment, the symbol " / " represents an "or" relationship between associated objects. For example, A / B represents A or B.
[0024] In this embodiment, terms such as "first" and "second" in the description and claims are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of the response messages.
[0025] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.
[0026] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more. For example, a plurality of processing units means two or more processing units; a plurality of elements means two or more elements.
[0027] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0028] As Figure 1 shown, the present application provides a memory - in - computing chip, including: an in - memory computing module and a digital computing module.
[0029] The in - memory computing module includes at least one memory - in - computing core. Each memory - in - computing core includes: a plurality of memory - in - computing operation cores arranged in parallel. Each operation core has an independent data input interface and an intermediate result output interface, and is configured to execute the data operation method proposed by the present application, specifically including storing a second compressed matrix, an orthogonal symmetric transformation matrix, and performing matrix - vector multiplication; a shared shift - add circuit, whose input end is connected to the output end of the corresponding memory - in - computing operation core.
[0030] Each memory - in - computing operation core includes: a memory array, a driving circuit, and a reading circuit; wherein, the memory array is composed of non - volatile memory cells and is configured to store the matrix conductance values after compression processing; the driving circuit is connected to the instruction output end of the processor and is configured to convert the digital control signal into a driving voltage signal and time - divisionally load it to the word - line nodes of the memory array; the reading circuit is connected to the bit - line nodes of the memory array and is configured to convert the analog current signal generated by the memory cells into a voltage signal that can be processed.
[0031] The shift - add circuit is configured to: receive multiplexed voltage signals output by each memory - in - computing operation core in a time - division manner, perform a displacement operation with bit - width expansion on each signal, accumulate and synthesize the multiplexed signals after displacement in a predetermined time sequence, and return the final operation result to the processor through the data bus.
[0032] The digital computing module includes a processor and a memory; the processor performs data interaction with the memory - in - computing core and the memory respectively through a bidirectional bus, and is configured to: execute the program stored in the memory. When the program stored in the memory is executed, the processor is configured to execute the data compression method proposed by the present application to generate compressed matrix parameters, and distribute matrix operation control instructions to the memory - in - computing core; the memory stores a computer program.
[0033] The non - volatile memory array includes but is not limited to: a cross - bar structure, including multiple word lines (column lines) and multiple bit lines (row lines). Among them, the multi - value memory array is composed of multiple programmable resistance units, and the types of resistance units include but are not limited to 1R, 1T1R, 2T2R, and the multi - value characteristics include but are not limited to 1 bit and multiple bits.
[0034] As Figure 2 shown, the present application provides a data compression method for a memory - in - computing architecture. The compression method includes: Step S110: Compress the data in sparse matrix format so that non-zero elements gather towards the main diagonal of the matrix, obtaining a first compressed matrix and a transformation matrix.
[0035] Preferably, consider the input symmetric sparse matrix as the adjacency matrix of a graph, and use a node reordering algorithm based on the graph structure that can compress the matrix bandwidth to generate a first compressed matrix and a transformation matrix. The node reordering algorithm for the compressed matrix is implemented in a digital computing module, specifically, it can be implemented in a processor or other compression processing methods. For the convenience of explanation, take the example of compressing matrix A using the Cuthill-McKee algorithm, where A is the input symmetric sparse matrix.
[0036] First, according to the Cuthill-McKee algorithm, compress the input symmetric sparse matrix to compress non-zero elements near the diagonal of the matrix. For example, as Figure 3 shown, after the input symmetric sparse matrix passes through the Cuthill-McKee algorithm, a transformation matrix is obtained, and then multiply the input symmetric sparse matrix by the transformation matrix and the transpose of the transformation matrix successively on the left and right to obtain the compressed matrix after compression.
[0037] Step S120: Symmetrize the transformation matrix to obtain an ortho-symmetric transformation matrix.
[0038] The symmetrization process is implemented in a digital computing module, specifically, it can be implemented in a processor.
[0039] Preferably, the symmetrization process includes: dividing the transformation matrix into multiple sub-blocks so that the number of non-zero elements in the off-diagonal sub-blocks is equal to that of their symmetric sub-blocks; performing elementary row and column transformations on each sub-block so that the non-zero elements in the diagonal sub-blocks are arranged along the main diagonal, and at the same time, the off-diagonal sub-blocks and their symmetric sub-blocks form an axisymmetric structure; reassembling all the block matrices together to form an ortho-symmetric transformation matrix.
[0040] In one embodiment, as Figure 3 shown, first, traverse the points near the envelope line of the first compressed matrix, and use its horizontal and vertical coordinates as the block markers of the transformation matrix; then, apply it to the transformation matrix to form a block transformation matrix, making it form blocks that are completely symmetric along the diagonal; then, verify whether the number of non-zero elements in the off-diagonal sub-blocks and their symmetric sub-blocks is equal. If equal, retain the block marker, otherwise, do not retain the block marker; then, symmetrize the non-zero elements in the block matrix, that is, move the non-zero elements in the diagonal sub-blocks to the diagonal of the same column, and assign the row coordinate index of the non-zero elements in the off-diagonal sub-blocks to the column coordinate index of the corresponding non-zero elements in the symmetric sub-block; finally, reassemble all the block matrices together to form an ortho-symmetric transformation matrix.
[0041] Step S130: Perform a similarity transformation on the sparse matrix according to the orthogonal symmetry transformation matrix to obtain a second compression matrix.
[0042] For generating the second compression matrix, it is implemented in the digital calculation module, and specifically, it can be implemented in a processor.
[0043] In one embodiment, as Figure 3 shown, multiply the input symmetric sparse matrix successively on the left by the orthogonal symmetry transformation matrix and on the right by the inverse of the orthogonal symmetry transformation matrix to obtain the second compression matrix.
[0044] Step S140: According to the density of the elements in the second compression matrix, divide it into sub-blocks where non-zero elements gather and zero-element sub-blocks.
[0045] For the division of the sub-blocks where non-zero elements gather, it is implemented in the digital calculation module, and specifically, it can be implemented in a processor.
[0046] Preferably, the block processing specifically includes: screening out the regions in the second compression matrix where the density of non-zero elements is higher than a preset threshold as high-density regions; dividing the high-density regions into rectangular sub-blocks with dimensions adapted to the memory array structure; and retaining the information on the relative position relationship between the sub-blocks.
[0047] In one embodiment, as Figure 3 shown, first, apply the final block marking to the compression matrix, and then mark the sub-blocks without non-zero elements as zero-element sub-blocks; then, mark the sub-blocks containing non-zero elements as sub-blocks where non-zero elements gather; then, gather the sub-blocks where non-zero elements gather into large sub-blocks in the order from top to bottom and from left to right; finally, form the sub-blocks where non-zero elements gather as shown by the dashed box in the new compression matrix, that is, the matrix that needs to be finally stored after compression.
[0048] Step S150: After representing the sub-blocks and the symmetry transformation matrix in binary encoding, split them to form binary bit-slice matrices.
[0049] For the binary encoding representation and the splitting of the bit-slice matrices, it is implemented in the digital calculation module, and specifically, it can be implemented in a processor. Preferably, the binary encoding and slicing include: converting the numerical data into complement code or original code representation; generating a bit-slice matrix by bit positions; where the most significant bit-slice corresponds to the highest resistance state characterization layer of the storage array.
[0050] In one embodiment, one of the block matrices that needs to be finally stored in the second compression matrix mapped in the current memory array is , for any one of its elements ( ) can be split into several binary numbers according to binary coding . Correspondingly, the block matrix can be split into several bit-slice matrices , …, .
[0051] Step S160: Deploy each bit-slice matrix in the memory-computing integrated memory array, and the resistance state combination of each storage unit represents the multi-bit value corresponding to the matrix bit.
[0052] The deployment of the bit-slice matrix is implemented in the memory-computing integrated core, specifically, it can be implemented on the memory array.
[0053] Preferably, the deployment of the bit-slice matrix is as follows: The bit-slices of each non-zero sub-block in the second compression matrix are respectively divided into layers in the order from the most significant bit to the least significant bit; allocate a storage array with the same size as the non-zero sub-block to store the corresponding layer bit-slices, where the number of states of the storage units in the th block storage array is , so that ; perform mapping on the bit-slice layer, so that each storage array simultaneously satisfies: 1) The storage units with the same coordinates on each block storage array store the corresponding bits of different bit-slice layers in the same group; 2) Different resistance states of the same storage unit represent different encodings of the data at the corresponding coordinates of the mapped bit-slice layer.
[0054] Preferably, the compressed matrix data structure is stored in the following three-part form: Bit-slice layer mapping relation table: Record the spatial coordinate mapping rules of each binary bit-slice layer in the storage array; Symmetric transformation matrix metadata: Store the parameters of the symmetric transformation matrix and the inverse matrix information; Block position index information: Record the row and column start addresses and sizes of each sub-block after block processing; The physical storage address of the data structure corresponds one-to-one with the spatial coordinates of the resistive random access memory array.
[0055] Taking the bit-slice matrix of the MSB as an example, the same column in it is stored on the same bit line of the memory array.
[0056] As Figure 4 shown, the present application provides a data operation method for a memory-computing integrated architecture, and the operation method includes: Step S210: Use the sparse matrix compression method to generate and deploy the second compression matrix and the corresponding orthogonal symmetric transformation matrix.
[0057] Step S220: Orthogonally multiply the input vector and the symmetric transformation matrix to obtain the transformed input vector.
[0058] The multiplication of the input vector and the orthogonal symmetric transformation matrix is implemented in the in-memory computing kernel, specifically, it can be implemented on the memory array.
[0059] In one embodiment, for the sake of convenience of explanation, all matrices are binary. The input vector is applied in parallel to the word lines of the memory array storing the orthogonal symmetric transformation matrix in the form of a driving voltage, and the multiplication operation of the orthogonal symmetric transformation matrix and the input vector is realized through Ohm's law and Kirchhoff's current law. The first current vector is obtained through the read circuit as the transformed input vector.
[0060] Step S230: Multiply the transformed input vector by the second compression matrix to obtain the transformed output vector.
[0061] The multiplication of the transformed input vector and the second compression matrix is implemented in the in-memory computing kernel, specifically, it can be implemented on the memory array.
[0062] In one embodiment, the elements in the transformed input vector corresponding to the row indices of the block matrix are used as inputs and applied in parallel to the word lines of the memory array storing the non-zero sub-blocks of the second compression matrix in the form of a driving voltage. The multiplication operation of the second compression matrix and the transformed input vector is performed through Ohm's law and Kirchhoff's current law. The second current vector is obtained through the read circuit as the transformed output vector.
[0063] Step S240: Multiply the transformed output vector by the orthogonal symmetric transformation matrix to obtain the actual output vector.
[0064] The multiplication of the transformed output vector and the orthogonal symmetric transformation matrix is implemented in the in-memory computing kernel, specifically, it can be implemented on the memory array.
[0065] In one embodiment, the transformed output vector is used as the input of the orthogonal symmetric transformation matrix and applied in parallel to the word lines of the memory array storing the orthogonal symmetric transformation matrix in the form of a driving voltage. The multiplication operation of the inverse matrix of the orthogonal symmetric transformation matrix and the transformed output vector is realized through Ohm's law and Kirchhoff's current law. The value of the current is obtained through the read circuit as the restored output vector.
[0066] Preferably, the matrix-vector multiplication operation is implemented through the following physical processes: a) Each vector element corresponds to a column of driving voltage of the memory array; b) Each matrix element is characterized by the conductive conductance value of the storage unit; c) The output current realizes the multiplication accumulation operation through column line summation.
[0067] Preferably, the inverse transformation operation includes: applying an input vector to the column line ends of an array in a memory array of an orthogonal symmetric transformation matrix; and reading output signals at the row line ends.
[0068] Preferably, the voltage signal conversion includes: converting a current signal into a voltage signal by using a transimpedance amplifier; and converting an analog voltage signal into a digital voltage signal by using an analog-to-digital converter circuit.
[0069] The above data compression method and operation method can be widely applied to all tasks involving deploying a sparse matrix on a memory-computation integrated architecture and performing matrix multiplication operations. Its application scenarios include, but are not limited to, solving sparse matrix equations, neural network inference, cryptographic accelerators, etc., and the system has high reconfigurability.
[0070] It should be understood that the above device is used to execute the method in the above embodiment. For the corresponding program modules in the device, their implementation principles and technical effects are similar to those described in the above method. The working process of the device can refer to the corresponding process in the above method, and will not be elaborated here.
[0071] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, the processor is caused to execute the method in the above embodiment.
[0072] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the method in the above embodiment.
[0073] It can be understood that the processor in the embodiment of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, graphics processing units (GPUs), neural network processing units (NPUs), digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0074] The method steps in the embodiments of this application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), dynamic random access memory (DRAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0075] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0076] It can be understood that the various digital numbers involved in the embodiments of this application are only for convenience of description and are not used to limit the scope of the embodiments of this application.
[0077] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A data compression method for a memory-compute integrated architecture, characterized in that Including: Compressing the data in sparse matrix format so that non-zero elements gather towards the main diagonal of the matrix, obtaining a first compressed matrix and a transformation matrix; Symmetrizing the transformation matrix to obtain an orthosymmetric transformation matrix; Performing a similarity transformation on the sparse matrix according to the orthosymmetric transformation matrix to obtain a second compressed matrix; Partitioning the second compressed matrix based on the density of matrix elements, obtaining sub-blocks with non-zero elements gathered and zero-element sub-blocks; Performing binary encoding and slicing on the data in each sub-block with non-zero elements gathered and the data in the orthosymmetric transformation matrix respectively, obtaining multiple bit-sliced matrices; Deploying each bit-sliced matrix in a memory-computing integrated memory array, and the resistance state combination of each storage unit represents the multi-bit value corresponding to the matrix bit.
2. The data compression method according to claim 1, characterized in that, The compressing the data in sparse matrix format so that non-zero elements gather towards the main diagonal of the matrix is specifically as follows: regarding the input sparse matrix as the adjacency matrix of a graph, and generating a compressed matrix and a transformation matrix through a node reordering algorithm based on the graph structure.
3. The data compression method according to claim 1, wherein The symmetrizing the transformation matrix to obtain an orthosymmetric transformation matrix is specifically as follows: Dividing the transformation matrix into multiple sub-blocks so that the number of non-zero elements in the off-diagonal sub-blocks is equal to that in their symmetric sub-blocks; Performing elementary row and column transformations on each sub-block so that the non-zero elements in the diagonal sub-blocks are arranged along the main diagonal, and at the same time the off-diagonal sub-blocks and their symmetric sub-blocks form an axisymmetric structure; Reassembling all the partitioned matrices together to form an orthosymmetric transformation matrix.
4. The data compression method according to claim 3, wherein The dividing the transformation matrix into multiple sub-blocks so that the number of non-zero elements in the off-diagonal sub-blocks is equal to that in their symmetric sub-blocks is specifically as follows: Traversing the points near the envelope line of the first compressed matrix, and using their horizontal and vertical coordinates as the partition marks of the transformation matrix; Applying it to the transformation matrix to form a partitioned transformation matrix, making it form partitions that are completely symmetric along the diagonal; Verifying whether the number of non-zero elements in the off-diagonal sub-blocks and their symmetric sub-blocks is equal. If equal, retain the partition mark, otherwise, do not retain the partition mark.
5. The data compression method according to claim 3, wherein The performing elementary row and column transformations on each sub-block so that the non-zero elements in the diagonal sub-blocks are arranged along the main diagonal, and at the same time the off-diagonal sub-blocks and their symmetric sub-blocks form an axisymmetric structure is specifically as follows: Moving the non-zero elements in the diagonal sub-blocks to the diagonals of the same column, and assigning the row coordinate index of the non-zero elements in the off-diagonal sub-blocks to the column coordinate index of the corresponding non-zero elements in the symmetric sub-blocks.
6. The data compression method according to claim 1, characterized in that The partitioning the second compressed matrix based on the density of matrix elements, obtaining sub-blocks with non-zero elements gathered and zero-element sub-blocks is specifically as follows: Screening out the regions in the second compressed matrix where the density of non-zero elements is higher than a preset threshold as high-density regions; Dividing the high-density regions into rectangular sub-blocks with sizes adapted to the memory array structure; Retaining the relative position relationship information between sub-blocks.
7. The data compression method according to claim 1, wherein The deploying each bit-sliced matrix in a memory-computing integrated memory array is specifically as follows: Divide the bit slices of each non-zero sub-block in the second compression matrix into layers in the order from the most significant bit to the least significant bit; Allocation A storage array in which the blocks are equal in size to the non-zero sub-blocks is used to store the corresponding layer slices, where the number of states of the storage units in the block storage array is , so that ; Perform mapping on the bit-slice layer so that each memory array simultaneously satisfies: 1) Memory cells with the same coordinates on each memory array store corresponding bits of different bit-slice layers within the same group; 2) Different resistance states of the same memory cell represent different encodings of the data at the corresponding coordinates of the bit-slice layer it maps to.
8. A data operation method for a memory-computation integrated architecture, characterized in that It includes: S1. Apply the input vector in the form of a driving voltage in parallel to the word lines of the memory array storing the orthogonal symmetric transformation matrix. Implement the multiplication operation of the orthogonal symmetric transformation matrix and the input vector through Ohm's law and Kirchhoff's current law. Obtain the first current vector through the reading circuit as the transformed input vector. S2. Use the elements corresponding to the row indices in the partitioned matrix in the transformed input vector as the input. Apply it in the form of a driving voltage in parallel to the word lines of the memory array storing the non-zero sub-blocks of the second compression matrix. Perform the multiplication operation of the second compression matrix and the transformed input vector through Ohm's law and Kirchhoff's current law. Obtain the second current vector through the reading circuit as the transformed output vector. S3. Use the transformed output vector as the input of the orthogonal symmetric transformation matrix. Apply it in the form of a driving voltage in parallel to the word lines of the memory array storing the orthogonal symmetric transformation matrix. Implement the multiplication operation of the inverse matrix of the orthogonal symmetric transformation matrix and the transformed output vector through Ohm's law and Kirchhoff's current law. Obtain the value of the current through the reading circuit as the restored output vector.
9. An in-memory computing chip, characterized in that, It includes: An in-memory computing module and a digital computing module; where The in-memory computing module includes at least one in-memory computing core. Each in-memory computing core includes: Multiple in-memory computing units arranged in parallel, each computing unit having an independent data input interface and an intermediate result output interface, configured to execute the data operation method as described in claim 8; A shared shift-accumulation circuit, whose input end is connected to the output end of the corresponding in-memory computing unit. Each in-memory computing unit includes: A memory array, composed of non-volatile memory cells, configured to store the compressed matrix conductance values; A driving circuit, connected to the instruction output end of the processor, configured to convert the digital control signal into a driving voltage signal and time-divisionally load it to the word line nodes of the memory array; A reading circuit, connected to the bit line nodes of the memory array, configured to convert the analog current signal generated by the memory cells into a processable voltage signal. The shift-accumulation circuit is configured to: Receive the multiplexed voltage signals output by each in-memory computing unit in a time-division manner, perform a bit-width extended displacement operation on each signal, accumulate and synthesize the displaced multiplexed signals according to a predetermined time sequence, and transmit the final operation result back to the processor through the data bus. The digital computing module includes a processor and a memory; the processor performs data interaction with the memory-in-computation kernel and the memory respectively through a bidirectional bus, and is configured to: execute the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to execute the data compression method described in any one of claims 1 to 7 to generate compressed matrix parameters, and distribute matrix operation control instructions to the memory-in-computation kernel; the memory stores a computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program runs on a processor, it causes the processor to enter the data compression method described in any one of claims 1 to 7, or the data operation method described in claim 8.