Graph Structure Data Processing Method, Accelerator, Storage Medium, and Program Product
By compressing the sparse feature matrix and parallel and delayed computing of the accelerator, the storage and computing efficiency problems of sparse feature matrix in the graph neural network are solved, and efficient graph structure data processing is realized to adapt to the computing needs of data of different sparseness.
Patent Information
- Application Number
- CN202510228598.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-02-28
AI Technical Summary
In the prior art, graph structure data processing has high requirements for system cache and is difficult to adapt to the computing needs of data of different sparseness. Especially in graph neural networks, the storage and computing efficiency problems of sparse feature matrix have not been effectively solved.
By compressing the sparse feature matrix, the compressed feature matrix and state feature matrix are generated, and parallel operations and delay operations are performed by accelerator to dynamically generate intermediate matrix to avoid the storage space of intermediate results and realize cache-free pipeline calculation.
It improves computing efficiency and system resource utilization, reduces the demand for storage space, adapts to the computing needs of data of different sparseness, and improves the processing capabilities of graph neural networks.
Smart Images

Figure CN119721119B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and deep learning, and in particular, to a method for processing graph-structured data, an accelerator, a storage medium, and a program product. Background Art
[0002] Graph-structured data is widely used in fields such as social network analysis, resource recommendation, knowledge graphs, supply chain management, and bioinformatics. In related technologies, graph neural networks are used to capture the topological structure and node features of graph-structured data, which can support relatively complex data analysis tasks. However, in related technologies, the data analysis process has high requirements for the system cache and is difficult to adapt to the computational requirements of data with different sparsities. Summary of the Invention
[0003] In view of the above problems, the present invention provides a method for processing graph-structured data, an accelerator, a device, a device, a storage medium, and a program product.
[0004] According to a first aspect of the present invention, there is provided a method for processing graph-structured data, which is applied to a graph neural network accelerator. The graph-structured data includes a sparse feature matrix corresponding to the graph-structured data to be processed, and the method includes: reading a compressed feature matrix and a state feature matrix, where the compressed feature matrix is obtained by compressing at least one of the row position information and numerical information of the elements in the sparse feature matrix based on a preset compression strategy, and the state feature matrix represents the state feature of the compressed feature matrix in the preprocessing; performing a first matrix operation on the compressed feature matrix and a weight matrix to obtain a first intermediate matrix, and performing a second matrix operation on the state feature matrix and the first intermediate matrix to obtain a second intermediate matrix, where the first matrix operation includes parallel operations and delay operations performed based on a plurality of gating algorithms of the accelerator, and the first intermediate matrix includes a plurality of parallel intermediate matrices obtained based on the parallel operations and a delay intermediate matrix obtained based on the delay operations; obtaining a target matrix according to the first intermediate matrix and the second intermediate matrix, so as to use the target matrix for a target task.
[0005] The second aspect of the present invention provides a graph structure data processing device, including: a reading module configured to read a compressed feature matrix and a status feature matrix, where the compressed feature matrix is obtained by compressing at least one of the row position information and numerical information of elements in a sparse feature matrix based on a preset compression strategy, and the status feature matrix characterizes the status feature of the compressed feature matrix in preprocessing; an operation module configured to perform a first matrix operation on the compressed feature matrix and a weight matrix to obtain a first intermediate matrix, and perform a second matrix operation on the status feature matrix and the first intermediate matrix to obtain a second intermediate matrix, where the first matrix operation includes parallel operation and delayed operation, and the first intermediate matrix includes a parallel intermediate matrix and a delayed intermediate matrix; a task execution module configured to obtain a target matrix based on the first intermediate matrix and the second intermediate matrix, so as to use the target matrix to perform a target task.
[0006] The third aspect of the present invention provides a graph neural network accelerator, including: a memory; a processor configured to execute the above-mentioned graph structure data processing method according to instructions stored in the memory.
[0007] The fourth aspect of the present invention provides an electronic device, including: one or more processors; a memory configured to store one or more computer programs, where the above-mentioned one or more processors execute the above-mentioned one or more computer programs to implement the steps of the above-mentioned method.
[0008] The fifth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instructions are stored, and when the above-mentioned computer program or instructions are executed by a processor, the steps of the above-mentioned method are implemented.
[0009] The sixth aspect of the present invention further provides a computer program product, including a computer program or instructions, and when the above-mentioned computer program or instructions are executed by a processor, the steps of the above-mentioned method are implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above-mentioned content and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0011] Figure 1A Shows an example diagram of compressing sparse data in a coordinate list format;
[0012] Figure 1B Shows a schematic structural diagram of a long short-term memory network unit;
[0013] Figure 2 Shows an application scenario diagram of the graph structure data processing method, accelerator, storage medium, and program product according to the embodiments of the present invention;
[0014] Figure 3Shows a flowchart of a graph structure data processing method according to an embodiment of the present invention;
[0015] Figure 4 Shows an example schematic diagram of the compression process of a sparse feature matrix according to an embodiment of the present invention;
[0016] Figure 5A Shows an example block diagram of a graph neural network accelerator according to an embodiment of the present invention;
[0017] Figure 5B Shows an example schematic diagram of the hardware architecture of an accelerator according to an embodiment of the present invention;
[0018] Figure 6 Shows a block diagram of the structure of a graph structure data processing device according to an embodiment of the present invention;
[0019] Figure 7 Shows a block diagram of an electronic device suitable for implementing the graph structure data processing method according to an embodiment of the present invention. Detailed Embodiments
[0020] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0021] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0023] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0024] In the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0025] In some examples, during the process of processing sparse data using the acceleration method of a graph neural network, the processor needs to automatically clear invalid zero values; or a compression algorithm is used to eliminate zero values in the data.
[0026] For example, the processor automatically filters non-zero data and only retains valid data for calculation. Although this method can output valid non-zero data, since the input is original sparse data, the output rate will be affected by the calculation process of eliminating zero values, and essentially does not really improve the effective rate of non-zero data entering the calculation unit.
[0027] For example, the row address is modified to a row end flag to reduce storage requirements, but this method does not consider the special case where all values in a row of sparse large data are all zero. The calculation process will still traverse these rows, resulting in unnecessary calculation overhead.
[0028] Figure 1A Shows an example diagram of compressing sparse data using a coordinate list format.
[0029] As Figure 1A shown, the data is compressed using the coordinate list format 100A, and then the compressed data is sent to the processor for calculation. Although this compression method can eliminate zero values in the data, it will simultaneously increase the row index 101 and column index 102 of non-zero data. The index values of the row index 101 and column index 102 occupy a large storage space, and the storage overhead of the index values is multiple times that of the data itself. If the data (value) 103 is one byte, for the address index of ultra-large data volume, generally 4 bytes are required, and a total of 8 bytes are required for both rows and columns corresponding to one byte of data, resulting in a large occupation of storage space.
[0030] In some examples, for the gating algorithm of a long short-term memory network, some calculation results need to be cached during the calculation process and wait until all corresponding calculation data is completed before calculating the next step of data.
[0031] Figure 1BShows a schematic structural diagram of a long short-term memory network unit.
[0032] As Figure 1B shown, in combination with the gating algorithm formula of the long short-term memory network, after inputting h (t-1) , i t , f t , g t , o t can be synchronously calculated. However, when calculating h t , it is necessary to first calculate c t , then calculate tanh(c t ), and then calculate h t by performing calculations with the previously calculated o t . In the related art, by keeping the o t value in the cache and waiting for the calculation result of tanh(c t ). If the data volume is small, the cache is usually acceptable. However, for the intermediate results of matrix calculations with a large data volume in graph neural networks, the demand for storage resources is large and will increase with the increase in the calculation parallelism.
[0033] Based on the above problems, the present invention provides a graph structure data processing method, accelerator, device, equipment, storage medium and program product, which are applied to a graph neural network accelerator. The graph structure data includes a sparse feature matrix corresponding to the graph structure data to be processed, and includes: reading a compressed feature matrix and a state feature matrix, where the compressed feature matrix is obtained by compressing the row position information and numerical information of at least one element in the sparse feature matrix based on a preset compression strategy, and the state feature matrix represents the state feature of the compressed feature matrix in the preprocessing; performing a first matrix operation on the compressed feature matrix and the weight matrix to obtain a first intermediate matrix, and performing a second matrix operation on the state feature matrix and the first intermediate matrix to obtain a second intermediate matrix, where the first matrix operation includes parallel operation and delayed operation, and the first intermediate matrix includes a parallel intermediate matrix and a delayed intermediate matrix; obtaining a target matrix according to the first intermediate matrix and the second intermediate matrix, so as to use the target matrix for a target task.
[0034] According to an embodiment of the present invention, a compressed feature matrix is obtained by compressing the row position information and numerical information of the elements in the sparse feature matrix, and then matrix operations are performed on the compressed compressed feature matrix, weight matrix, and state feature matrix. The sparsity of the sparse feature matrix in the graph-structured data and the density of other matrices (including the weight matrix, state feature matrix, and first intermediate matrix) are comprehensively considered in the operation process, and the calculation characteristics of sparse matrices and dense matrices are compatible. Since the first intermediate matrix is obtained by performing parallel operations and delayed operations on the compressed feature matrix and the weight matrix based on multiple gating algorithms of the accelerator, the first intermediate matrix is dynamically generated according to requirements, realizing cacheless pipelining calculation, avoiding the occupation of storage space by intermediate results, and further improving the calculation efficiency and the utilization rate of system resources.
[0035] Figure 2 FIG. shows an application scenario diagram of a graph-structured data processing method, accelerator, storage medium, and program product according to an embodiment of the present invention.
[0036] As Figure 2 shown, the application scenario according to this embodiment may include a terminal device 201, a server 202, an external storage unit 203, and a network 204. The network 204 is used to provide a medium for communication links between the terminal device 201, the server 202, and the external storage unit 203. The network 204 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0037] The user can use the terminal device 201 to interact with the server 202 through the network 204 to receive or send messages, etc. The terminal device 201 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.
[0038] The server 202 can be used to process graph-structured data. For example, different types of feature matrices are read from the external storage unit 203, and matrix operations are performed on different matrices according to operation instructions, and the operation results are sent to the external storage unit 203 for storage, or the operation results are displayed through the terminal device 201. For example, the server 202 can recommend resources to the user according to the operation results, or analyze relevant physical examination data to achieve classification and prediction of causes of diseases.
[0039] It should be noted that the graph structure data processing method provided by the embodiments of the present invention can generally be executed by the server 202. Correspondingly, the graph structure data processing device provided by the embodiments of the present invention can generally be set in the server 202. The graph structure data processing method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 202 and capable of communicating with the server 202. Correspondingly, the graph structure data processing device provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 202 and capable of communicating with the server 202.
[0040] It should be understood that the numbers of the terminal devices, the network, and the servers in FIG. 1 are merely illustrative. According to the implementation requirements, any number of terminal devices, external storage units, networks, and servers can be provided.
[0041] Figure 3 The flowchart of the graph structure data processing method according to the embodiments of the present invention is shown.
[0042] As Figure 3 shown, the graph structure data processing method of this embodiment is applied to a graph neural network accelerator. The graph structure data includes a sparse feature matrix corresponding to the graph structure data to be processed. The graph structure data processing method includes operations S310 to S330.
[0043] In operation S310, a compressed feature matrix and a state feature matrix are read. The compressed feature matrix is obtained by compressing at least one of the row position information and the numerical information of the elements in the sparse feature matrix based on a preset compression policy. The state feature matrix represents the state feature of the compressed feature matrix in the preprocessing.
[0044] In the embodiments of the present invention, the external storage unit can be a memory corresponding to the accelerator for storing the data to be processed or the processed result data (including but not limited to graph structure data, compressed feature matrix, and state feature matrix). The accelerator can be a processor that processes the graph data structure using a preset algorithm. The row position information of the element can represent the position of the element in the current row information, including the row end position or the non-row end position. The numerical information can represent that the element is a zero value or a non-zero value. The state feature in the preprocessing can represent the state feature obtained in the previous cycle corresponding to the current calculation cycle in the graph structure data processing method.
[0045] In operation S320, a first matrix operation is performed on the compressed feature matrix and the weight matrix to obtain a first intermediate matrix, and a second matrix operation is performed on the state feature matrix and the first intermediate matrix to obtain a second intermediate matrix. The first matrix operation includes parallel operation and delay operation. The first intermediate matrix includes a parallel intermediate matrix and a delay intermediate matrix.
[0046] In an embodiment of the present invention, the method for processing graph-structured data may be a graph data processing algorithm based on a Long Short-Term Memory (LSTM) network. Corresponding matrix operations are performed on different matrices according to the gating algorithm and related state algorithms in the long short-term memory network. There can be multiple gating algorithms, and different intermediate matrices can be obtained according to different gating algorithms. The related state algorithms may include a memory cell state algorithm and a hidden state algorithm.
[0047] For example, based on multiple gating algorithms of the accelerator and the operation duration information corresponding to the gating algorithms, parallel operations and delayed operations can be respectively performed on the compressed feature matrix and the weight matrix at different times according to the cacheless pipeline calculation requirements, to obtain a first intermediate matrix. For example, based on the memory cell state algorithm, a second matrix operation is performed on the state feature matrix and the obtained first intermediate matrix to obtain a second intermediate matrix.
[0048] In operation S330, a target matrix is obtained based on the first intermediate matrix and the second intermediate matrix, so as to use the target matrix for a target task.
[0049] In an embodiment of the present invention, the target matrix may be the final output result matrix obtained after performing multiple cycle operations. After obtaining the first intermediate matrix and the second intermediate matrix, a matrix multiplication operation can be performed on the first intermediate matrix and the second intermediate matrix using an activation function to obtain the target matrix. After obtaining the target matrix, according to actual requirements, resource recommendation, event prediction, or relationship reasoning can be performed on the target object based on the target matrix.
[0050] For example, through the graph-structured data processing method, the interaction sequence between users and items and the social relationships between users are analyzed to achieve more accurate resource recommendation. For example, by analyzing the time series characteristics of user behavior and combining with a graph neural network to extract the relationship structure between users, the behavior pattern of users or potential friends can be predicted. For example, by analyzing the time dynamic changes of biological sequence data (such as gene sequences) and combining with a graph neural network, the structural information of a biological network (such as a protein interaction network) can be processed, and the hidden patterns in the biological network can be analyzed for disease prediction and drug discovery.
[0051] According to an embodiment of the present invention, a compressed feature matrix is obtained by compressing the row position information and numerical information of elements in a sparse feature matrix, and then matrix operations are performed on the compressed compressed feature matrix, weight matrix, and state feature matrix. The operation process comprehensively considers the sparsity of the sparse feature matrix in graph-structured data and the density of other matrices (including the weight matrix, state feature matrix, and first intermediate matrix), and is compatible with the calculation characteristics of sparse matrices and dense matrices. Since the first intermediate matrix is obtained by performing parallel operations and latency operations on the compressed feature matrix and the weight matrix based on multiple gating algorithms of the accelerator, the first intermediate matrix is dynamically generated according to requirements, realizing cacheless pipelined computing, avoiding the occupation of storage space by intermediate results, and further improving the computing efficiency and the utilization rate of system resources.
[0052] It can be understood that the method for processing graph-structured data has been described above, and the compression method of the sparse feature matrix will be described below.
[0053] According to an embodiment of the present invention, the features of an element include at least one of an all-zero row feature and a non-all-zero row feature, and the non-all-zero row feature includes at least one non-zero element; the processing method further includes: in the case where the feature of the current element is an all-zero row feature or a non-all-zero row feature, determining a label value corresponding to the all-zero row feature or determining a label value corresponding to the non-zero element based on the label correspondence relationship, where the label correspondence relationship represents the relationship between multiple label values and the all-zero row feature and non-zero elements.
[0054] In an embodiment of the present invention, considering the problems of uncertain input data length, large amount of data for single matrix calculation, and irregular data sparsity in graph neural networks, the matrix multiplication in LSTM is performed in a row-by-row multiplication manner. For example, taking a row (such as L1) in the sparse feature matrix with only two non-zero elements as an example, the first non-zero element a in L1 is multiplied and added with the corresponding weight row W l1 (including multiple weight values) in the weight matrix, and the multiplication and addition result of the second non-zero element b and the corresponding weight row W l2 is accumulated, and then the first row data of the final data result can be calculated. By this method, there is no limit to the length of the sparse feature matrix to be input, and only the non-zero data needs to be screened for effective calculation.
[0055] In the related art, before performing row multiplication on a sparse feature matrix, it is necessary to extract the valid data (non-zero data) in the sparse feature matrix (for example, automatically filter through a processor or compress the data through compression software), but these methods are difficult to effectively solve the problem of storage space occupation. To solve this technical problem, according to the characteristics of row multiplication, after each complete row multiplication calculation, a row of valid data of the output result can be obtained. Therefore, for the row index value (tag value), it is only necessary to determine which data is the last data of the current row.
[0056] In an embodiment of the present invention, each element in the sparse feature matrix can include a full-zero row feature and a non-full-zero row feature based on whether the value is zero or non-zero. The full-zero row feature can represent the feature that all elements in the row corresponding to the target element are zero elements, and the non-full-zero row feature can represent the feature that the elements in the row corresponding to the target element include both zero elements and non-zero elements. The representation method of the tag value can be determined according to the actual situation. For example, bit positions can be used to represent the full-zero row feature and the non-full-zero row feature of the element.
[0057] According to an embodiment of the present invention, based on the characteristics of matrix matrix row multiplication, by recording the index of the last data in each row, the boundary of each row can be directly determined, avoiding additional traversal operations, and the row operation of the sparse matrix is more efficient, especially in the calculation of large-scale sparse matrices. In matrix multiplication, only the non-zero elements in each row need to be processed, avoiding meaningless calculations of zero elements, which can reduce the amount of calculation and improve the calculation efficiency.
[0058] According to an embodiment of the present invention, the tag value includes a first tag value and a second tag value; determining the tag value corresponding to the full-zero row feature or determining the tag value corresponding to the non-zero element based on the tag correspondence relationship includes: when the feature of the current element is the full-zero row feature, determining the tag value corresponding to the full-zero row feature as the second tag value; when the feature of the current element is the non-full-zero row feature, determining the tag value corresponding to the non-zero element as the first tag value based on the numerical information and the row position information.
[0059] In an embodiment of the present invention, the first tag value can be the tag value corresponding to the non-zero element in the non-full-zero row feature; the second tag value can be the tag value of the row corresponding to the full-zero row feature. For example, 0 represents a non-full-zero row, and 1 represents a full-zero row; if the current row is a full-zero row, no calculation is required, but a row of all 0 data (represented as 10) needs to be added during matrix operation; for a non-full-zero row, two bit positions bit can be used to represent it. The first bit, 0 represents a non-full-zero row, 1 represents a full-zero row, and the second bit, 0 represents that it is not the last data in this row, and 1 represents that it is the last data in this row.
[0060] According to an embodiment of the present invention, the method further includes: processing zero elements in the non-all-zero row features, and determining a compressed feature matrix based on a first tag value and a second tag value.
[0061] In an embodiment of the present invention, considering the resource overhead problem caused by traversing zero elements in the non-all-zero row features during the calculation process. The present invention updates and processes the zero elements in the non-all-zero row features, and then compresses the sparse matrix based on different tag values to obtain a compressed feature matrix.
[0062] According to an embodiment of the present invention, processing zero elements in the non-all-zero row features and determining a compressed feature matrix based on a first tag value and a second tag value includes: when the feature of the current element is a non-all-zero row feature, deleting the zero elements in the current row; sequentially converting non-zero elements into a first tag value, and converting the non-all-zero row feature into a second tag value to obtain a compressed feature matrix.
[0063] Figure 4 An example schematic diagram showing the compression process of a sparse feature matrix according to an embodiment of the present invention is shown.
[0064] As Figure 4 shown, in 400, the row information in the original sparse feature matrix 401 can be updated from the original four bytes to two bits in the compressed feature matrix 402. For the left bit, 0 indicates a non-all-zero row, and 1 indicates an all-zero row, which is compatible with the case of all-zero rows in the sparse feature matrix; for the right bit, 0 indicates a non-last element of the current row, and 1 indicates the last element of the current row.
[0065] When performing a row data multiplication operation and an accumulation operation, if the element is not the last element of the current row, the operation and accumulation operations continue. When the last element of the current row is recognized, the result of the accumulation operation is output as a row of the final output result, the accumulation cache is cleared, and the row multiplication operation of the next row is performed. Thus, the storage space of the index value of the row after compressing the data can be compressed to 1 / 16 of the original, and the bit width of the row data does not increase with the increase in the data volume, and is indicated as 2 bits.
[0066] In an embodiment of the present invention, by using two bits to represent the row information in the compressed feature matrix, the last element of each row can be quickly determined, so that only necessary valid information needs to be processed in the row multiplication operation, which is compatible with the operation requirements of the sparse matrix and the dense matrix of the graph structure data. By reducing unnecessary data storage and operations, the efficiency and performance of matrix operations are further improved.
[0067] In an embodiment of the present invention, the computing unit structure for matrix multiplication may include multiple multipliers and adders. Matrix operations can be performed through a hardware accelerator (such as a field-programmable gate array). For example, by inputting the data to be input into the multiplier for multiplication operations, with the multiplier corresponding to a row or a column of the matrix, multiple multiplication operations can be simultaneously executed through parallel operation. Each multiplier can receive an element from a different matrix for multiplication operation; thus, the adder is used to accumulate the results from the corresponding multipliers, and multiple adders can perform parallel operations to handle multiple accumulation operations simultaneously. Through the multiplication and addition operations of matrix elements, each time only one element in the sparse matrix and a row in the corresponding weight matrix are multiplied and then accumulated until a row of data in the sparse feature matrix is calculated. After detecting the end-of-row flag of the current element, it can indicate that a row of the matrix has been calculated and output, and the cache in the computing unit is cleared.
[0068] According to an embodiment of the present invention, through Figure 4 the compression method and the above computing unit structure shown, the compressed feature matrix after compression can be used. Without decompression, it can directly read and perform matrix multiplication and addition operations in sequence from the memory (such as a double data rate synchronous dynamic random access memory). Since the data has no line breaks and selective reading, the read-write bandwidth of the memory can be maximized, and the data only needs to be read once to complete all matrix calculations, achieving a reduction in the number of memory accesses and further improving the efficiency of data processing and calculation.
[0069] It can be understood that the above has described how to compress the sparse feature matrix. Next, the reading method of the feature matrix will be described.
[0070] According to an embodiment of the present invention, reading the compressed feature matrix and the state feature matrix includes: based on the storage address information, sequentially reading the elements in the compressed feature matrix from the first storage unit. When the read tag value is the end-of-row tag value, the target element and the target row information are obtained. The end-of-row tag value is used to indicate that the element is the last non-zero element in the target row information; reading the state feature matrix from the second storage unit based on the storage address information and the mapping relationship. The mapping relationship represents the corresponding relationship between the state features in the state feature matrix and the target element.
[0071] In an embodiment of the present invention, the first storage unit can be used to store graph structure data, the tag values corresponding to the elements, and the compressed feature matrix after compression. The intermediate matrix obtained during the matrix operation on the compressed feature matrix can be stored in the second storage unit. For example, the second storage unit can be used to store the state feature matrix corresponding to the LSTM algorithm, which can include the memory cell state matrix and the hidden state matrix.
[0072] For example, based on the storage address information, sequentially reading elements in the compressed feature matrix from the first storage unit may include: setting an index or pointer in the first storage unit to point to the starting position of the data pointer array; sequentially reading the values of non-zero elements according to the addresses in the data pointer array, and when reading each element, detecting whether it is an end-of-line tag value. The end-of-line tag value is usually a specific marker (such as 1) used to indicate the end of a row; in the case of detecting the end-of-line tag value, record the value of the current element and its row information.
[0073] According to an embodiment of the present invention, by reading elements in the compressed feature matrix based on the storage address information and obtaining the information of the current element and its row when the end-of-line tag value is read, large-scale sparse data can be effectively processed and analyzed.
[0074] According to an embodiment of the present invention, the processing method further includes: when the target matrix is obtained, clearing the operation caches corresponding to the first matrix operation and the second matrix operation in the internal storage unit of the accelerator, and performing the matrix operation of the next cycle, where the internal storage unit is the on-chip storage unit of the accelerator.
[0075] In an embodiment of the present invention, after a matrix operation cycle is completed, there are some intermediate calculation results and temporary data in the internal storage unit of the processor. The manner of clearing the operation caches corresponding to the first matrix operation and the second matrix operation in the internal storage unit of the accelerator can be specifically determined according to the actual situation.
[0076] For example, it can be achieved by resetting or clearing the address space of the internal storage unit, or directly writing new data to overwrite the old data before the start of a new operation cycle, or by sending a specific instruction to the accelerator to instruct it to clear the internal storage unit, so as to perform the matrix operation of the next cycle. After clearing the old operation cache, the new matrix data can be loaded into the internal storage unit, and the accelerator performs operations on the new matrix data according to a preset algorithm (such as matrix multiplication operation).
[0077] According to an embodiment of the present invention, the state feature matrix includes a hidden state matrix, and the algorithm of the accelerator is a gating algorithm, which includes a first gating algorithm and a second gating algorithm; performing a first matrix operation on the compressed feature matrix and the weight matrix to obtain a first intermediate matrix, and performing a second matrix operation on the state feature matrix and the first intermediate matrix to obtain a second intermediate matrix, including: determining a first moment for parallel operation, and determining a second moment for the second matrix operation based on the first moment, and determining a third moment for delayed operation based on the second moment; performing a parallel operation on the compressed feature matrix and the weight matrix based on the first gating algorithm at the first moment to obtain a plurality of parallel intermediate matrices; obtaining the second intermediate matrix by using the plurality of parallel intermediate matrices and the hidden state matrix at the second moment, and performing a delayed operation on the compressed feature matrix and the weight matrix based on the second gating algorithm at the third moment to obtain a delayed intermediate matrix, so that the time information for obtaining the second intermediate matrix is the same as the time information for obtaining the delayed intermediate matrix.
[0078] In an embodiment of the present invention, considering the gating algorithm based on the long short-term memory network, matrix operations need to be calculated in a certain order and cannot achieve full parallelism. To solve this technical problem, the present invention combines the gating algorithm of the long short-term memory network, the calculation characteristics of row multiplication, and the advantage of directly using the compressed feature matrix without decoding, and realizes corresponding matrix operations at multiple moments by using a single data delay unit, so as to align the intermediate calculation results of the matrix operations in time and meet the condition of calculating the next calculation data without caching.
[0079] In an embodiment of the present invention, based on the long short-term memory network calculation structure, a shift register can be used. By controlling the input delay of data, the time information of different intermediate matrix operation results can be made the same, so as to avoid using other data cache units. The ways to achieve single data input delay can include: determining a plurality of time information for different matrix operations, including a first moment for parallel operation and a second moment for the second matrix operation, so that a third moment for delayed operation can be determined based on the second moment. After determining different moments, the input delay of data can be achieved by always controlling, serial input and output, cascading use, reset and enable control.
[0080] For example, the clock frequency of the shift register can be adjusted by a clock signal to control the speed of data movement, thereby achieving the control of input delay. Or by introducing data bits one by one at the data input end, data can be shifted into the register at the required moment, thereby achieving precise control of the data input timing. For example, through reset control, the register content can be cleared at the target moment, and enable control can start or stop the data shift when needed.
[0081] According to an embodiment of the present invention, by controlling the input delay of single data to align the intermediate results in time, it is possible to ensure data pipelining calculation without caching the intermediate result matrix, improve the calculation efficiency, reduce the overall cache requirement, and the required delay unit will not increase the resource loss as the calculation parallelism increases.
[0082] According to an embodiment of the present invention, the first gating algorithm includes a first row multiplication operation and a first addition operation. The weight matrix includes weight row information corresponding to non-zero elements, and the weight row information includes a plurality of weight values; at the first moment, the first gating algorithm is used to perform a parallel operation on the compressed feature matrix and the weight matrix to obtain a plurality of parallel intermediate matrices, including: at the first moment, performing a first row multiplication operation on at least one non-zero element and the weight row information to obtain a plurality of first intermediate eigenvalues; performing a first addition operation on the plurality of first intermediate eigenvalues to obtain a plurality of parallel intermediate matrices.
[0083] In an embodiment of the present invention, the first gating algorithm may include an input gate algorithm, a forget gate algorithm, and a related state algorithm in a long short-term memory network. At the first moment, a first row multiplication operation may be performed on non-zero elements in the compressed feature matrix and the weight information to obtain a plurality of first intermediate eigenvalues, and then a first addition operation is performed on the plurality of first intermediate eigenvalues to obtain a plurality of parallel intermediate matrices. For example, for the input gate algorithm, at time t, multiply h (t−1) with the weight information W hi , and at the same time multiply x t with the weight information W ii to perform the first multiplication operation to obtain a plurality of first intermediate eigenvalues. On this basis, perform a first addition operation on W ii x t , W hi h (t−1) and other intermediate eigenvalues or the dense feature matrix to obtain the intermediate input matrix i t . At the same time, based on the corresponding algorithm, the intermediate forget matrix f t and the intermediate state matrix g t corresponding to the forget gate algorithm and the candidate cell state algorithm can be obtained respectively, and then the parallel intermediate matrices are obtained.
[0084] According to an embodiment of the present invention, the parallel intermediate matrices include an intermediate input matrix, an intermediate forget matrix, and an intermediate state matrix; at the second moment, using the plurality of parallel intermediate matrices and the hidden state matrix to determine the second intermediate matrix, including: at the second moment, performing a multiplication operation on the intermediate input matrix, the intermediate forget matrix, the intermediate state matrix, and the hidden state matrix to obtain the second intermediate matrix.
[0085] In an embodiment of the present invention, the second moment may be the same moment as a certain moment when the parallel intermediate matrix is calculated. After obtaining the intermediate input matrix, the intermediate forgetting matrix, and the intermediate state matrix, the intermediate input matrix and the intermediate state matrix may be subjected to a matrix multiplication operation to obtain a first result, and the obtained hidden state matrix and the intermediate forgetting matrix may be subjected to a matrix multiplication operation to obtain a second result, so as to perform a matrix addition operation on the first result and the second result to obtain a second intermediate matrix.
[0086] For example, multiply the intermediate input matrix i t by the intermediate state matrix g t to perform a matrix multiplication operation to obtain a first result i t g t , and multiply the hidden state matrix c (t−1) by the intermediate forgetting matrix f t to perform a matrix multiplication operation to obtain a second result f t c (t−1) .
[0087] According to an embodiment of the present invention, the second gating algorithm includes a second row multiplication operation and a second addition operation; performing a delayed operation on the compressed feature matrix and the weight matrix based on the second gating algorithm at the third moment to obtain a delayed intermediate matrix includes: performing a second row multiplication operation on at least one non-zero element and the weight row information at the third moment to obtain a plurality of second intermediate eigenvalues; performing a second addition operation on the plurality of second intermediate eigenvalues to obtain a delayed intermediate matrix.
[0088] In an embodiment of the present invention, the second gating algorithm may represent an output gate algorithm. After performing a second row multiplication operation on the non-zero element and the weight row information by the output gate algorithm and the single data delay unit to obtain a plurality of second intermediate eigenvalues, a second addition operation is performed, and then a delayed intermediate matrix is obtained, so that the moment of obtaining the delayed intermediate matrix is the same as the moment of obtaining the second intermediate matrix.
[0089] According to an embodiment of the present invention, by making the moment of obtaining the delayed intermediate matrix the same as the moment of obtaining the second intermediate matrix based on the gating algorithm and the single data delay unit, it is possible to ensure the time synchronization when data flows between computing units, thereby reducing the waiting time and improving the overall computing efficiency. Since there is no need for an additional cache to store intermediate results, the demand for storage resources can be reduced, thereby reducing the cost and power consumption of the system. The single data delay unit helps to optimize the data flow, making the data flow more smoothly during the calculation process and reducing the bottleneck of data transmission. This design makes the system easier to expand. Each computing unit can independently process data without relying on a centralized cache mechanism. By reducing the dependence on the cache, system errors caused by cache consistency problems can be reduced, further improving the reliability of the system.
[0090] According to an embodiment of the present invention, the graph neural network accelerator includes an arithmetic array, and the arithmetic array includes a plurality of arithmetic units arranged in a one-dimensional direction, and the number of arithmetic units is greater than or equal to the number of columns included in the weight matrix.
[0091] In an embodiment of the present invention, considering the characteristic that the compressed feature matrix in the present invention has less restriction on the matrix length during the matrix operation process, it only needs to meet the storage requirements of the weight matrix corresponding to the non-zero elements. By setting a plurality of arithmetic units in the accelerator according to the actual situation, the arithmetic units are arranged in a one-dimensional direction, and the number of arithmetic units is designed to be greater than or equal to the number of columns included in the weight matrix. In this way, it can be ensured that when performing matrix multiplication operations, the arithmetic array can process all columns of the weight matrix at one time, thereby improving the calculation efficiency.
[0092] In a feasible embodiment, the graph structure data processing method may further include: storing the weight matrix column by column in the internal memory; parallelly reading the target weights at the target positions in the column information stored in the internal memory to determine the weight rows corresponding to the features.
[0093] According to an embodiment of the present invention, by reusing the weight matrix on the processor, the reuse of the weight matrix is maximized, the number of memory accesses can be reduced, the data reuse rate can be improved, and the power consumption can be greatly reduced. At the same time, the reuse of the weight matrix helps to optimize the data stream, making the data flow more smoothly during the calculation process and reducing the bottleneck of data transmission.
[0094] Figure 5A An example block diagram of a graph neural network accelerator according to an embodiment of the present invention is shown.
[0095] As Figure 5A shown, the graph neural network accelerator 500A of this embodiment includes: a memory 51; a processor 52 configured to execute the above-mentioned graph structure data processing method according to the instructions stored in the memory 51.
[0096] The memory 51 may store different types of instructions and data. The instructions may include read instructions, data input delay instructions, and arithmetic instructions. The data may include graph structure data, sparse feature matrices, compressed feature matrices, state feature matrices, and target matrices obtained through operations.
[0097] The processor 52 may include a processing unit 521 and a storage unit 522. The processing unit 521 is used to perform corresponding data processing according to various instructions. The storage unit 522 may be an on-chip storage unit integrated on the processing unit 521, and the storage unit 522 may also be used to store the weight matrix and the parameters of the graph neural network model.
[0098] For example, the processor 52 may execute the graph structure data processing method according to the instructions stored in the memory 51. For example, the processor 52 performs matrix operations using a multiplier and an adder based on a gating algorithm.
[0099] Figure 5B FIG. shows a schematic diagram of the hardware architecture of an accelerator according to an embodiment of the present invention.
[0100] As Figure 5B shown, in 500B, the processor 52 inputs a data input delay instruction according to the data stored in the storage unit 522, and sends a first operation instruction to the forget gate operation unit 5211, the input gate operation unit 5212, and the first state operation unit 5213 in the processing unit 521 based on the first gating algorithm at the first moment; sends a second operation instruction to the second state operation unit 5214 in the processing unit 521 at the second moment, and sends a third operation instruction to the output gate operation unit in the processing unit 521 based on the delay unit 523 at the third moment based on the second gating algorithm, so as to obtain an intermediate operation result at the same moment for the calculation of the third state operation unit 5215, obtain an operation result, and output the operation result to a Double Data Rate SDRAM (DDR). According to actual requirements, the storage unit 522 may include multiple storage units that meet different transmission rates, such as storage unit 1, storage unit 2, storage unit 3, and storage unit 4.
[0101] According to an embodiment of the present disclosure, by inputting a data input delay instruction according to the data stored in the storage unit 522, the operation process meets the sequential calculation requirements of matrix operations, and there is no need to use an internal cache, directly implementing cacheless pipelining calculation and avoiding calculation delay. Through the collaborative work of the memory 51 and the processor 52, the graph neural network accelerator can efficiently execute the graph structure data processing method, thereby realizing the accelerated operation of the graph neural network.
[0102] Based on the above graph structure data processing method, the present invention also provides a graph structure data processing device. The following will be combined with Figure 6 to describe this device in detail.
[0103] Figure 6 FIG. shows a block diagram of the graph structure data processing device according to an embodiment of the present invention.
[0104] As Figure 6 shown, the graph structure data processing device 600 of this embodiment includes a reading module 610, an operation module 620, and a task execution module 630.
[0105] A reading module 610 is configured to read a compressed feature matrix and a status feature matrix. The compressed feature matrix is obtained by compressing the row position information and numerical information of at least one element in the sparse feature matrix based on a preset compression strategy. The status feature matrix characterizes the status features of the compressed feature matrix in preprocessing. In one embodiment, the reading module 610 may be configured to perform the operation S210 described above, which will not be elaborated herein.
[0106] An operation module 620 is configured to perform a first matrix operation on the compressed feature matrix and a weight matrix to obtain a first intermediate matrix, and perform a second matrix operation on the status feature matrix and the first intermediate matrix to obtain a second intermediate matrix. The first matrix operation includes parallel operation and latency operation, and the first intermediate matrix includes a parallel intermediate matrix and a latency intermediate matrix. In one embodiment, the operation module 620 may be configured to perform the operation S220 described above, which will not be elaborated herein.
[0107] A task execution module 630 is configured to obtain a target matrix based on the first intermediate matrix and the second intermediate matrix, so as to perform a target task using the target matrix. In one embodiment, the task execution module 630 may be configured to perform the operation S230 described above, which will not be elaborated herein.
[0108] According to an embodiment of the present invention, based on the reading module 610, the operation module 620, and the task execution module 630 in the graph structure data processing device, a compressed feature matrix is obtained by compressing the row position information and numerical information of elements in the sparse feature matrix, and then matrix operations are performed on the compressed compressed feature matrix, the weight matrix, and the status feature matrix. The operation process comprehensively considers the sparsity of the sparse feature matrix in the graph structure data and the density of other matrices (including the weight matrix, the status feature matrix, and the first intermediate matrix), and is compatible with the calculation characteristics of sparse matrices and dense matrices. Since the first intermediate matrix is obtained by performing parallel operation and latency operation on the compressed feature matrix and the weight matrix based on multiple gating algorithms of the accelerator, the first intermediate matrix is dynamically generated according to requirements, realizing cacheless pipelined calculation, avoiding the occupation of storage space by intermediate results, and further improving the calculation efficiency and the utilization rate of system resources.
[0109] According to an embodiment of the present invention, the features of an element include at least one of a full-zero row feature and a non-full-zero row feature. The non-full-zero row feature includes at least one non-zero element. The apparatus further includes: a label value determination module, configured to, when the feature of the current element is a full-zero row feature or a non-full-zero row feature, determine a label value corresponding to the full-zero row feature or a label value corresponding to the non-zero element based on a label correspondence relationship, where the label correspondence relationship characterizes the relationship between multiple label values and the full-zero row feature and the non-zero element.
[0110] According to an embodiment of the present invention, the tag value includes a first tag value and a second tag value; the tag value determination module includes: a second tag value determination sub-module and a first tag value determination sub-module. The second tag value determination sub-module is configured to determine, when the feature of the current element is the all-zero row feature, that the tag value corresponding to the all-zero row feature is the second tag value; the first tag value determination sub-module is configured to determine, when the feature of the current element is a non-all-zero row feature, that the tag value corresponding to the non-zero element is the first tag value based on the numerical information and the row position information.
[0111] According to an embodiment of the present invention, the apparatus further includes: a zero element processing module, configured to process the zero elements in the non-all-zero row feature and determine a compressed feature matrix based on the first tag value and the second tag value.
[0112] According to an embodiment of the present invention, the zero element processing module includes: an element deletion sub-module and an element conversion sub-module. The element deletion sub-module is configured to delete the zero elements in the current row when the feature of the current element is a non-all-zero row feature; the element conversion sub-module is configured to sequentially convert the non-zero elements into the first tag value and convert the non-all-zero row feature into the second tag value to obtain a compressed feature matrix.
[0113] According to an embodiment of the present invention, the reading module 610 includes: an element reading sub-module and a matrix reading sub-module. The element reading sub-module is configured to sequentially read the elements in the compressed feature matrix from the first storage unit based on the storage address information, and obtain a target element and target row information when the read tag value is a row end tag value, where the row end tag value is used to indicate that the element is the last non-zero element in the target row information; the matrix reading sub-module is configured to read a state feature matrix from the second storage unit based on the storage address information and a mapping relationship, where the mapping relationship represents the correspondence between the state features in the state feature matrix and the target element.
[0114] According to an embodiment of the present invention, the processing apparatus further includes: a clearing module, configured to, when a target matrix is obtained, clear the operation caches corresponding to the first matrix operation and the second matrix operation in the internal storage unit of the accelerator and perform the matrix operation in the next cycle, where the internal storage unit is the on-chip storage unit of the accelerator.
[0115] According to an embodiment of the present invention, the state feature matrix includes a hidden state matrix, and the algorithm of the accelerator is a gating algorithm, where the gating algorithm includes a first gating algorithm and a second gating algorithm; the operation module 620 includes: a time determination sub-module, configured to determine a first time for parallel operation, and based on the first time, determine a second time for a second matrix operation, and based on the second time, determine a third time for a delay operation; a parallel operation sub-module, configured to perform a parallel operation on the compressed feature matrix and the weight matrix based on the first gating algorithm at the first time to obtain a plurality of parallel intermediate matrices; a delay operation sub-module, configured to use the plurality of parallel intermediate matrices and the hidden state matrix to obtain a second intermediate matrix at the second time, and perform a delay operation on the compressed feature matrix and the weight matrix based on the second gating algorithm at the third time to obtain a delay intermediate matrix, such that the time information for obtaining the second intermediate matrix is the same as the time information for obtaining the delay intermediate matrix.
[0116] According to an embodiment of the present invention, the first gating algorithm includes a first row multiplication operation and a first addition operation, the weight matrix includes weight row information corresponding to non-zero elements, and the weight row information includes a plurality of weight values; the parallel operation sub-module includes: a row multiplication operation unit and an addition operation unit. The row multiplication operation unit is configured to perform a first row multiplication operation on at least one non-zero element and the weight row information at the first time to obtain a plurality of first intermediate eigenvalues; the addition operation unit is configured to perform a first addition operation on the plurality of first intermediate eigenvalues to obtain a plurality of parallel intermediate matrices.
[0117] According to an embodiment of the present invention, the parallel intermediate matrices include an intermediate input matrix, an intermediate forgetting matrix, and an intermediate state matrix; the delay operation sub-module includes: a multiplication operation unit, configured to perform a multiplication operation on the intermediate input matrix, the intermediate forgetting matrix, the intermediate state matrix, and the hidden state matrix at the second time to obtain a second intermediate matrix.
[0118] According to an embodiment of the present invention, the second gating algorithm includes a second row multiplication operation and a second addition operation; the delay operation sub-module further includes: a second row multiplication operation unit and a second addition operation unit. The second row multiplication operation unit is configured to perform a second row multiplication operation on at least one non-zero element and the weight row information at the third time to obtain a plurality of second intermediate eigenvalues; the second addition operation unit is configured to perform a second addition operation on the plurality of second intermediate eigenvalues to obtain a delay intermediate matrix.
[0119] According to an embodiment of the present invention, the graph neural network accelerator includes an operation array, and the operation array includes a plurality of operation units arranged in a one-dimensional direction, and the number of operation units is greater than or equal to the number of columns included in the weight matrix.
[0120] According to an embodiment of the present invention, any multiple of the reading module 610, the operation module 620, and the task execution module 630 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the reading module 610, the operation module 620, and the task execution module 630 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the reading module 610, the operation module 620, and the task execution module 630 may be at least partially implemented as a computer program module, which can execute corresponding functions when the computer program module is run.
[0121] Figure 7 The block diagram of an electronic device suitable for implementing the graph structure data processing method according to an embodiment of the present invention is shown.
[0122] As Figure 7 shown, the electronic device 700 according to an embodiment of the present invention includes a first processor 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage part 708 into the random access memory (RAM) 703. The first processor 701 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related processor group, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The first processor 701 may also include on-board memory for caching purposes. The first processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0123] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The first processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The first processor 701 performs various operations of the method flow according to an embodiment of the present invention by executing programs in the ROM 702 and / or the RAM 703. It should be noted that the programs may also be stored in one or more memories other than the ROM 702 and the RAM 703. The first processor 701 may also perform various operations of the method flow according to an embodiment of the present invention by executing programs stored in the one or more memories.
[0124] According to an embodiment of the present invention, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A driver 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 710 as needed so that a computer program read from it can be installed into the storage portion 708 as needed.
[0125] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to an embodiment of the present invention is implemented.
[0126] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the ROM 702 and / or RAM 703 and / or ROM 702 and RAM 703 described above.
[0127] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the graph structure data processing method provided by the embodiment of the present invention.
[0128] When the computer program is executed by the first processor 701, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0129] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0130] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or be installed from the removable medium 711. When the computer program is executed by the first processor 701, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0131] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0133] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0134] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A method for processing graph-structured data, applied to a graph neural network accelerator. The graph neural network accelerator includes a processor and a memory. The graph neural network accelerator is a processor that processes graph data structures using a preset algorithm. The graph-structured data includes a sparse feature matrix corresponding to the graph-structured data to be processed. The graph-structured data includes at least one of physical examination data, interaction sequences, social relationship data, behavioral time series, and biological sequences. It is characterized in that The method includes: Using a processor to read a compressed feature matrix and a status feature matrix from a memory, where the compressed feature matrix is obtained by compressing at least one of the row position information and numerical information of the sparse feature matrix based on a preset compression strategy, and the status feature matrix represents the status feature of the compressed feature matrix in preprocessing; Performing a first matrix operation on the compressed feature matrix and a weight matrix to obtain a first intermediate matrix, and performing a second matrix operation on the status feature matrix and the first intermediate matrix to obtain a second intermediate matrix, including: Determining a first time for parallel operation, determining a second time for the second matrix operation based on the first time, and determining a third time for delayed operation based on the second time, where the weight matrix includes weight row information corresponding to non-zero elements; Performing a first row multiplication operation in the first gating algorithm on at least one non-zero element and the weight row information at the first time to obtain a plurality of first intermediate eigenvalues; performing a first addition operation in the first gating algorithm on the plurality of first intermediate eigenvalues to obtain a plurality of parallel intermediate matrices; Obtaining a second intermediate matrix using the plurality of parallel intermediate matrices and a hidden state matrix at the second time, and performing a delayed operation on the compressed feature matrix and the weight matrix based on the second gating algorithm at the third time to obtain a delayed intermediate matrix, such that the time information for obtaining the second intermediate matrix is the same as the time information for obtaining the delayed intermediate matrix, where the first matrix operation includes parallel operation and delayed operation, and the first intermediate matrix includes parallel intermediate matrices and delayed intermediate matrices; Obtaining a target matrix according to the first intermediate matrix and the second intermediate matrix for using the target matrix to perform a target task; Wherein, the features of the element include at least one of all-zero row feature and non-all-zero row feature, and the non-all-zero row feature includes at least one non-zero element; The method further includes: When the feature of the current element is an all-zero row feature or a non-all-zero row feature, determining a label value corresponding to the all-zero row feature or determining a label value corresponding to the non-zero element based on a label correspondence relationship, where the label correspondence relationship represents the relationship between a plurality of label values and the all-zero row feature and non-zero elements.
2. The processing method according to claim 1, wherein The label value includes a first label value and a second label value; The determining a label value corresponding to the all-zero row feature or determining a label value corresponding to the non-zero element based on the label correspondence relationship includes: When the feature of the current element is the all-zero row feature, determining the label value corresponding to the all-zero row feature as the second label value; When the feature of the current element is the non-all-zero row feature, determining the label value corresponding to the non-zero element as the first label value based on the numerical information and the row position information.
3. The processing method according to claim 2, characterized in that The processing method further includes: Processing zero elements in the non-all-zero row feature and determining the compressed feature matrix based on the first label value and the second label value.
4. The processing method according to claim 3, characterized in that The processing zero elements in the non-all-zero row feature and determining the compressed feature matrix based on the first label value and the second label value includes: When the feature of the current element is the non-all-zero row feature, deleting the zero elements in the current row; Convert the non-zero elements to the first tag value in sequence, and convert the non-all-zero row features to the second tag value, to obtain the compressed feature matrix.
5. The processing method according to any one of claims 1 to 4, characterized in that The reading of the compressed feature matrix and the state feature matrix includes: Based on the storage address information, sequentially read the elements in the compressed feature matrix from the first storage unit. When the read tag value is the end-of-row tag value, obtain the target element and the target row information, where the end-of-row tag value is used to indicate that the element is the last non-zero element in the target row information; Read the state feature matrix from the second storage unit based on the storage address information and the mapping relationship, where the mapping relationship represents the correspondence between the state features in the state feature matrix and the target element.
6. The processing method according to claim 1, characterized in that The processing method further includes: When the target matrix is obtained, clear the operation caches corresponding to the first matrix operation and the second matrix operation in the internal storage unit of the accelerator, and perform the matrix operation in the next cycle, where the internal storage unit is the on-chip storage unit of the accelerator.
7. The processing method according to claim 1, characterized in that The parallel intermediate matrix includes an intermediate input matrix, an intermediate forgetting matrix, and an intermediate state matrix; Determining the second intermediate matrix using the plurality of parallel intermediate matrices and the hidden state matrix at the second moment includes: Performing a multiplication operation on the intermediate input matrix, the intermediate forgetting matrix, the intermediate state matrix, and the hidden state matrix at the second moment to obtain the second intermediate matrix.
8. The processing method according to claim 1, characterized in that, The second gating algorithm includes a second row multiplication operation and a second addition operation; Performing the delay operation on the compressed feature matrix and the weight matrix based on the second gating algorithm at the third moment to obtain the delayed intermediate matrix includes: Performing the second row multiplication operation on at least one of the non-zero elements and the weight row information at the third moment to obtain a plurality of second intermediate feature values; Performing the second addition operation on the plurality of second intermediate feature values to obtain the delayed intermediate matrix.
9. The processing method according to claim 1, characterized in that The graph neural network accelerator includes an operation array, and the operation array includes a plurality of operation units arranged in a one-dimensional direction, and the number of operation units is greater than or equal to the number of columns included in the weight matrix.
10. A graph neural network accelerator, characterized in that, Including: A memory; A processor configured to execute the method according to any one of claims 1 to 9 according to the instructions stored in the memory.
11. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.
12. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Compression LSTM accelerator and acceleration method based on FPGA
CN113222133A
Graph neural network acceleration method and graph neural network acceleration structure
CN118690803A