Cache access method and related graph neural network system
Patent Information
- Application Number
- CN202111373408.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-11-19
AI Technical Summary
因此,在图神经网络系统中,使用通用的高速缓存访问方法无法得到很高的效率,造成整体训练及推理时间大幅地增加
[0006]本申请所公开的高速缓存的访问方法以及相关图神经网络系统能提升图神经网络系统中对高速缓存的访问效率,进而降低整体训练时间。
Smart Images

Figure CN116151338B_ABST
Abstract
Description
Technical Field
[0001] This application relates to a cache, and more particularly to a cache access method and a related graph neural network system. Background Technology
[0002] When training a Graph Neural Network (GNN), the access process is highly discrete and random. Therefore, using general-purpose cache access methods in GNN systems is inefficient, significantly increasing overall training and inference time. Thus, how to plan the cache in a GNN system and optimize its access methods has become a pressing issue in this field. Summary of the Invention
[0003] One of the purposes of this application is to disclose a cache access method and a related graph neural network system to solve the above problems.
[0004] One embodiment of this application discloses a cache access method. The cache is used to reduce the average time of a graph neural network processor accessing memory. The graph neural network processor is used to perform operations on a graph neural network, which is stored in the memory in a compressed sparse row format. The method includes: receiving the address corresponding to a node in the graph neural network and the type of the address; when the type is one of a first type and a second type, performing a search based on the marker column of the address comparison degree lookup table to obtain at least the degree of the node, wherein the degree is the number of edges of the node; determining whether the degree is greater than a preset value and obtaining a determination result; and based on the determination result, deciding whether to perform a search in the cache in a region corresponding to the type, wherein the cache includes at least a first region and a second region, wherein the first region corresponds to the first type and the second region corresponds to the second type, the first region is used to store edge-related information, and the second region is used to store attribute-related information; wherein the first type indicates that the desired information is edge-related information of the node, and the second type indicates that the desired information is attribute-related information of the node.
[0005] One embodiment of this application discloses a graph neural network system, including: a graph neural network processor for executing the above-described method; a degree lookup table for searching based on the labeled field; a cache including at least a first region and a second region, wherein the first region corresponds to the first type and the second region corresponds to the second type, the first region is used to store edge-related information and the second region is used to store attribute-related information; and the memory.
[0006] The cache access method and related graph neural network system disclosed in this application can improve the cache access efficiency in graph neural network systems, thereby reducing the overall training time. Attached Figure Description
[0007] Figure 1 This is a schematic diagram of a graph neural network stored in memory in a compressed sparse row format.
[0008] Figure 2 This is a schematic diagram of an embodiment of the graph neural network system of this application.
[0009] Figure 3 This is a schematic diagram illustrating an embodiment of the degree lookup table, region lookup table, and cache in the graph neural network system of this application.
[0010] Figure 4 This is a flowchart illustrating an embodiment of the cache access method of this application.
[0011] Figure 5 This is a flowchart of an embodiment of the cache access method of this application, specifically for a method of type first.
[0012] Figure 6 This is a flowchart of an embodiment of the cache access method of this application, specifically for the second type of method.
[0013] Figure 7 This is a flowchart of an embodiment of the cache access method of this application, specifically for a method of type three. Detailed Implementation
[0014] The following disclosure provides various implementations or examples that can be used to achieve different features of this disclosure. Specific examples of components and configurations described below are for simplification purposes. It is understood that these descriptions are illustrative only and are not intended to limit the scope of this disclosure. For example, in the following description, forming a first feature on or over a second feature may include, in some embodiments, the first and second features being in direct contact with each other; and may also include, in some embodiments, additional components being formed between the first and second features, such that the first and second features may not be in direct contact. Furthermore, component symbols and / or reference numerals may be reused in multiple embodiments of this disclosure. Such reuse is for the purpose of brevity and clarity and does not in itself represent a relationship between the different embodiments and / or configurations discussed.
[0015] Furthermore, the use of spatially relative terms, such as "below," "below," "lower than," "above," "above," and similar terms, may be for the convenience of describing the relationship between one component or feature depicted in the figure and one or more other components or features. These spatially relative terms, in addition to the orientation shown in the figure, also encompass various different orientations of the device during use or operation. The device may be placed in other orientations (e.g., rotated 90 degrees or in other orientations), and these spatially relative descriptive terms should be interpreted accordingly.
[0016] While the numerical ranges and parameters used to define the broader scope of this application are approximate values, the relevant values in the specific embodiments have been presented as precisely as possible. However, any value inevitably contains standard deviations due to individual test methods. Here, "identical" generally means that the actual value is within plus or minus 10%, 5%, 1%, or 0.5% of a particular value or range. Alternatively, the term "identical" means that the actual value falls within the acceptable standard error of the mean, as determined by those skilled in the art to which this application pertains. It is understood that, except for experimental examples, or unless explicitly stated otherwise, all ranges, quantities, values, and percentages used herein (e.g., used to describe material usage, duration, temperature, operating conditions, quantity ratios, and others similar) are modified with "identical". Therefore, unless otherwise stated, the numerical parameters disclosed in this specification and the accompanying claims are approximate values and are subject to change as needed. At a minimum, these numerical parameters should be understood as the indicated significant digits and values obtained by applying general rounding. In this context, a range of values is expressed as a distance from one endpoint to the other or between the two endpoints; unless otherwise stated, all ranges of values herein include the endpoints.
[0017] Figure 1 This diagram illustrates a Graph Neural Network (GNN) stored in memory using compressed sparse row format (CSR). A GNN can contain multiple nodes, each determining its neighbors based on its corresponding edges. Furthermore, each node possesses attribute information. To obtain the complete data of a GNN in memory, a method must first be established to retrieve the complete data of any node and all its first-order neighbors in memory, and this method will serve as the basis for obtaining the complete data of the GNN. Under the storage condition of compressed sparse row format, the above-mentioned basic method consists of three steps.
[0018] In the first step, the system first determines which node to retrieve data from based on the requirements, i.e., it determines the "index" of the node to be retrieved (hereinafter referred to as the root node). Based on the "index," the value of the "offset" corresponding to the "index" and the value of the next "offset" can be obtained in column 102. Next, in the second step, based on the "offset" of the "index," a starting position can be obtained in column 104. This starting position is used to indicate the starting position of the "edge" information of the root node in column 104. In other words, the value stored in column 102 points to a specific position in column 104. Specifically, column 102 continuously stores the "index" information of all adjacent nodes of the root node, starting from the starting position.
[0019] Since the number of adjacent nodes is unknown, in accordance with the principles of compressed sparse line formatting, the next starting position needs to be determined based on the value of the next "offset". This ensures that the data before the next starting position consists of the "index" information of the adjacent nodes of the root node. Finally, in the third step, based on the "index" information of all adjacent nodes in column 104, the "attributes" of each adjacent node can be obtained in column 106.
[0020] For example, to obtain Figure 1 The attributes of all adjacent nodes of node ② (index ②) are stored in column 102. Since node ② is the root node, the offset of node ② is 2. In column 102, the next offset after offset 2 is 5. Therefore, in column 104, the edges of node ② are stored starting at position 2, and edges starting at position 5 are not. Thus, the information from position 2 to position (5-1) contains the edges of node ②, including all edges of node ②, i.e., the indices of all adjacent nodes of node ②, namely nodes ①, ③, and ④. In column 106, the attributes of node ① ("ABCDEFG"), node ③ ("BCBCBCB"), and node ④ ("ASDFGHJ") can be obtained.
[0021] Based on observations and extensive practical experience with graph neural networks, this application arrives at a general conclusion: in graph neural networks, the greater the number of edges (i.e., the number of first-order adjacent nodes, also known as "degree") of a node, the higher the probability of that node being accessed. Based on this principle, this application optimizes the planning and access methods of caches in graph neural network systems, and the details will be explained below.
[0022] Figure 2This is a schematic diagram of an embodiment of the graph neural network system of this application. The graph neural network system 200 includes a graph neural network processor 202, a cache 204, memory 206, a degree lookup table 208, and a region lookup table 210. The graph neural network processor 202 is used to perform operations on the graph neural network. The cache 204 is used to reduce the average time for the graph neural network processor 202 to access memory 206.
[0023] Figure 3 This is a schematic diagram illustrating an embodiment of the degree lookup table 208, region lookup table 210, and cache 204 in the graph neural network system 200 of this application. First, the cache 204 will be described. The cache 204 includes a first region 212, a second region 214, a third region 216, and a data array 218. The extent of each region can be determined by the graph neural network processor 202 or upper-layer software, and the extents of each region do not overlap. The first region 212, the second region 214, and the third region 216 each include a "label" field and a "index" field. The "label" field is used to provide a comparison function during lookup. The index is used to point to a specific location range in the data array 218.
[0024] The degree lookup table 208 includes a "tag" field, an "offset" field, and a "degree" field. The "tag" field provides a comparison function during the search, and specifically, the tag value corresponds to a node. The "offset" is related to the starting position of the "edge" of the corresponding tagged node in memory 206. The "degree" is the number of "edges" of the corresponding tagged node. In this embodiment, the eviction policy of the degree lookup table 208 is a combination of Least Recently Used (LRU) and the "degree" value; for example, the less recently used and the smaller the degree, the earlier it is deleted.
[0025] The region lookup table 210 contains a "type" field and a "region" field. In this embodiment, the "type" field contains three different "types": a first type, a second type, and a third type. The "region" field in the region lookup table 210 stores the specific region range of each region in the cache 204. The "region" corresponding to the first type is the first region 212 in the cache 204; the "region" corresponding to the second type is the second region 214 in the cache 204; and the "region" corresponding to the third type is the third region 216 in the cache 204. The first region 212 is used to store information related to "edges"; the second region 214 is used to store information related to "attributes"; and the third region 216 is used to store information related to coalesce "edges". As explained earlier regarding the compressed sparse line format, since the "edge" information is stored contiguously in memory, the third region 216 stores indicators of "edge" information read from memory 206 due to spatial locality.
[0026] In this embodiment, the eviction policy for the first region 212 and the second region 214 is Least Recently Used; the eviction policy for the third region 216 is Eviction after Use. For example, once all the nodes contained in the "edge" related information of a certain cache line have been accessed, the data in this cache line is deleted.
[0027] When the graph neural network processor 202 wants to obtain data, it issues a request containing information 201 and information 203. Information 201 contains "type" information, and information 203 contains "address" information. In this embodiment, the "type" information is used to search for the "type" field in the region lookup table 210. When the "type" is the first type, it indicates a desire to obtain the "edge" of the node corresponding to the "address"; when the "type" is the second type, it indicates a desire to obtain the "attribute" of the node corresponding to the "address"; details of the third type will be explained later.
[0028] Figure 4 This is a flowchart illustrating an embodiment of the cache access method of this application. In method 400, firstly, in step 402, the "address" (i.e., information 203) corresponding to a node in the graph neural network and the "type" (i.e., information 201) of the "address" are received from the graph neural network processor 202.
[0029] Next, in step 404, the subsequent operation is determined based on the "type". Generally speaking, when the "type" is either the first type or the second type, a search is performed based on the "mark" field in the "address" comparison lookup table 208 to obtain at least the "degree" of the root node. The essence of this application is that in step 406, it is determined whether the "degree" is greater than a preset value, and a determination result is obtained. That is, when the "degree" is greater than the preset value, it means that the root node has a higher probability of being accessed. Therefore, according to the eviction policy of the first region 212 and the second region 214, the probability that the relevant data of the root node is stored in the first region 212 and the second region 214 is higher; when the "degree" is not greater than the preset value, the probability that the relevant data of the root node is stored in the first region 212 and the second region 214 is lower. Therefore, based on this principle, in step 408, the determination result can determine whether to perform a search in the "region" corresponding to the "type" in the cache 204. For example, if it is determined that the probability of the root node's related data being stored in the "region" of the "type" in the cache 204 is low, directly accessing the memory 206 may yield a more efficient result.
[0030] Figures 5 to 7 The terms "type" for the first type, the second type, and the third type will be described in detail. Figure 5 The flowchart 500 illustrates a specific embodiment when the "type" is the first type.
[0031] First, in step 502, the graph neural network processor 202 receives the "address" (i.e., information 203) corresponding to a node in the graph neural network and the "type" (i.e., information 201) of the "address", wherein the "type" is the first type, indicating that the graph neural network processor 202 wants to obtain the information of the "edge" of the root node, that is, it wants to obtain the "index" of all first-order adjacent nodes of the root node.
[0032] In step 504, the search is performed based on the "mark" field in the "address" comparison lookup table 208.
[0033] In step 506, the result of the search in lookup table 208 is obtained. If lookup table 208 is hit, then proceed to step 508.
[0034] In step 508, the "degree" and "offset" of the root node are obtained from the degree lookup table 208.
[0035] In step 510, it is determined whether the "degree" is greater than a preset value. If so, proceed to step 512.
[0036] In step 512, the "offset" is compared with the "mark" field in the first region 212 of the cache 204 to perform a lookup. In some embodiments, before looking up the first region 212, the specific region range of the first region 212 in the cache 204 is determined using the region lookup table 210.
[0037] In step 514, the search result for the first region 212 is obtained. If the first region 212 is matched, then proceed to step 516.
[0038] In step 516, the first index corresponding to the "offset" is obtained from the "index" field of the first region 212.
[0039] In step 518, the information of the "edge" of the root node is read from the data array 218 according to the first index.
[0040] Please return to step 514. If the first region 212 is missing, proceed to step 520.
[0041] In step 520, memory 206 is accessed to obtain information about the "edges" of the root node.
[0042] Please return to step 510. If it is determined that the "degree" is not greater than the preset value, proceed to step 522.
[0043] In step 522, the "offset" is compared with the "mark" field in the third region 216 of the cache 204 to perform a lookup. In some embodiments, before looking up the third region 216, the specific region range of the third region 216 in the cache 204 is determined using the region lookup table 210.
[0044] In step 524, the result of searching the third region 216 is obtained. If the third region 216 is found, then proceed to step 526.
[0045] In step 526, the third indicator corresponding to the "offset" is obtained from the "Indicator" field of the third region 216.
[0046] In step 528, the information of the "edge" of the root node is read from the data array 218 according to the third index.
[0047] Please return to step 524. If the third region 216 is missing, proceed to step 520.
[0048] Please return to step 506. If lookup table 208 is missing, proceed to step 530.
[0049] In step 530, memory 206 is accessed to obtain the "offset" of the root node.
[0050] After obtaining the "indices" of all first-order neighboring nodes of the root node, the graph neural network processor 202 may also want to obtain the "attributes" of each first-order neighboring node. Therefore, it is necessary to use... Figure 6 Examples of implementations. Figure 6 The middle section illustrates a flowchart 600 of a specific embodiment when the "type" is the second type.
[0051] First, in step 602, the graph neural network processor 202 receives the "address" (i.e., information 203) corresponding to a node in the graph neural network and the "type" (i.e., information 201) of the "address", wherein the "type" is the second type, indicating that the graph neural network processor 202 wants to obtain the "attributes" of the root node.
[0052] In step 604, the search is performed based on the "mark" field in the "address" comparison lookup table 208.
[0053] In step 606, the result of searching lookup table 208 is obtained. If lookup table 208 is found, then proceed to step 608.
[0054] In step 608, the "degree" of the root node is obtained from the degree lookup table 208.
[0055] In step 610, it is determined whether the "degree" is greater than a preset value. If so, proceed to step 612.
[0056] In step 612, the lookup is performed by comparing the "address" with the "mark" field in the second region 214 of the cache 204. In some embodiments, before looking up the third region 216, the specific region range of the second region 214 in the cache 204 is determined using the region lookup table 210.
[0057] In step 614, the result of searching the second region 214 is obtained. If the second region 214 is matched, then proceed to step 616.
[0058] In step 616, a second indicator corresponding to the "address" is obtained from the "Indicator" field of the second region 214.
[0059] In step 618, the "attributes" of the root node are read from the data array 218 according to the second indicator.
[0060] Please return to step 614. If the second region 214 is missing, proceed to step 620.
[0061] In step 620, memory 206 is accessed to obtain the "attributes" of the root node.
[0062] Please return to step 610. If it is determined that the "degree" is not greater than the preset value, proceed to step 620.
[0063] Please return to step 606. If lookup table 208 is missing, proceed to step 620.
[0064] Please review Figure 5 If the degree lookup table 208 is missing, step 530 will be executed to directly access memory 206 to obtain the "offset" of the root node. The graph neural network processor 202 also requires an additional process to obtain the "edge" information of the root node, i.e., to obtain the "index" of all first-order adjacent nodes of the root node. Therefore, it needs to use... Figure 7 Examples of implementations. Figure 7 The middle section illustrates a flowchart 700 of a specific embodiment when the "type" is the third type.
[0065] First, in step 702, the graph neural network processor 202 receives the "address" (i.e., information 203) corresponding to a node in the graph neural network and the "type" (i.e., information 201) of the "address," where the "type" is the third type, indicating that the graph neural network processor 202 already has the "offset" of the root node and wants to obtain the information of the "edge" of the root node based on the "offset." At this time, the content of the "address" is the "offset" of the root node.
[0066] In step 704, the lookup is performed by comparing the "address" with the "mark" field in the third region 216 of the cache 204.
[0067] In step 706, the result of searching the third region 216 is obtained. If the third region 216 is matched, then proceed to step 708.
[0068] In step 708, a third index corresponding to the "address" is obtained from the "index" field of the third region 216. In some embodiments, before searching the third region 216, the specific region range of the third region 216 in the cache 204 is determined using the region lookup table 210.
[0069] In step 710, the information of the "edge" of the root node is read from the data array 218 according to the third indicator.
[0070] Please return to step 706. If the third region 216 is missing, proceed to step 712.
[0071] In step 712, memory 206 is accessed to obtain information about the "edges" of the root node.
[0072] It should be noted that the degree lookup table 208 and the region lookup table 210 of this application are located in other memory or caches besides the cache 204. Furthermore, in the embodiments of this application, the cache 204 is configured in a fully associative manner, but this application is not limited thereto.
[0073] The cache access methods 400 / 500 / 600 / 700 of this application and the associated graph neural network system 200 can improve the access efficiency of the cache 204 in the graph neural network system 200, thereby reducing the overall training time.
[0074] The foregoing description briefly outlines the features of certain embodiments of this application, enabling those skilled in the art to more fully understand the various forms of this disclosure. Those skilled in the art will readily recognize that this disclosure serves as a basis for designing or modifying other processes and structures to achieve the same objectives and / or advantages as the embodiments described herein. Those skilled in the art should understand that these equivalent embodiments remain within the spirit and scope of this disclosure, and various changes, substitutions, and modifications can be made without departing from the spirit and scope of this disclosure.
Claims
1. A method for accessing a cache, wherein the cache is used to reduce the average time for a graph neural network processor to access memory, the graph neural network processor being used to perform operations on a graph neural network, characterized in that, The graph neural network is stored in the memory in a compressed sparse row format, and the method includes: The system receives the address corresponding to a node in the graph neural network and the type of the address, wherein the address includes the offset of the node, and the offset is used to determine the position of the edge of the node in the memory; When the type is either the first type or the second type, a search is performed based on the marker field in the address comparison degree lookup table to obtain at least the degree of the node, wherein the degree is the number of edges of the node; Determine whether the degree is greater than a preset value and obtain the determination result; and Based on the judgment result, it is determined whether to perform a search in the region corresponding to the type in the cache, wherein the cache includes at least a first region, a second region, and a third region, wherein the first region corresponds to the first type, the second region corresponds to the second type, the first region is used to store edge-related information, the second region is used to store attribute-related information, and the third region is used to store clustered edge-related information. If the degree is greater than the preset value, a search is performed in the region corresponding to the type; if the degree is less than or equal to the preset value and the type is the first type, a search is performed in the third region; if the degree is less than or equal to the preset value and the type is the second type, a search is performed in memory. The first type indicates that the desired information is edge-related information of the node, and the second type indicates that the desired information is attribute-related information of the node.
2. The method as described in claim 1, characterized in that, The step of obtaining the degree of the node by comparing the address with the marker field in the degree lookup table, where the type is the first type, includes: When the degree lookup table is hit, the degree and offset of the node are obtained from the degree lookup table, wherein the offset is related to the starting position of the node's edge in memory.
3. The method as described in claim 2, characterized in that, Based on the judgment result, the step of determining whether to perform a search in the cache region corresponding to the type includes: When the judgment result indicates that the degree corresponding to the address is greater than the preset value, the cache is searched by comparing the offset with the marker field in the first region.
4. The method as described in claim 3, characterized in that, Also includes: When the first region is hit, a first index corresponding to the offset is obtained from the first region. The first index is used to point to the location in the data array of the cache where the edge information of the node is stored. as well as Based on the first indicator, the edge information of the node is read from the data array.
5. The method as described in claim 3, characterized in that, Also includes: When the first region is missing, the memory is accessed to obtain the edge information of the node.
6. The method as described in claim 2, characterized in that, Based on the judgment result, the step of determining whether to perform a search in the cache region corresponding to the type includes: When the determination result indicates that the degree corresponding to the address is not greater than the preset value, a search is performed in the cache based on the offset compared with the marker field in the third region; and The method further includes: When the third region is hit, a third index corresponding to the offset is obtained from the third region. This third index points to the location in the cached data array where the edge information of the node is stored. Based on the third indicator, the edge information of the node is read from the data array.
7. The method as described in claim 6, characterized in that, Also includes: When the third region is missing, access the memory to obtain the edge information of the node.
8. The method as described in claim 1, characterized in that, The type is the first type, and the method further includes: When the degree lookup table is missing, the memory is accessed to obtain the offset of the node, wherein the offset is related to the starting position of the node's edge in the memory.
9. The method as described in claim 1, characterized in that, The step of obtaining the degree of the node by comparing the address with the marker field in the degree lookup table, where the type is the second type, includes: When the degree lookup table is matched, the degree of the node is obtained from the degree lookup table.
10. The method as described in claim 9, characterized in that, Based on the judgment result, the step of determining whether to perform a search in the cache region corresponding to the type includes: When the determination result indicates that the degree corresponding to the address is greater than the preset value, a search is performed in the cache based on the address comparison with the marker field in the second region; and The method further includes: When the second region is hit, a second index corresponding to the address is obtained from the second region. This second index points to the location in the cache data array where the node's attributes are stored. Based on the second index, the node's attributes are read from the data array. When the second region is missing or when the judgment result indicates that the degree of the corresponding address is not greater than the preset value, the memory is accessed to obtain the attributes of the node.
11. The method as described in claim 1, characterized in that, The type is the second type, and the method further includes: When the degree lookup table is missing, the memory is accessed to obtain the attributes of the node.
12. The method as described in claim 1, characterized in that, The method further includes: When the type is the third type, the address is the offset of the node, and what is desired is the edge-related information of the node, wherein the offset is related to the starting position of the edge of the node in the memory; In the cache, the lookup is performed by comparing the address with the marker field in the third region; When the third region is hit, a third index corresponding to the address is obtained from the third region. The third index is used to point to the location in the data array of the cache where the edge information of the node is stored, and the edge information of the node is read from the data array based on the third index. When the third region is missing, access the memory to obtain the edge information of the node.
13. A graph neural network system, characterized in that, include: A graph neural network processor for performing the method as described in any one of claims 1 to 12; The degree lookup table is used to search based on the marked field; The cache includes at least a first region and a second region, wherein the first region corresponds to the first type and the second region corresponds to the second type. The first region is used to store edge-related information and the second region is used to store attribute-related information. as well as The memory; The first type indicates that the desired information is edge-related information of the node, and the second type indicates that the desired information is attribute-related information of the node.
14. The graph neural network system as described in claim 13, characterized in that, The cache also includes a third region, which is used to store information related to the clustered edges.
15. The graph neural network system as described in claim 14, characterized in that, It also includes a region lookup table, which contains a type field and a region field, wherein the type field includes the first type, the second type and the third type, and the region field includes the range of the first region in the cache, the range of the second region in the cache and the range of the third region in the cache, wherein the graph neural network processor uses the region lookup table to obtain the first region corresponding to the first type, the second region corresponding to the second type and the third region corresponding to the third type.
Citation Information
Patent Citations
Data storage method based on compressed graph, storage medium, storage device and server
CN110389953A
On-chip storage system and method for graph neural network application
CN111695685A