Sampler and apparatus for graph neural network model execution
By dividing the number of all neighbor nodes around the specified node into multiple equal sub-intervals and using random numbers to determine the neighbor nodes to be sampled, the complexity caused by the disordered state of random numbers and the uneven sampling results in the existing sampler is solved, and a more efficient and uniform sampling effect is achieved.
Patent Information
- Application Number
- CN202110231273.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-02
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-03-02
AI Technical Summary
When existing samplers randomly sample neighbor nodes of neural network models, due to the disordered state of random numbers, the storage and comparison complexity is caused, and the sampling results cannot ensure uniform distribution.
A sampler is designed to determine the neighbor node to be sampled by dividing the number of all neighbor nodes around a specified node into multiple equal sub-intervals and obtaining a third integer value within these sub-intervals using a random number.
The uniform sampling of the neighbor nodes around the specified node is achieved, reducing the complexity of storage and comparison, improving sampling efficiency, and reducing storage overhead.
Smart Images

Figure CN114997380B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and particularly to a sampler and a device for a graph neural network model. Background Art
[0002] In recent years, deep learning models have been widely developed in aspects such as image classification, speech recognition, natural language processing, etc., and a series of successful applications have been achieved. However, in more and more real-world scenarios, data is represented in the form of a graph (Graph). A graph not only contains entities but also the dependencies between entities, such as social networks, protein molecular structures, customer relationships on e-commerce platforms, and so on. Figure 1 is an exemplary social network relationship graph. As shown in the figure, the social network relationship graph includes entities and the dependencies between entities (the dependencies are represented by connecting edges in the figure). Such as Figure 1 The various graphs shown are called graph data. The entities on the graph are called nodes in the graph data, and the nodes directly or indirectly connected to a certain node are called the neighbor nodes of that node.
[0003] With the in-depth research of the industry on graph data and neural learning models, deep learning models for processing graph data have begun to emerge. Such deep learning models are collectively referred to as graph neural network (GNN) models.
[0004] In order to accelerate the execution of graph neural network models, the industry has designed dedicated hardware acceleration units for graph neural network models to execute graph neural network models. Such acceleration units are called graph neural network acceleration units. At the same time, the industry is also continuously improving graph neural network acceleration units.
[0005] In a graph neural network acceleration unit, various samplers are included. Among them, a typical sampler is used to complete the random sampling of neighbor nodes around a specified node of a graph neural network model. The sampled node information is used to construct the embedding expression (or called embedding vector) of the specified node. The main reason for the sampler to only sample partial neighbor node information of the specified node is that if all neighbor node information is sampled, since the number of neighbor nodes of many nodes will increase exponentially with the increase of the graph order (or number of layers), etc., the embedding expressions of these nodes will become too complex and require a large amount of storage overhead.
[0006] Figure 2 is a conceptual schematic diagram of an existing sampler. Refer to Figure 2, the sampler 200 includes a random number generator 202 and an execution component 203. When sampling is required for a specified node, the external system provides the index values of the N neighbor nodes of the node to the execution component 203. At the same time, the random number generator 202 generates K random numbers, and then the execution component 203 compares the K random numbers with the index values of the N neighbor nodes of the node to determine K neighbor nodes to be sampled, where N and K are positive integers greater than 1. However, for such a sampler, the inventor found that due to the disordered state of the K random numbers, the execution component 203 needs to always store the index values of the N neighbor nodes for comparison, and the finally obtained K neighbor nodes cannot ensure uniform distribution around the specified node. Summary of the Invention
[0007] The object of the present disclosure is to provide a sampler and a device for executing a graph neural network model to solve the technical problems existing in the prior art.
[0008] According to the first aspect of the embodiments of the present disclosure, a sampler is provided. The sampler is used to complete random sampling of neighbor nodes around a specified node of a graph neural network model. The sampler includes:
[0009] A random number generator for generating a plurality of random numbers;
[0010] A calculation unit for performing a mathematical operation, where the mathematical operation represents dividing a numerical range between zero and a first integer value into a plurality of equal sub-intervals based on a second integer value, and obtaining a plurality of third integer values within the plurality of sub-intervals based on the plurality of random numbers, where the first integer value represents the number of all neighbor nodes around the specified node, and the second integer value represents the number of neighbor nodes to be sampled for the specified node;
[0011] An execution component for determining neighbor nodes to be sampled from all neighbor nodes around the specified node according to the plurality of third integer values.
[0012] Optionally, the execution component includes:
[0013] An input queue, composed of input storage units, for storing the index values of all neighbor nodes around the specified node;
[0014] An output queue, composed of output storage units;
[0015] A comparison enabling unit for comparing the plurality of third integer values with the index values in the input queue, and when they are the same as a first index value, outputting an enabling signal, and driving the output storage unit to write the first index value output by the input storage unit into the corresponding output storage unit through the enabling signal.
[0016] Optionally, the input queue includes only one input storage unit, and the index values of all the neighbor nodes are stored in the input storage unit in ascending order over a plurality of clock cycles.
[0017] The comparison enabling unit further includes:
[0018] For a newly received third integer value, repeatedly compare the third integer value with the index values in the input storage unit over a plurality of clock cycles, and only when the third integer value is the same as the first index value in the input storage unit, receive another new third integer value in the next clock cycle.
[0019] Optionally, the random number generator generates the plurality of random numbers in the first sub-interval of the plurality of sub-intervals, and the calculation unit maps the plurality of random numbers to the plurality of sub-intervals.
[0020] Optionally, the calculation unit performs the mathematical operation of formula (1):
[0021] sum I = round(random I + N*J / K) Formula (1);
[0022] where round represents a rounding operation, N represents the first integer value, K represents the second integer value, sum I represents a plurality of third integer values, random I represents a random number between [0, N / K], J ∈ {0, 1, 2, 3,..., K - 1}, I ∈ {0, 1, 2, 3,..., K - 1}, both N and K are positive integers greater than 1, and N is greater than K.
[0023] Optionally, the random number generator generates one random number in each sub-interval, thereby obtaining the plurality of random numbers.
[0024] Optionally, the calculation unit performs the mathematical operation of formula (2):
[0025] sum I = round(random I + N / K) Formula (2);
[0026] where round represents a rounding operation, N represents the first integer value, K represents the second integer value, sum I represents a plurality of third integer values, random IRepresents a random number within the corresponding sub - interval, \(I\in\{0,1,2,3,\ldots,K - 1\}\), where both \(N\) and \(K\) are positive integers greater than 1, and \(N\gt K\).
[0027] Optionally, the input queue includes a plurality of storage units, and the comparison enabling unit includes:
[0028] For a newly received third integer value, repeatedly compare the third integer value with each index value in the plurality of input storage units within a plurality of clock cycles. Only after the third integer value is the same as the first index value in the storage unit, will another new third integer value be received in the next clock cycle.
[0029] Optionally, both the input queue and the output queue are first - in - first - out queues.
[0030] According to a second aspect of the embodiments of the present disclosure, there is provided a processing device including a plurality of different operators implemented in hardware to execute various instructions of a graph neural network model, wherein the operator includes the sampler described in any one of the above.
[0031] According to a third aspect of the embodiments of the present disclosure, there is provided a device including:
[0032] At least one of the above - mentioned processing devices;
[0033] A bus channel for receiving various instructions of the graph neural network model;
[0034] A command processor for parsing the various instructions and sending the parsed commands to the at least one processing device for processing.
[0035] According to a fourth aspect of the embodiments of the present disclosure, there is provided a method for randomly sampling neighbor nodes around a specified node of a graph neural network model, the method including:
[0036] Generating a plurality of random numbers;
[0037] Performing a mathematical operation, where the mathematical operation represents dividing the numerical range between zero and a first integer value into a plurality of equal sub - intervals based on a second integer value, and obtaining a plurality of third integer values within the plurality of sub - intervals based on the plurality of random numbers, where the first integer value represents the number of all neighbor nodes around the specified node, and the second integer value represents the number of neighbor nodes to be sampled for the specified node;
[0038] Determining the neighbor nodes to be sampled from all neighbor nodes around the specified node according to the plurality of third integer values.
[0039] Optionally, sampling a partial number of neighbor nodes from all neighbor nodes around the specified node according to the multiple third integer values includes:
[0040] Receiving and storing index values of all neighbor nodes around the specified node; comparing each third integer value of the multiple third integer values with the index values of all neighbor nodes of the specified node, and when it is the same as a first index value, outputting the first index value as the index value of the neighbor node to be sampled, and sampling accordingly.
[0041] Optionally, generating the multiple random numbers includes: generating the multiple random numbers in the first sub-interval of the multiple sub-intervals, and the mathematical operation includes: mapping the multiple random numbers to the multiple sub-intervals.
[0042] Optionally, generating the multiple random numbers includes: generating one random number in each sub-interval, thereby obtaining the multiple random numbers.
[0043] The sampler provided by the embodiments of the present disclosure, for a specified node, first divides the numerical range defined by zero and the number of all neighbor nodes around the specified node into multiple sub-intervals, and then determines the neighbor nodes to be sampled for the specified node according to the multiple random numbers obtained in the multiple sub-intervals, so as to achieve uniform sampling of the neighbor nodes around the specified node. Description of the Drawings
[0044] Through the description of the embodiments of the present disclosure with reference to the following drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0045] Figure 1 is an exemplary social network relationship graph;
[0046] Figure 2 is a schematic diagram of a sampler in the prior art;
[0047] Figure 3 is a hierarchical structure diagram of a data center;
[0048] Figure 4 is a three-dimensional structure diagram of a data center;
[0049] Figure 5 is a schematic diagram of the structure of a cloud server;
[0050] Figure 6 is a schematic diagram of the structure of a GNN acceleration unit;
[0051] Figure 7 is a schematic diagram of the structure of a GNN core according to the embodiments of the present disclosure;
[0052] Figure 8It is a schematic structural diagram of a sampler provided by an embodiment of the present disclosure;
[0053] Figure 9 Used to illustrate K intervals into which N neighbor nodes are divided;
[0054] Figure 10 Based on Figure 9 It is a structural diagram of a variant embodiment generated;
[0055] Figure 11 Used to illustrate the waveform diagrams of the clock signal CLK and the enable signal EN;
[0056] Figure 12 Used to illustrate a bar chart of the performance indicators of the sampler in the prior art and the embodiment of the present disclosure;
[0057] Figure 13 It is a flowchart of the sampling method provided by the embodiment of the present disclosure. Detailed implementation manners
[0058] The following describes the present disclosure based on embodiments, but the present disclosure is not limited to these embodiments. In the following detailed description of the present disclosure, some specific details are described in detail. Those skilled in the art can fully understand the present disclosure without the description of these details. In order to avoid obscuring the essence of the present disclosure, well-known methods, processes, and procedures are not described in detail. Additionally, the drawings are not necessarily drawn to scale.
[0059] The following terms are used herein:
[0060] Acceleration unit: Also known as a neural network acceleration unit, in view of the low efficiency of general-purpose processors in some specialized fields (such as processing images, processing various operations of neural networks, etc.), a processing unit designed to improve the data processing speed in these specialized fields. It often works in conjunction with a general-purpose processor CPU, accepts the control of the general-purpose processor, and executes some specific-purpose or specific-field processing to improve the computer processing efficiency in specific-purpose or specific fields.
[0061] On-chip memory: Memory used separately within the main core or secondary core and cannot be shared.
[0062] Command processor: A command interface between the acceleration unit and the central processing unit that drives the acceleration unit to work. The command processor receives the instructions sent by the central processing unit for the acceleration unit to execute, distributes these instructions to each core in the acceleration unit for execution. Additionally, it is responsible for the synchronization of each core in the acceleration unit.
[0063] Graph Neural Network Model: By organically combining connections and symbols, it not only enables deep learning models to be applied to non-Euclidean structures such as graphs but also endows deep learning models with a certain causal reasoning ability. Graph neural networks extend existing neural networks for processing graph data. In graph data, each node is defined by its characteristics and related nodes, and edges represent the relationships between nodes.
[0064] Data center
[0065] Figure 3 The hierarchical structure diagram of a data center showing a scenario applied in an embodiment of the present disclosure is presented.
[0066] A data center is a specific network of devices for global collaboration, used to transfer, accelerate, display, compute, and store data information on the Internet network infrastructure. In future development, data centers will also become assets for enterprise competition. With the widespread application of data centers, artificial intelligence and the like are increasingly applied to data centers. As an important technology of artificial intelligence, neural networks have been widely applied to big data analysis and operations in data centers.
[0067] In traditional large data centers, the network structure is usually Figure 1 the three-layer structure shown, namely the hierarchical inter-networking model. This model consists of the following three layers:
[0068] Access Layer 103: Sometimes also called the edge layer, it includes access switches 130 and each server 140 connected to the access switches. Each server 140 is a processing and storage entity in the data center, and a large amount of data processing and storage in the data center are completed by these servers 140. The access switch 130 is a switch used to allow these servers to access the data center. One access switch 130 accesses multiple servers 140. The access switches 130 are usually located at the top of the rack, so they are also called Top of Rack switches, and they are physically connected to the servers.
[0069] Aggregation Layer 102: Sometimes also called the distribution layer, it includes aggregation switches 120. Each aggregation switch 120 connects multiple access switches and provides other services at the same time, such as firewalls, intrusion detection, network analysis, etc.
[0070] Core Layer 101: It includes the core switch 110. The core switch 110 provides high-speed forwarding for the packets entering and leaving the data center and provides connectivity for multiple aggregation layers. The network of the entire data center is divided into an L3 layer routing network and an L2 layer routing network. The core switch 110 usually provides a flexible L3 layer routing network for the network of the entire data center.
[0071] Under normal circumstances, the aggregation switch 120 is the demarcation point between the L2 and L3 layer routing networks. Below the aggregation switch 120 is the L2 network, and above is the L3 network. Each group of aggregation switches manages a Point of Delivery (POD). Each POD is an independent VLAN network. When a server migrates within a POD, it does not need to modify its IP address and default gateway because one POD corresponds to one L2 broadcast domain.
[0072] The Spanning Tree Protocol (STP) is usually used between the switch 120 and the access switch 130. STP makes only one aggregation layer switch 120 available for a VLAN network, and other aggregation layer switches 120 are used only when a failure occurs (the dotted lines in the above figure). That is to say, at the aggregation layer, horizontal expansion cannot be achieved because even if multiple aggregation switches 120 are added, only one is working.
[0073] Figure 4 shows Figure 2 the physical connections of the components in the hierarchical data center. As Figure 2 shown, one core switch 110 is connected to multiple aggregation switches 120, one aggregation switch 120 is connected to multiple access switches 130, and one access switch 130 accesses multiple servers 140.
[0074] Cloud server
[0075] The cloud server 140 is the real device in the data center. Since the cloud server 140 runs at high speed to execute various tasks such as matrix calculation, image processing, machine learning, compression, search and sorting, etc., in order to be able to efficiently complete the above various tasks, the cloud server 140 usually includes a central processing unit (CPU) and various acceleration units, such as Figure 5 shown. The acceleration unit is, for example, one of the acceleration units of various neural networks, the data transfer unit (DTU), the graphics processing unit (GPU), the application-specific integrated circuit (ASIC), and the field-programmable gate array (FPGA). The following Figure 3 takes the example shown to introduce each acceleration unit separately.
[0076] Data Transmission Unit (DTU) 260: It is a wireless terminal device specifically used to convert serial port data into IP data or convert IP data into serial port data for transmission through a wireless communication network. The main function of the DTU is to transmit the data of remote devices back to the background center wirelessly. At the front end, the DTU is connected to the customer's device through an interface. After the DTU is powered on and running, it first registers to the mobile GPRS network, and then establishes a socket connection with the background center set in the DTU. The background center serves as the server of the socket connection, and the DTU is the client of the socket connection. Therefore, the DTU and the background software are used together. After the connection is established, the front-end device and the background center can perform wireless data transmission through the DTU.
[0077] Graphics Processing Unit (GPU) 240: It is a processor specifically designed for image and graphics-related operations. By using the GPU, the disadvantage of too little space for computing units in the CPU is overcome. A large number of computing units dedicated to graphics computing are adopted, which reduces the dependence of the graphics card on the CPU and undertakes some computationally intensive image processing tasks originally borne by the CPU.
[0078] Application Specific Integrated Circuit (ASIC): It refers to an integrated circuit designed and manufactured according to the requirements of specific users and the needs of specific electronic systems. Since this kind of integrated circuit is customized according to user requirements, its structure often adapts to the requirements of specific users.
[0079] Field Programmable Gate Array (FPGA): It is a further development product based on programmable devices such as PAL and GAL. It appears as a semi-custom circuit in the field of Application Specific Integrated Circuits (ASIC), which not only solves the deficiencies of custom circuits but also overcomes the drawback of limited gate circuits in the original programmable devices.
[0080] Graph Neural Network Acceleration Unit 230: It is a general term for acceleration units dedicated to graph neural network models. It can be a neural network model for Euclidean Structure Data, or a neural network model for processing non-Euclidean structure data (such as graph data). This article discusses a graph neural network accelerator for processing graph data. The graph neural network model (including executable code and graph data) can be stored in the memory 210, and the scheduling unit 220 deploys the graph neural network model to the graph neural network acceleration unit 230 for execution. Specifically, the scheduling unit 220 can inform the graph neural network acceleration unit 230 in the form of instructions of the storage location of the executable code of the graph neural network model in the memory 210. Then, the graph neural network acceleration unit 230 can address according to these locations and load the executable instructions into the high-speed memory thereon. The scheduling unit 220 can also send the executable code of the graph neural network to the graph neural network acceleration unit 230 in the form of instructions. The graph neural network acceleration unit 230 receives the executable code and loads it into the high-speed memory thereon. Similarly, the graph neural network acceleration unit 230 can also obtain graph data in the above manner. After the graph neural network acceleration unit 230 obtains the executable code and graph data, it executes the executable code and feeds back the execution result.
[0081] The following combines Figure 6 , and specifically illustrates how the scheduling unit controls the acceleration unit to work. It should be understood that regardless of the acceleration unit applicable to any neural network model, the scheduling unit generally drives the acceleration unit to work in the same mode.
[0082] The following combines Figure 6 the internal structure diagrams of the processing unit and the graph neural network acceleration unit 230 to illustrate how the processing unit controls the acceleration unit to work. As Figure 6 shown, the processing unit 220 includes multiple processor cores 222 and a cache 221 shared by the multiple processor cores 222. Each processor core 222 includes an instruction fetch unit 203, an instruction decoding unit 224, an instruction issuing unit 225, and an instruction execution unit 226.
[0083] The instruction fetch unit 223 is used to transfer the instruction to be executed from the memory 210 to the instruction register (which can be one of the registers in the register file 229 Figure 6 shown for storing instructions), and receive the next instruction fetch address or calculate the next instruction fetch address according to the instruction fetch algorithm. The instruction fetch algorithm includes, for example: incrementing or decrementing the address according to the instruction length.
[0084] After fetching the instruction, the processing unit 220 enters the instruction decoding stage. The instruction decoding unit 224 decodes the fetched instruction according to a predetermined instruction format to obtain the operand acquisition information required by the fetched instruction, so as to prepare for the operation of the instruction execution unit 225. The operand acquisition information points to, for example, an immediate value, a register, or other software / hardware that can provide a source operand.
[0085] The instruction issue unit 225 is located between the instruction decoding unit 224 and the instruction execution unit 226 and is used for instruction scheduling and control to efficiently distribute each instruction to different instruction execution units 226, making it possible to perform parallel operations on multiple instructions.
[0086] After the instruction issue unit 225 issues the instruction to the instruction execution unit 226, the instruction execution unit 226 starts to execute the instruction. However, if the instruction execution unit 226 determines that the instruction should be executed by the acceleration unit, it forwards the instruction to the corresponding acceleration unit for execution. For example, if the instruction is a graph neural network inference or graph neural network training instruction, the instruction execution unit 226 no longer executes the instruction, but sends the instruction to the graph neural network acceleration unit 230 through the bus, and the graph neural network acceleration unit 230 executes the instruction.
[0087] The graph neural network acceleration unit 230 internally includes one or more GNN cores, a command processor 237, a direct memory access mechanism 235, and a bus channel 231.
[0088] The bus channel 231 is the channel for instructions to enter and exit the graph neural network acceleration unit 230 from the bus. According to different mechanisms, the bus channel 231 may include a PCIE channel 232, an I2C channel 233, and a JTAG channel 234.
[0089] PCIe, that is, PCI-Express, is a high-speed serial computer expansion bus standard proposed by Intel in 2001, aiming to replace the old PCI, PCI-X, and AGP bus standards. PCIe belongs to high-speed serial point-to-point dual-channel high-bandwidth transmission. The devices connected are allocated exclusive channel bandwidth and do not share bus bandwidth. It mainly supports functions such as active power management, error reporting, end-to-end reliable transmission, hot plugging, and quality of service. Its main advantage is high data transmission rate and considerable development potential. Currently, most of the PCIe buses are PCIe GEN3, but the embodiments of the present disclosure can also adopt PCIe GEN4, that is, a bus channel that follows the PCI-Express 4.0 standard.
[0090] The I2C channel 233 is a simple, two-way, two-wire synchronous serial bus channel developed by Philips. It only requires two wires to transmit information between the devices connected to the bus.
[0091] JTAG is short for Joint Test Action Group and is the common name for IEEE Standard 1149.1, named Standard Test Access Port and Boundary-Scan Architecture. This standard is used to verify the design and test the functionality of printed circuit boards produced. JTAG was officially standardized by IEEE document 1149.1-1990 in 1990. In 1994, a supplementary document was added to describe the Boundary-Scan Description Language (BSDL). Since then, this standard has been widely adopted by electronics companies worldwide. Boundary scan has almost become a synonym for JTAG. The JTAG channel 234 is a bus channel that follows this standard.
[0092] The Direct Memory Access (DMA) mechanism 235 is a function provided by some computer bus architectures that enables data to be directly written from an attached device (such as an external memory) into the on-chip memory of the graph neural network acceleration unit 230. This method greatly improves the efficiency of data access compared to the method where all data transfers between devices have to go through the command processor 237. Because of such a mechanism, the cores of the graph neural network acceleration unit 230 can directly access the memory 210 to read parameters (such as weight parameters of each node) in the deep learning model, etc., greatly improving the data access efficiency. Although the direct memory access mechanism 235 is shown in the figure as being located between the processor 237 and the bus channel 231, the design of the graph neural network acceleration unit 230 is not limited to this. In some hardware designs, each GNN core can include a direct memory access mechanism 235, so that the GNN core can directly read data from the attached device and write it into the on-chip memory of the graph neural network acceleration unit 230 without going through the command processor 237.
[0093] The command processor 237 distributes the instructions sent from the processing unit 220 to the graph neural network acceleration unit 230 to the GNN cores 236 for execution. The instruction execution unit 226 sends the instructions to be executed that require the graph neural network acceleration unit 230 to execute to the graph neural network acceleration unit 230 or the instruction execution unit 226 informs the storage location of the instructions to be executed on the memory 210. After the instruction sequence to be executed enters through the bus channel 231, it is cached in the command processor 237. The command processor 237 selects a GNN core and distributes the instruction sequence to it for execution. The instructions to be executed come from the compiled deep learning model. It should be understood that the instruction sequence to be executed can include the instructions to be executed in the processing unit 220 and the instructions to be executed that require the graph neural network acceleration unit 230.
[0094] GNN core
[0095] Figure 7 It is a schematic structural diagram of a GNN core. The GNN core 236 is the core part of implementing a graph neural network model through hardware, including a scheduler 604, a buffer 608, registers 606, various operators such as a sampler, a pooling operator, a convolutional operator, an activation operator, etc., and a register update unit 607. An operator is a hardware unit that executes a certain or certain instructions of a graph neural network model.
[0096] The scheduler 604 receives instructions from the outside and triggers one or more graph neural network operators according to the instructions. A graph neural network operator is a hardware unit used to execute executable instructions in a graph neural network model. Message queues 620 are used to transfer data between graph neural network operators. For example, the scheduler 604 triggers operator 1 to execute, operator 1 transmits intermediate data into the message queue 620, and operator 2 takes out the intermediate data from the message queue 620 as input data for execution. The message queue 620 represents the general term for message queues between various operators. However, in fact, different message queues are used to transfer intermediate data between different operators. At the same time, the execution result of the operator is written into the result queue (also included in the message queue 620), and the register update unit 607 takes out the operator execution result from the result queue and updates the corresponding status register, result register, and / or status register accordingly. The scheduler 604 can also send various requests to the outside, and these requests are sent out through the command processor 237 and via the bus channel 231. For example, the scheduler 604 can send a data loading request to the scheduling unit 220. The scheduling unit 220 obtains the data access address and transmits the access address to the scheduler 604. The scheduler 604 provides the data access address to the acceleration unit GNN core, and the acceleration unit GNN core can provide control to the direct memory access mechanism 609, and the direct memory access mechanism 609 controls the data loading.
[0097] It should be noted that the graph neural network model referred to in this article is a general term for all models of neural networks applied to graph data. However, according to different adopted technologies and classification methods, graph neural network models can be divided into different categories. For example, from the perspective of propagation methods, graph neural network models can be divided into graph convolutional neural network (GCN) models, graph attention network (GAT, abbreviated to distinguish from GAN) models, Graph LSTM models, etc. Therefore, a graph neural network acceleration unit is usually dedicated to accelerating a certain type of graph neural network model, and different graph neural network acceleration units can design different hardware operators. However, generally, a graph neural network acceleration unit will include a typical sampler, which is used to complete the random sampling of neighbor nodes around a specified node of a graph neural network model, and the partial node information obtained by sampling is used to construct the embedding expression of the specified node.
[0098] Sampler
[0099] Figure 8 It is a structural diagram of a sampler provided by an embodiment of the present disclosure. As shown in the figure, the sampler 800 includes a computing unit 801, a random number generator 802, and an execution component 803.
[0100] The random number generator 802 is used to generate random numbers.
[0101] The computing unit 801 is used to perform specific mathematical operations. Specifically, the specific mathematical operation means dividing the numerical range between zero and the first integer value into multiple equal sub-intervals based on the second integer value, and obtaining multiple third integer values within the multiple sub-intervals based on multiple random numbers received from the random number generator 802, where the first integer value represents the number of multiple neighbor nodes of a specified node, and the second integer value represents the number of neighbor nodes to be sampled for the specified node.
[0102] The execution component 803 is used to determine the neighbor nodes to be sampled from all the neighbor nodes around the specified node according to the multiple third integer values.
[0103] For the sampler provided by the embodiment of the present disclosure, for a specified node, first divide the numerical range defined by zero and the number of all neighbor nodes around the specified node into multiple sub-intervals, and then determine the neighbor nodes to be sampled for the specified node according to multiple random numbers obtained within the multiple sub-intervals, so as to achieve uniform sampling of the neighbor nodes around the specified node.
[0104] In addition, since the sampler samples the neighbor nodes around the specified node in the order of the magnitudes of multiple random numbers, when the sampling operation corresponding to a random number is completed, the index value of the neighbor node within the sub-interval corresponding to the random number can no longer be stored in the execution component, which can save the storage overhead in the execution component.
[0105] In one embodiment, the execution component 803 includes a comparison enable unit 804, an input queue 805, and an output queue 806. The input queue 805 is composed of input storage units and is used to store the index values of multiple neighbor nodes of the specified node. The output queue 806 is composed of output storage units. Here, the storage units are divided into input storage units and output storage units only for convenience of description, and there is no essential difference between the two. The comparison enable unit 804 is used to perform the following steps: continuously receive the third integer value from the computing unit 801, and compare the third integer value with the index value in the input storage unit. When it is the same as the first index value, output an enable signal, and drive the output storage unit to write the first index value output by the input storage unit into the corresponding output storage unit in the output queue 806 through the enable signal.
[0106] In summary, in this embodiment, the numerical range between zero and the first integer value (the number of multiple neighbor nodes of a specified node) is divided into multiple equal sub-intervals, and then a random integer is obtained within each sub-interval. Subsequently, the index value identical to each random integer is found among the multiple index values of the multiple neighbor nodes and written into the output queue. Finally, each output storage unit of the output queue stores the index value of the neighbor node to be sampled. Since in this embodiment, a random integer is obtained within each sub-interval, and the index value identical to each random integer is used as the index value of the sampled neighbor node. In other words, the index value of each sampled neighbor node is located within each sub-interval, so the sampled neighbor nodes are relatively evenly distributed, which is beneficial to balancing the sampling bias.
[0107] Continue to refer to Figure 8 As shown, in one embodiment, the calculation unit 801 includes a divider 8011 and an adder 8012. The divider 8011 is used to perform a division operation. We denote the first integer value as N and the second integer value as K, then the division operation is: step_len = N / K. The random number generator 802 is essentially a pseudo-random number generator (i.e., generating random numbers within a specified range). The random number generator 802 generates K random numbers random with value ranges between [0, step_len], [step_len, 2 * step_len], ……, [(K - 1) * step_len, K * step_len]. The adder 8012 is used to add the operation result step_len output by the divider 8011 and the random number random output by the random number generator 802 to obtain the sum sum, and output sum to the comparison enable unit 804. The number of random numbers random is K. Therefore, the calculation unit 801 will output K sums.
[0108] The mathematical operations of the calculation unit 802 can be expressed by the following formula.
[0109] step_len = N / K Formula (1)
[0110] sum I = round(random I + step_len) Formula (2)
[0111] Wherein, N and K are positive integers greater than 1 and N is greater than K, random I represents the random number output each time, I ∈ {0, 1, 2, 3, …, K - 1}, and round is a rounding operation.
[0112] In another embodiment, the random number generator 802 generates K random numbers random within the range [0, step_len], and the calculation unit 801 further needs to include a multiplier (not shown). The digital operations composed of the multiplier, adder, and divider are represented by formulas (3)-(5):
[0113] step_len = N / K Formula (3)
[0114] product = step_len * J Formula (4)
[0115] sum = round(random + product) Formula (5)
[0116] Where N and K are positive integers greater than 1 and N is greater than K, J is an integer, and the value range of J satisfies the formula J ∈ {0, 1, 2, 3,..., K - 1}, and round is a rounding operation.
[0117] Next, based on Figure 9 the working process of the sampler will be further described. Figure 9 It is used to indicate that N neighbor nodes are divided into K sub-intervals.
[0118] According to this embodiment, the external system stores the index values of N neighbor nodes of a certain node into the input queue 805. This storage process is serial, that is, one neighbor node's index value is stored per clock cycle, and a total of N clock cycles are consumed to complete the storage operation of the index values of N neighbor nodes. During these N clock cycles, the random number generator 802 performs K random number generation operations and obtains K random numbers random. During these N cycles, the calculation unit 801 outputs K sums. Refer to Figure 9As shown, the K sums are K random integers within [0, step_len], [step_len, 2*step_len], [2*step_len, 3*step_len], …, [(K-1)*step_len, K*step_len]. During these N clock cycles, for each newly calculated sum received from the calculation unit 801 by the comparison enable unit 804, the comparison enable unit 804 compares the sum with each index value in the input queue 805. When the first index value equal to the sum is obtained, an enable signal EN is sent to a certain output storage unit, and the enable signal EN will drive the output storage unit to write the first index value output by the input storage unit into the output storage unit. Only after finding the first index value equal to it will the next newly calculated sum be received. Further, if the N index values are input into the input queue 805 in ascending order, then when comparing the K gradually increasing sums with the N gradually increasing index values, only the latest sum needs to be compared with the latest index value. Thus, for N neighbor nodes, it takes no more than N clock cycles to obtain the indices of the K sampled neighbor nodes.
[0119] In an alternative embodiment, both the input queue 805 and the output queue 807 are first-in-first-out queues. A first-in-first-out queue means that the index value that enters the queue first will also be output from the queue first. For example, the index value that enters the input queue 805 first is greater than the index value that enters the input queue 805 later.
[0120] Figure 10 is based on Figure 9 The structural diagram of the embodiment variation generated. In Figure 10 the only difference is that the input queue only includes one input storage unit 807, that is, it can only store one index value.
[0121] Figure 11 A waveform diagram for schematically showing the clock signal CLK and the enable signal EN. Based on Figure 11 to exemplarily illustrate Figure 10Signal control flow of the sampler. In the first clock cycle of CLK, the external system stores index1 in the input storage unit 807. At the same time, the comparison enable unit 804 compares the index value index1 in the input storage unit 807 with the received sum1, and they are different. Then, in the second clock cycle, the external system stores index2 in the input storage unit 807. At the same time, the comparison enable unit 804 compares sum1 with index2 in the input storage unit 807, and they are the same. Then, a high-level enable signal EN is output as the write enable signal for the output queue 806 to drive the output queue 806 to write the index value index2 into the corresponding output storage unit of the output queue 806. Then, in the third clock cycle, the external system stores index3 in the input storage unit 807. At the same time, the comparison enable unit 804 receives sum2 and compares sum2 with index3, and they are different. Then, in the fourth clock cycle, the external system stores index4 in the input storage unit 807. At the same time, the comparison enable unit 804 compares sum2 with index4, and they are the same. A high-level enable signal EN is generated as the write enable signal for the output queue 806 to drive the output queue 806 to write the index value index4 into the corresponding output storage unit of the output queue 806, and so on. However, it should be emphasized here that when the external system stores the index value in the input storage unit 807, it needs to be stored in ascending order, that is, index1 < index2 < index3 < index4. In summary, the sampler provided by the embodiment of the present disclosure completes the sampling operation of K neighbor nodes within N clock cycles, while the prior art needs to complete the sampling operation of N neighbor nodes within (N + K). Therefore, the embodiment of the present disclosure improves the sampling efficiency, reduces the latency, and the sampled neighbor nodes are more uniform.
[0122] The sampler of the embodiment of the present disclosure only uses one storage unit to store the index values of neighbor nodes, while the sampler provided by the prior art needs N storage units to store the index values of neighbor nodes. Therefore, the embodiment of the present disclosure reduces the memory occupancy and helps to reduce the manufacturing cost.
[0123] Based on the laboratory environment, we verified the sampler of the prior art and the sampler of the embodiment of the present disclosure, and the bar chart as shown in Figure 12 can be obtained. Among them, BaseLine represents the sampler of the prior art, Invented represents the sampler of the embodiment of the present disclosure. The left bar chart represents the algorithm performance, and the right represents the number of registers. After the experiment, it can be determined that the algorithm performance has increased by 91.9% and 23% of the registers have been saved. The number of neighbor nodes of the sampler in this experiment is NMAX = 270, and the maximum value of the neighbor nodes to be sampled is KMAX = 25.
[0124] Corresponding to the above hardware-implemented sampler, an embodiment of the present disclosure also provides a software-implemented sampling method, which is used to complete random sampling of neighbor nodes around a specified node of a graph neural network model. As Figure 13 shown, the method includes the following steps.
[0125] Step S101, generate a plurality of random numbers.
[0126] Step S102, perform a mathematical operation to obtain a plurality of third integer values. The mathematical operation represents dividing the numerical range between zero and a first integer value into a plurality of equal sub-intervals based on a second integer value, and obtaining a plurality of third integer values within the plurality of sub-intervals based on the plurality of random numbers, where the first integer value represents the number of all neighbor nodes around the specified node, and the second integer value represents the number of neighbor nodes to be sampled for the specified node.
[0127] Step S103, determine the neighbor nodes to be sampled from all neighbor nodes around the specified node according to the plurality of third integer values.
[0128] In one embodiment, step S103 includes: receiving and storing the index values of all neighbor nodes around the specified node; and comparing each of the plurality of third integer values with the index values of all neighbor nodes of the specified node. When it is the same as a first index value, output the first index value as the index value of the neighbor node to be sampled, and sample accordingly.
[0129] In one embodiment, step S101 includes: generating a plurality of random numbers in the first sub-interval of the plurality of sub-intervals, and the mathematical operation is used to map the plurality of random numbers to the plurality of sub-intervals respectively. This mapping can be completed using the following formula (6).
[0130] sum I = round(random I + N * J / K) Formula (6);
[0131] where round represents a rounding operation, N represents the number of neighbor nodes, K represents the number of intervals to be divided, J ∈ {0, 1, 2, 3,..., K - 1}, I ∈ {0, 1, 2, 3,..., K - 1}, and both N and K are positive integers greater than 1, and N is greater than K.
[0132] Commercial value of the embodiments of the present disclosure
[0133] The sampler provided by the embodiments of the present disclosure is applied to a graph neural network acceleration unit and can uniformly sample neighbor nodes. In a further embodiment, the number of memories can also be reduced and the sampling efficiency can be improved, thereby reducing the manufacturing cost of the graph neural network acceleration unit. Therefore, the sampler provided by the embodiments of the present disclosure and the graph neural network acceleration unit including such a sampler should have application prospects and commercial value.
[0134] Those skilled in the art can understand that the present disclosure can be implemented as a system, a method, and a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely, complete hardware, complete software (including firmware, resident software, microcode), and can also be implemented in the form of a combination of software and hardware. In addition, in some embodiments, the present disclosure can also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.
[0135] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium is, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium include: an electrical connection of one or more specific wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical memory, a magnetic memory, or any suitable combination of the above. In this article, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by, or in combination with, a processing unit, apparatus, or device.
[0136] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any other suitable combination. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by, or in combination with, an instruction system, apparatus, or device.
[0137] The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., and any suitable combination of the above.
[0138] The computer program code for implementing the embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as JAVA and C++, and may also include conventional procedural programming languages such as C. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).
[0139] The foregoing are only the preferred embodiments of the present disclosure and are not intended to limit the present disclosure. For those skilled in the art, various modifications and changes can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A sampler, which is used to randomly sample the neighbor nodes around a specified node of a graph neural network model, and the sampler is implemented in hardware. It includes: A random number generator for generating a plurality of random numbers; A computing unit including a divider and an adder for performing mathematical operations, where the mathematical operations represent dividing the numerical range between zero and a first integer value into a plurality of equal sub-intervals based on a second integer value, and obtaining a plurality of third integer values within the plurality of sub-intervals based on the plurality of random numbers, wherein the first integer value represents the number of all neighbor nodes around the specified node, and the second integer value represents the number of neighbor nodes to be sampled for the specified node; An execution component for determining the neighbor nodes to be sampled from all the neighbor nodes around the specified node according to the plurality of third integer values.
2. The sampler according to claim 1, wherein, the execution component includes: An input queue composed of input storage units for storing the index values of all neighbor nodes around the specified node; An output queue composed of output storage units; A comparison enabling unit for comparing the plurality of third integer values with the index values in the input queue, and when they are the same as a first index value, outputting an enabling signal, and driving the output storage unit to write the first index value output by the input storage unit into the corresponding output storage unit through the enabling signal.
3. The sampler according to claim 2, wherein, the input queue only includes one input storage unit, and the index values of all the neighbor nodes are stored in the input storage unit in ascending order in a plurality of clock cycles, then the comparison enabling unit further includes: For a newly received third integer value, repeatedly comparing the third integer value with the index value in the input storage unit in a plurality of clock cycles, and only when the third integer value is the same as the first index value in the input storage unit, receiving another new third integer value in the next clock cycle.
4. The sampler according to claim 1, wherein, the random number generator generates the plurality of random numbers in the first sub-interval of the plurality of sub-intervals, and the computing unit maps the plurality of random numbers to the plurality of sub-intervals.
5. The sampler according to claim 4, wherein, the computing unit performs the mathematical operations of formula (1) to map the plurality of random numbers to the plurality of sub-intervals: Formula (1); Among them, represents a rounding operation, N represents the first integer value, and K represents the second integer value. represents multiple third integer values. represents a random number between [0, N / K], J ∈ {0, 1, 2, 3, …, K - 1}, I ∈ {0, 1, 2, 3, …, K - 1}, both N and K are positive integers greater than 1, and N is greater than K.
6. The sampler according to claim 1, wherein, the random number generator generates one random number in each sub-interval, thereby obtaining the plurality of random numbers.
7. The sampler according to claim 6, wherein, the computing unit executes the mathematical operations of formula (2): Formula (2); where round represents a rounding operation, N represents the first integer value, K represents the second integer value, represents multiple third integer values, represents a random number within the corresponding sub-interval, I ∈ {0, 1, 2, 3, …, K−1}, and both N and K are positive integers greater than 1, and N is greater than K.
8. The sampler according to claim 2, wherein, the input queue includes a plurality of storage units, then the comparison enabling unit includes: For a newly received third integer value, repeatedly compare the third integer value with each index value in the multiple input storage units over multiple clock cycles. Only after the third integer value is the same as the first index value in the storage unit will another new third integer value be received in the next clock cycle.
9. The sampler according to claim 2, wherein, both the input queue and the output queue are first-in, first-out queues.
10. A processing device includes a plurality of different operators implemented in hardware to execute various instructions of a graph neural network model, wherein, the operator includes the sampler according to any one of claims 1 to 9.
11. A device, comprising: at least one processing device according to claim 10; a bus channel for receiving various instructions of the graph neural network model; a command processor for parsing the various instructions and sending the parsed commands to at least one of the processing devices for processing.
12. A method for performing random sampling of neighbor nodes around a specified node of a graph neural network model, applied to a sampler in a graph neural network accelerator for executing a graph neural network model for processing graph data, the graph data including graph data in any of the following scenarios: image processing, the method comprises: generating a plurality of random numbers; performing a mathematical operation, the mathematical operation representing dividing a numerical range between zero and a first integer value into a plurality of equal sub-intervals based on a second integer value, and obtaining a plurality of third integer values within the plurality of sub-intervals based on the plurality of random numbers, wherein the first integer value represents the number of all neighbor nodes around the specified node, and the second integer value represents the number of neighbor nodes to be sampled for the specified node; determining the neighbor nodes to be sampled from all the neighbor nodes around the specified node according to the plurality of third integer values.
13. The method according to claim 12, wherein, the determining the neighbor nodes to be sampled from all the neighbor nodes around the specified node according to the plurality of third integer values includes: receiving and storing the index values of all the neighbor nodes around the specified node; comparing each of the plurality of third integer values with the index values of all the neighbor nodes of the specified node, and when it is the same as the first index value, outputting the first index value and recording it as the index value of the neighbor node to be sampled.
14. The method according to claim 12, wherein, the generating a plurality of random numbers includes: generating the plurality of random numbers in the first sub-interval of the plurality of sub-intervals, and the mathematical operation includes: mapping the plurality of random numbers to the plurality of sub-intervals.
15. The method according to claim 12, wherein, the generating a plurality of random numbers includes: generating one random number in each sub-interval to obtain the plurality of random numbers.
Citation Information
Patent Citations
Training sample screening method and device, electronic equipment and storage medium
CN111881936A
Vehicle-mounted laser point cloud marking classification method based on graph structure and attention mechanism
CN112070054A