Processing unit and method for configuring the same
Through a configurable and reconfigurable processing architecture, combined with FPGA and NVM, the shortcomings of existing processing units between computing intensity and memory capacity are solved, and efficient computing resource configuration and memory management are achieved, which is suitable for applications such as graphics computing and neural network computing.
Patent Information
- Application Number
- CN202080102792.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-18
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2040-09-18
AI Technical Summary
Existing central processing units (CPUs) provide high memory capacity and high communication bandwidth but low computing intensity, while graphics processing units (GPUs) provide high computing intensity but low memory capacity and communication bandwidth. There is a lack of a flexible processing architecture that can simultaneously provide high computing intensity, high memory capacity and high communication bandwidth.
A configurable and reconfigurable processing architecture is adopted, through network coupling of core processing elements and auxiliary processing elements, combined with field programmable gate arrays (FPGAs) and non-volatile memories (NVMs), and computing and memory resources are dynamically configured through configuration bitstreams to achieve flexible coupling of core and auxiliary processing elements.
It provides a flexible processing unit architecture that can be configured and reconfigured according to different application requirements, improves computing efficiency, is suitable for graphics computing and neural network computing, and reduces costs.
Smart Images

Figure CN115803724B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of computer technology, and more particularly to a processing unit and a method for configuring a processing unit. Background Art
[0002] Computing systems have made significant contributions to the advancement of modern society and are used in numerous applications to achieve beneficial results. Numerous devices, such as desktop personal computers (PCs), laptops, tablets, netbooks, smartphones, and servers, have facilitated increased productivity and reduced costs in communicating and analyzing data across most areas of entertainment, education, business, and science. Many technologies and applications require processing units optimized for large datasets, with high computational intensity and high memory bandwidth. Other technologies and applications require processing units optimized for smaller datasets, utilizing lightweight computing and low latency. Conventional central processing units (CPUs) typically offer high memory capacity and high communication bandwidth, but relatively low computational intensity. In contrast, conventional graphics processing units (GPUs) typically offer higher computational intensity but relatively low memory capacity and low communication bandwidth. Therefore, there is a continuing need for a flexible processing architecture that can provide high computational intensity, high memory capacity, and high communication bandwidth. Summary of the Invention
[0003] The present disclosure is best understood by referring to the following description and accompanying drawings, which illustrate embodiments of the present disclosure directed to a configurable and reconfigurable processing architecture that can be advantageously used for graphics computing, neural network computing, graph neural network computing, and the like.
[0004] In one embodiment, the processing unit may include a core processing element and multiple auxiliary processing elements coupled together via one or more networks. The core processing element may include processing logic, such as, but not limited to, a field programmable gate array (FPGA) and a non-volatile memory (NVM), such as, but not limited to, a flash memory (FLASH) memory. The processing logic and non-volatile memory (NVM) of the core processing element may have configurable computing and memory resources. The multiple auxiliary processing elements may each include processing logic, such as, but not limited to, a field programmable gate array (FPGA) and a non-volatile memory (NVM), such as, but not limited to, a flash memory (FLASH) memory. The processing logic and non-volatile memory (NVM) of the auxiliary processing elements may have configurable memory resources and connection topology. The processing unit may also include one or more input / output interfaces for receiving one or more bitstreams for configuring and reconfiguring the processing unit.
[0005] In one embodiment, configuring the processing unit may include accessing one or more configuration bitstreams. Accessing one or more configuration bitstreams may include retrieving one or more bitstreams from a non-volatile memory (NVM) of multiple auxiliary processing elements and a core processing element, or receiving one or more configuration bitstreams from a host device at one or more input / output interfaces. The core processing element, multiple auxiliary processing elements, and one or more networks may be configured based on one or more configuration bitstreams. The configuration of the processing unit may include, but is not limited to, the configuration of computing resources of the core processing element and memory management of multiple auxiliary processing elements. One or more applications may then be executed on the configured processing unit, including but not limited to one or more long short-term memory (LSTM) NN applications, one or more multi-layer perception (MLP) NN applications, one or more gated recurrent unit (GRU) NN applications, one or more attention model (attention) NN applications, one or more sum / mean NN applications, and the like.
[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Embodiments of the present disclosure are illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate similar elements and in which:
[0008] Figure 1 A block diagram of a processing unit according to some aspects of the present disclosure is shown.
[0009] Figure 2 A flowchart illustrating a method of configuring a processing unit according to some aspects of the present disclosure is shown.
[0010] Figure 3A and 3B Exemplary configurations of processing units according to some aspects of the present disclosure are described.
[0011] Figure 4A 、 4B 4C illustrate exemplary configurations of processing units according to aspects of the present disclosure. DETAILED DESCRIPTION
[0012] Reference will now be made in detail to embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. Although the present disclosure will be described in conjunction with these embodiments, it will be understood that they are not intended to limit the present disclosure to these embodiments. On the contrary, the present invention is intended to encompass alternatives, modifications, and equivalents that may be included within the scope of the invention as defined by the appended claims. In addition, in the following detailed description of the present disclosure, many specific details are set forth in order to provide a comprehensive understanding of the present disclosure. However, it will be understood that the present disclosure can be practiced without these specific details. In other cases, well-known methods, processes, components, and circuits are not described in detail to avoid unnecessarily obscuring some aspects of the present disclosure.
[0013] Some embodiments of the present disclosure below are presented in the form of routines, modules, logic blocks, and other symbolic representations of operations on data within one or more electronic devices. Description and representation are the means used by those skilled in the art to most effectively convey the content of their work to others skilled in the art. Routines, modules, logic blocks, etc. are generally considered herein to be self-consistent sequences of processes or instructions that lead to desired results. These processes involve physical operations on physical quantities. Typically, although not necessarily, these physical operations take the form of electrical or magnetic signals that can be stored, transmitted, compared, and otherwise manipulated in electronic devices. For convenience, and in conjunction with common usage, with respect to the embodiments of the present disclosure, these signals are referred to as data, bits, values, elements, symbols, characters, terms, numbers, character strings, etc.
[0014] However, it should be remembered that these terms are to be interpreted as referring to physical operations and quantities and are merely convenient labels and will be further interpreted according to terminology commonly used in the art. Unless otherwise clearly indicated from the discussion below, it should be understood that throughout the discussion of this disclosure, discussions utilizing terms such as "receiving" and the like refer to actions and processes of electronic devices (such as electronic computing devices that manipulate and transform data). Data is represented as physical (e.g., electronic) quantities in the logic circuits, registers, memory, etc. of the electronic device and is converted to other data similarly represented as physical quantities in the electronic device.
[0015] In this application, the use of transitional conjunctions is intended to include conjunctions. The use of definite or indefinite articles is not intended to indicate cardinality. In particular, reference to "the" or "an" object is intended to indicate one of a possible plurality of such objects. The use of terms such as "include" and "comprising" specifies the presence of the element being described, but does not exclude the presence or addition of one or more other elements or groups thereof. It should also be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used herein to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the embodiments. It should also be understood that when an element is referred to as being "coupled" to another element, it may be directly or indirectly connected to the other element, or there may be intervening elements. In contrast, when an element is referred to as being "directly connected" to another element, there are no intervening elements. It should also be understood that the term "and or" includes any and all combinations of one or more related elements. It should also be understood that the wording and terminology used herein are for descriptive purposes and should not be construed as limiting.
[0016] Figure 1 A block diagram of a processing unit according to some aspects of the present disclosure is shown. The processing unit 100 may include a core processing element 110 and multiple auxiliary processing elements 120a-120m coupled together via one or more networks 130a-130n. In one embodiment, the core processing element 110 and the multiple auxiliary processing elements 120a-120m may be coupled together via one or more configurable mesh networks 130a-130n. In the mesh networks 130a-130n, the core processing element 110 and the multiple auxiliary processing elements 120a-120m may be configurably coupled together, dynamically and non-hierarchically coupled to as many other elements as possible, and cooperate with each other to efficiently route data, instructions, control signals, etc. The processing unit 110 may also include one or more input / output interfaces 140. The one or more input / output interfaces 140 may be configured to couple the processing unit 100 to a host device 150, etc. In one implementation, the core processing element 110, the plurality of auxiliary processing elements 120a-120m, and the one or more networks 130a-130n may be implemented in a system-in-package (SiP). In another implementation, the core processing element may be implemented as a peripheral card, such as a Peripheral Component Interface Express (PCIe) card. The plurality of auxiliary processing elements 120a-120m may be implemented as one or more add-in peripheral cards, such as a Peripheral Component Interface Express (PCIe) card, each of which includes one or more auxiliary processing units 120a-120m.
[0017] Core processing element 110 may include configurable computing logic and non-volatile memory (NVM). In one embodiment, the configurable computing logic and NVM of core processing element 110 may be tightly coupled. In one embodiment, core processing element 110 may include a field programmable gate array (FPGA) and flash memory (FLASH). Core processing element 110 may include configurable graphics computing capabilities. Core processing element 110 may also include configurable data flow capabilities.
[0018] The plurality of auxiliary processing elements 120a-120m may include computing logic and non-volatile memory (NVM). The computing logic and non-volatile memory (NVM) of the auxiliary processing elements 120a-120m may be smaller than the computing logic and non-volatile memory (NVM) of the core processing element 110. In one embodiment, the configurable computing logic and NVM of each auxiliary processing element 120a-120m may be tightly coupled together. In one implementation, each auxiliary processing element 120a-120m may include a field programmable gate array (FPGA) and flash memory (FLASH). The auxiliary processing elements 120a-120m may include multiple memory channels. In one embodiment, the plurality of auxiliary processing elements 120a-120m may be homogeneous, having the same computing logic, the same non-volatile memory, the same memory capacity, and the same memory channels between the computing logic and the non-volatile memory. In another embodiment, the plurality of auxiliary processing elements 120a-120m may be heterogeneous, where one or more auxiliary processing elements have different computational logic, different non-volatile memory, different memory capacity, different number of memory channels, etc., than one or more other auxiliary processing elements.
[0019] The core processing element 110 can be configured based on one or more configuration bitstreams. In one embodiment, the computing resources of the computing logic of the core processing element 110 can be configured based on one or more configuration bitstreams. For example, the allocation of computing resources of the core processing element 110 can be configured for a single graph neural network (GNN) application or multiple GNN applications. Computing resources can also be allocated for inference, training and or optimization of sampling patterns of one or more graph neural network (GNN) applications. In another example, one or more functions of the computing resources of the core processing element 110 can be configured. One or more functions can be configured to implement a long short-term memory (LSTM) model, a multi-layer perception (MLP), a gated recurrent unit (GRU), an attention model, and / or mean and or other similar general neural network operations. In yet another example, the neural network node sampling strategy of the core processing element 110 can be configured. The node sampling strategy can include, but is not limited to, random and iterative nodes, random or iterative edges, and / or edge weights, random, maximum K, in-degree and window of neighboring nodes. In yet another example, the operation (OP) instruction set of the core processing element 110 can be configured. The operation (OP) instruction set may include, but is not limited to, large, medium, and small OP instruction sets. In another embodiment, the memory resources of the core processing element 110 may be configured based on one or more configuration bitstreams. For example, the memory capacity, number of memory channels, topology, etc. of the core processing element 110 may be configured.
[0020] The plurality of auxiliary processing elements 120a-120m may also be configured based on one or more configuration bitstreams. In one embodiment, the memory management of the plurality of auxiliary processing elements 120a-120m may be configured based on one or more configuration bitstreams. For example, the memory capacity of the processing elements 120a-120m may be configured. In another embodiment, the network interfaces of the plurality of auxiliary processing elements 120a-120m may be configured based on one or more configuration bitstreams. For example, the auxiliary processing elements 120a-120m may be configured to have one or more serial connections and / or one or more parallel connections.
[0021] One or more networks 130a-130n may also be configured based on one or more configuration bitstreams. In one implementation, the one or more configuration bitstreams may directly configure the communication links between the plurality of auxiliary processing elements 120a-120m and the core processing element 110. In another embodiment, the one or more networks 130a-130n may be indirectly configured by configuring one or more serial interfaces and / or one or more parallel interfaces of the auxiliary processing elements 120a-120m and the core processing element 110.
[0022] In one embodiment, each of the one or more configuration bitstreams may be stored in a non-volatile memory (NVM) of each of the plurality of auxiliary processing elements 120a-120m and the corresponding core processing element 110. Processing unit 100 may be configured or reconfigured by loading different configuration bitstreams into the NVM of each of the plurality of auxiliary processing elements 120a-120m and the corresponding core processing element 110 via one or more input / output interfaces 140. In another embodiment, one or more configuration bitstreams may be received from a host device 150. Host device 150 may include a software daemon 160 and non-volatile memory (NVM) 170. NVM 170 of host device 150 may store any number of configuration bitstreams. Software daemon 160 of host device 150 may receive user application programming interface (API) calls 180 to configure or reconfigure processing unit 100. In response to the application programming interface call 180, the software daemon 160 can configure or reconfigure the plurality of auxiliary processing elements 120a-120m and the core processing element 110 through the one or more input / output interfaces 140 of the processing unit 100 based on one or more configuration bitstreams stored in the non-volatile memory 170 of the host 150.
[0023] Now refer to Figure 2 , a method for configuring a processing unit according to some aspects of the present disclosure is shown. The method for configuring a processing unit may begin by accessing one or more configuration bitstreams at 210. In one embodiment, accessing the one or more configuration bitstreams may include retrieving configuration bitstreams stored in non-volatile memory (NVM) of each of a plurality of auxiliary processing elements and a corresponding core processing element of the processing unit. In another embodiment, accessing the one or more configuration bitstreams may include receiving the one or more configuration bitstreams on one or more input / output interfaces of the processing unit. For example, in response to an application programming interface (API) call received by a host device, the host device may send one or more configuration bitstreams on one or more input / output interfaces of the processing unit. The one or more configuration bitstreams may include configuration parameters associated with executing one or more given applications on the processing unit. For example, the one or more configuration bitstreams may include configuration parameters associated with executing one or more neural network (NN) applications on the processing unit.
[0024] At 220, the core processing element of the processing unit can be configured based on one or more configuration bitstreams. In one embodiment, the computing logic of the core processing element can be configured based on one or more configuration bitstreams. For example, the allocation of computing resources of the computing logic can be configured for a single graph neural network (GNN) application or multiple GNN applications. Computing resources can also be allocated for optimization of reasoning, training, sampling, and or similar modes. In another example, one or more functions of the computing logic of the core processing element can be configured. One or more functions can be configured to implement a long short-term memory (LSTM) model, a multi-layer perception (MLP), a gated recurrent unit (GRU), an attention model, and / or mean and or other similar graph neural network operations. In another example, the neural network node sampling strategy of the core processing element can be configured. The node sampling strategy can include, but is not limited to, random and iterative nodes, random or iterative edges, and / or edge weights, random, maximum K, in-degree, and window of adjacent nodes. In another example, the operation instruction set of the core processing element can be configured. The operation (OP) instruction set can include, but is not limited to, large, medium, and small OP instruction sets. In another embodiment, the memory resources of the core processing element may be configured based on one or more configuration bitstreams. For example, the memory capacity, number of memory channels, topology, etc. of the core processing element may be configured.
[0025] At 230, multiple auxiliary processing elements of the processing unit can be configured based on one or more configuration bitstreams. In one embodiment, memory management of the multiple auxiliary processing elements can be configured based on the one or more configuration bitstreams. For example, the memory capacity of the processing element can be configured. In another embodiment, network interfaces of the multiple auxiliary processing elements can be configured based on the one or more configuration bitstreams. For example, the auxiliary processing elements can be configured with one or more serial communication connections and / or one or more parallel communication connections.
[0026] At 240, networks coupling the plurality of auxiliary processing elements to the core processing element may be configured based on one or more configuration bitstreams. In one embodiment, the one or more networks may be configured directly based on the one or more configuration bitstreams. In another embodiment, the one or more networks may be configured indirectly through configuration of one or more serial interfaces and / or one or more parallel interfaces of the auxiliary processing elements and the core processing element.
[0027] At 250, one or more given applications may be executed on the configured processing unit. In one embodiment, one or more neural network (NN) applications may be executed on the configured processing unit, such as, but not limited to, one or more long short-term memory (LSTM) model NN applications, one or more multi-layer perceptron (MLP) NN applications, one or more gated recurrent unit (GRU) NN applications, one or more attention model NN applications, one or more sum / mean neural network applications, etc.
[0028] The process of 210-250 can be repeated to reconfigure the processing unit at 260 one or more times. For example, the processing unit can be configured at 210-250 to execute a first neural network application. Then, the processing unit can be reconfigured at 260 to execute a second neural network application. In another example, the processing unit can be configured to execute a training mode of one or more graph neural network (GNN) applications at night. Then, the processing unit can be reconfigured to execute an inference mode of one or more graph neural network (GNN) applications during the day. In training mode, the processing unit can be configured to process large data sets, providing high computational intensity and high data bandwidth. In inference mode, the processing unit can be configured to process smaller data sets, utilizing lighter weight computations and high data transfer rates.
[0029] Now refer to Figure 3A and 3B , illustrates an exemplary configuration of a processing unit according to some aspects of the present disclosure. In one embodiment, the processing unit 100 may be configured to perform the following Figure 3A The LSTM model neural network application shown in FIG. LSTM configuration can be optimized for training, sampling, or inference. Then, the processing unit 100 can be reconfigured to perform the following Figure 3B The attention model neural network application shown. Similarly, the attention model configuration can be optimized for training, sampling, or inference.
[0030] Now refer to Figure 4A 、 4B 4C, illustrate exemplary configurations of processing units according to some aspects of the present disclosure. In one embodiment, processing unit 100 may be configured to execute multiple applications. For example, a first portion of core processing element 110 and a first subset of auxiliary processing elements 1201-1203 may be configured to execute a first application 410. A second portion of core processing element 110 and a second subset of auxiliary processing elements 1204-1209 may be configured to execute a second application 420, such as Figure 4A In another example, a portion of a set of auxiliary processing elements 1201-1203 may be shared between the first and second applications 430, 440, as shown. Figure 4BIn another example, the auxiliary processing elements 1201-1205 for executing the first application 450 may have a different configuration than the auxiliary processing elements 1206-1208 for executing the second application 460, as shown in FIG. Figure 4C As shown. Figure 4C Different configurations for interconnecting auxiliary processing elements 120 and core processing elements 110 are described, but these configurations may also differ in terms of memory capacity, memory bandwidth, logical functionality, etc. Furthermore, the above examples are non-limiting and are intended to illustrate that core processing element 110 and multiple auxiliary processing elements 120 may be configured in any number of different configurations to execute one or more applications in one or more modes.
[0031] Some aspects of the present disclosure advantageously provide a configurable processing unit. The architecture of the processing unit can be advantageously configured and reconfigured for specific applications. The configurable architecture provides a flexible processing unit for executing multiple applications and / or execution modes. The processing unit also provides improved performance for a given set of one or more applications by configuring the architecture based on a specific set of one or more applications. Some aspects of the present disclosure can also advantageously reduce costs by providing a configurable architecture that can be optimized for specific applications or for a wider range of applications.
[0032] The foregoing descriptions of specific embodiments of the present disclosure have been presented for purposes of illustration and description. They are not intended to be exhaustive or to limit the disclosure to the precise forms disclosed, and many modifications and variations are apparent in light of the above teachings. The embodiments are chosen and described in order to best explain the principles of the present disclosure and its practical application, thereby enabling others skilled in the art to best utilize the present disclosure and various embodiments with various modifications suitable for the specific use contemplated. The scope of the invention is intended to be defined by the appended claims and their equivalents.
Claims
1. A processing unit, wherein: The processing unit includes: a core processing element comprising processing logic and non-volatile memory having configurable compute resources including graphics computing capabilities and memory resources including configurable data flow; a plurality of auxiliary processing elements, each auxiliary processing element comprising processing logic having configurable memory resources and connection topology and non-volatile memory, the non-volatile memory of the auxiliary processing element storing a configuration bitstream; one or more configurable mesh networks configured to dynamically couple the plurality of auxiliary processing elements to the core processing element, a connection topology of the mesh networks being dynamically adjusted based on the configuration bitstream; One or more input / output interfaces are configured to receive the configuration bitstream from a host device to configure or reconfigure the computing resources and memory resources of the core processing element and the memory resources and connection topology of the auxiliary processing element.
2. The processing unit according to claim 1, wherein: The processing logic of the core processing element includes a field programmable gate array; and The non-volatile memory of the core processing element includes flash memory.
3. The processing unit according to claim 1, wherein: The processing logic of the plurality of auxiliary processing elements comprises a field programmable gate array; and The non-volatile memory of the auxiliary processing element includes a flash memory.
4. The processing unit according to claim 1, wherein: The core processing elements are configured based on one or more configuration bitstreams; The plurality of auxiliary processing elements are configured based on the one or more configuration bitstreams; as well as The one or more networks are configured based on the one or more configuration bitstreams.
5. The processing unit according to claim 4, wherein: Respective ones of the one or more configuration bitstreams are stored in non-volatile memory of respective ones of the plurality of auxiliary processing elements and the corresponding processing element.
6. A method for configuring a processing unit, wherein: The processing unit is the processing unit according to any one of claims 1 to 5, and the method includes: access one or more configuration bitstreams; configuring core processing elements of the processing unit based on the one or more configuration bitstreams; configuring a plurality of auxiliary processing elements of the processing unit based on the one or more configuration bitstreams; configuring one or more networks coupling the plurality of auxiliary processing elements to the core processing element based on the one or more configuration bitstreams; and One or more given applications are executed on the processing unit.
7. The method according to claim 6, wherein: Accessing the one or more bitstreams includes retrieving the one or more configuration bitstreams from non-volatile memory of the plurality of auxiliary processing elements and the core processing element of the processing unit.
8. The method according to claim 6, wherein: Accessing the one or more bitstreams includes receiving the one or more configuration bitstreams on one or more input / output interfaces of the processing unit.
9. The method according to claim 8, wherein Accessing the one or more bitstreams includes receiving the one or more configuration bitstreams from a host device at one or more input / output interfaces of the processing unit.
10. The method according to claim 6, wherein: Configuring the core processing element includes configuring computational logic of the core processing element based on the one or more configuration bitstreams.
11. The method according to claim 10, wherein: Configuring the core processing element includes one or more of the following: configuring allocation of one or more computing resources of the computing logic; configuring one or more functions of the computing logic; Configuring a node sampling strategy for the computational logic; A set of operating instructions for configuring the computing logic; and Configuring memory resources of the core processing element.
12. The method according to claim 11, wherein Configuring the core processing element includes allocating computing resources for a single resident graph neural network application or a plurality of resident graph neural network applications based on the one or more configuration bitstreams.
13. The method according to claim 11, wherein Configuring the memory resources of the core processing element includes configuring a plurality of memory channels in the core processing element based on the one or more configuration bitstreams.
14. The method according to claim 11, wherein Configuring the memory resources of the core processing element includes configuring memory capacity in the core processing element based on the one or more configuration bitstreams.
15. The method according to claim 6, wherein Configuring the plurality of auxiliary processing elements includes one or more of the following: configuring memory management of the plurality of auxiliary processing elements; and One or more network interfaces of the plurality of auxiliary processing elements are configured.
Citation Information
Patent Citations
Multicore wireless and media signal processor (MSP)
US20080288728A1