Reuse of Processing Elements of Artificial Intelligence Processors

By introducing a hijacking control circuit into the AI ​​processor and switching the input of the MAC unit as an alternative input vector, the problem of reduced processing speed and increased power consumption caused by the inactivity of the MAC unit in the AI ​​processor is solved, and higher processing efficiency and lower power consumption are achieved.

CN112561046BActive Publication Date: 2025-06-27MICRON TECHNOLOGY INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010932806.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-10
Filing Date
2020-09-08
Publication Date
2025-06-27
Estimated Expiration
2040-09-08

AI Technical Summary

Technical Problem

Due to the inactivity of the MAC unit in AI processors, processing speed is reduced and power consumption is increased.

Method used

By introducing a hijacking control circuit, the output of the neural network layer is analyzed. If the output has not changed, the input of the MAC unit is switched as an alternative input vector, and the input vector retrieved from the remote data source.

Benefits of technology

Effectively utilizing idle MAC units improves the processing efficiency of AI processors and reduces power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112561046B_ABST
    Figure CN112561046B_ABST
Patent Text Reader

Abstract

This application relates to the reuse of processing elements of an artificial intelligence processor. The disclosed embodiments relate to an improved control circuit for an artificial intelligence processor. In one embodiment, a device is disclosed that includes: a processing element including processing means configured to receive a first set of vectors; a hijacking control circuit configured to replace the first set of vectors with a second set of vectors in response to detecting that the processing element is idle; and a Processing Element Control Circuit (PECC) that stores a set of values representing the second set of vectors, the set of values retrieved from a remote data source.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Copyright Notice

[0002] This application contains copyrighted material. The copyright owner does not object to anyone faxing copies of the patented invention in the form in which it appears in the Patent and Trademark Office files or records, but reserves all copyright rights in other respects. Technical Field

[0003] The disclosed embodiments relate to artificial intelligence (AI) processors, and more particularly, to improving the performance of such processors by reusing inactive multiply-accumulate (MAC) units. Background Art

[0004] With the increase in AI and machine learning (ML) applications, dedicated AI processors have been developed to increase the processing speed of AI algorithms such as neural network (NN) algorithms. Generally, such processors incorporate a large number of identical processing elements such as MAC units. In an NN algorithm, these MAC units are either active (processing) or inactive (idle) based on the processing output of a previous MAC unit. This is due to the multi-layer nature of neural networks. Since one or more MAC units in an AI processor are idle due to the lack of activation by a previous MAC unit, for any given clock cycle, multiple computing units of the AI processor are wasted. This inactivity results in a reduced throughput for a limited task and an increased power consumption. Summary of the Invention

[0005] The disclosed embodiments solve these and other technical problems by providing a mechanism for reusing idle processing elements such as MAC units. In the illustrated embodiments, a hijack control circuit that selectively switches the inputs of a given MAC unit is introduced into an AI processor. This hijack control circuit analyzes the output of a given neural network layer over two clock cycles. If the output does not change, then the hijack control circuit switches the input of the MAC unit to an alternative input vector.

[0006] This alternative input vector is managed by a processing element control circuit. The processing element control circuit manages a table of input vectors received from a cloud platform in some embodiments. After receiving an indication that the input of a MAC unit should be switched, the processing element control circuit selects a new set of input vectors and provides them to the MAC unit for processing.

[0007] In one embodiment, a device is disclosed that includes: a processing element including processing means configured to receive a first set of vectors; a hijacking control circuit configured to replace the first set of vectors with a second set of vectors in response to detecting that the processing element is idle; and a processing element control circuit (PECC) that stores a set of values representing the second set of vectors, the set of values being retrieved from a remote data source.

[0008] In another embodiment, a method includes: receiving, at a processing element including processing means, a first set of vectors; storing, by a processing element control circuit (PECC), a set of values representing a second set of vectors, the set of values being retrieved from a remote data source; and replacing, by a hijacking control circuit, the first set of vectors with the second set of vectors in response to detecting that the processing element is idle. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The foregoing and other objects, features, and advantages of the present invention will be apparent from the following description of embodiments as illustrated in the accompanying drawings, in which reference characters refer to the same parts throughout the views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the invention.

[0010] Figure 1 is a block diagram illustrating an artificial neuron according to some embodiments of the present invention.

[0011] Figure 2 is a block diagram of a neural network according to some embodiments of the present invention.

[0012] Figure 3 is a block diagram of a multiplier unit according to an embodiment of the present invention.

[0013] Figure 4 is a block diagram of a neural network processing element according to an embodiment of the present invention.

[0014] Figure 5 is a block diagram of a hijacking control circuit according to an embodiment of the present invention.

[0015] Figure 6 is a block diagram illustrating a processing element control circuit according to some embodiments of the present invention.

[0016] Figure 7 is a flowchart illustrating a method for reusing a processing element in an AI processor according to some embodiments of the present invention.

[0017] Figure 8 is a flowchart illustrating a method for detecting that a processing element is idle according to some embodiments of the present invention.

[0018] Figure 9is a flowchart illustrating a method for managing an input vector table according to some embodiments of the present invention. Detailed Description

[0019] Figure 1 is a block diagram of an artificial neuron according to some embodiments of the present invention.

[0020] In the illustrated embodiments, the artificial neuron includes digital processing elements that mimic the behavior of biological neurons in the human brain. Multiple similar neurons are used to form an artificial neural network, as Figure 2 illustrated. In some embodiments, the neuron includes a software construct; however, in the illustrated embodiments, the artificial neuron includes hardware processing elements.

[0021] The illustrated artificial neuron includes a processing element for processing an input vector x, the vector x including values x1, x2, … x n . Each of the values of x is associated with a corresponding weight w1, w2, … w n (forming a weight vector w). Thus, the input to the artificial neuron includes weighted inputs. In some embodiments, the artificial neuron is further configured to receive a bias input b. In some embodiments, this bias input is hard-coded as a logic high value (i.e., a value of 1). In some embodiments, the bias value is selected to provide a default output value of the neuron when all inputs are logic low (i.e., zero).

[0022] As illustrated, the artificial neuron includes a two-stage processing pipeline. During the first stage, a summator 102 is used to sum the products of the values of the input vector x and the weight vector w. In some embodiments, the sum is added with the bias input b (if implemented). Thus, the resulting value forms a scalar value provided by the summator 102 to a threshold unit 104.

[0023] As illustrated, the threshold unit 104 receives the scalar output of the summator 102 and generates an output value. Generally, the threshold unit 104 compares the output with a threshold and if the value exceeds the threshold, then outputs a first constant value (i.e., 1) and if the value is less than the threshold, then outputs a second constant value (i.e., 0). In some embodiments, the threshold unit 104 utilizes a linear function, a sigmoid function, a hyperbolic tangent function, a rectified linear function, or other types of functions. In some embodiments, the threshold unit 104 may also be referred to as an activation unit or an activation function.

[0024] Figure 2 is a block diagram of a neural network according to some embodiments of the present invention.

[0025] As illustrated, the neural network includes multiple artificial neurons connected in multiple layers ( Figure 1as described). In the described embodiment, a three-layer neural network is described. The network includes an input layer 202 and an output layer 206. The input layer and the output layer generally do not include artificial neurons, but rather handle the input and output of vectors, respectively.

[0026] In contrast, the network includes three "hidden" layers 204a, 204b, 204c, each including six artificial neurons. Each neuron receives the output of each processing element in the earlier layer. These layers are called hidden layers because they are generally not accessible to other processing elements (or users). The described network depicts three hidden layers each having six neurons, however, any number of hidden layers and any number of neurons per layer can be used.

[0027] In the described embodiment, the input vector is provided to the hidden layers 204a, 204b, 204c via the input layer 202. Each value of the input vector is provided to each neuron of the first hidden layer 204a that calculates a temporary vector, and this temporary vector is transmitted to the second layer 204b that performs a similar operation and publishes a second temporary vector to the final hidden layer 204c. The final hidden layer 204c processes the second temporary vector and generates an output vector that is transmitted to the output layer 206. In the described embodiment, the neural network is trained to adjust the values of the weights used in each layer. Further details of the neural network are not provided herein and any existing or future neural network having a similar structure can also be used.

[0028] Figure 3 is a block diagram of a multiplier unit according to an embodiment of the present invention.

[0029] In the described embodiment, the multiplier unit receives the layer output (from a hidden layer or an input layer) via a register 302 driven by a system clock. The register 302 stores the layer output for a predetermined number of clock cycles. As described, the layer output (x) is then combined with a weight vector (w) that is associated with a given layer implemented by the multiplier unit. Then, the two vectors (x, w) are provided to the multiplier 304.

[0030] In response, the multiplier 304 executes the input vectors and multiplies the input vectors. In the described embodiment, this multiplication includes multiplying each value of the input vector by the corresponding weight value. Thus, if x = {x1, x2,... x n} and w = {w1, w2,... w n}, then the resulting output vector y includes y = {x1w1, x2w2,... x n w n}.

[0031] In some embodiments, multiplier 304 may include a Booth multiplier, although other multipliers may be used. In some embodiments, multiplier 304 includes a plurality of individual multipliers to multiply each value of input vector x by weight vector w. Thus, the input vector and weight vector (x, w) are 64-bit vectors, and multiplier 304 may include 64 individual multipliers. As described above but not shown, the multiplied values are then accumulated (or added) to generate an output value for further processing by a subsequent layer.

[0032] Figure 4 is a block diagram of a neural network processing element according to one embodiment of the present invention.

[0033] As in Figure 3 the output (input vector x) of a neural network layer is received at register 402 operating according to a system clock. The combined input vector (x) and weight vector (w) are then transmitted to processing element 404.

[0034] Compared with Figure 3 the illustrated processing element 404 includes two multiplexers 408, 410 and register 412. Additionally, processing element 404 receives signals from hijack control circuit 416 and processing element control circuit 414. Processing element 404 is further configured to transmit data to processing element control circuit 414. As illustrated, processing element 404 may include the nth processing element among many elements implemented in an AI processor. Herein, the n subscript is omitted unless distinguishing signals is required.

[0035] The first multiplexer 408 receives two inputs. The first input includes the x and w vectors. The second input includes an external data signal including two vectors a and b 418. In one embodiment, vectors a and b 418 are the same length as the x and w vectors. As will be described, vectors a and b 418 may include any data and are not necessarily limited to vectors used in neural network processing. Generally, vectors a and b 418 include any two vectors in which element-wise multiplication is desired. The first multiplexer 408 is controlled by a hijack ctrl signal 422 generated by hijack control circuit 416. This circuit 416 is depicted (and described) in the Figure 5 description of and the disclosure is not repeated herein. Generally, the first multiplexer 408 operates to switch between neural network vectors x and w and remote vectors a and b 418.

[0036] The output of the first multiplexer 408 is transmitted to a processing device 406, which includes a multiplier in the illustrated embodiment. In one embodiment, the processing device 406 performs an element-wise multiplication (Hadamard product) on the received vectors. Thus, the multiplier depends on hijack ctrl The value of signal 422 multiplies the x and w vectors or the a and b vectors 418.

[0037] In the illustrated embodiment, the processing device 406 outputs the element-wise multiplication result to a second multiplexer 410 and a register 412. When hijack ctrl 422 is disabled, the second multiplexer 410 is configured to output the element-wise product of x and w. Additionally, the register 412 is configured to store the element-wise product of x and w for at least one clock cycle. Thus, when hijack ctrl 422 is disabled, the processing element 404 moves to normal processing of the neural network layer vectors.

[0038] When hijack ctrl 422 is enabled, the second multiplexer 410 uses the output of the register 412 as the output value of the processing element 404. Since the register latches the previously computed output value (the element-wise product of x and w), the processing element simulates the repeated computation of this product while computing the element-wise product of a and b 418, as previously described.

[0039] As illustrated, the output of the processing device 406 is also wired to a processing element control circuit 414 to provide the output value y 420 to the processing element control circuit 414 in each clock cycle. As will be described in Figure 6 the description of, the processing element control circuit 414 includes logic for selectively ignoring the output value computed using the x and w vectors while capturing the product of the a and b 418 vectors. Specifically, hijack ctrl 422 is used to selectively capture the product of the a and b vectors 418. Although the description generally assumes that all operations require one clock cycle, the circuit is not limited to such embodiments. In fact, based on the cycle rate of the processing device 406 or other components, a clock divider can be used to latch values for longer than one clock cycle.

[0040] Furthermore, although the foregoing description describes the use of the processing device 406, other processing devices (e.g., adders) can be used in a similar manner. In some embodiments, the processing device 406 may alternatively only require one input. In these embodiments, the processing device 406 will include a unary processing element and the input to the multiplexer 408 will include a single vector. For example, in these embodiments, the processing device 406 may include a shifter or a similar unary device.

[0041] In the illustrated circuit, processing element 404 multiplies the input vectors from the neural network layer while latching the values. When processing element 404 is switched to process alternative values, the register that latched the previous neural network output is used to drive the output while the output of the multiplier is routed to processing element control circuit 414. Thus, processing element 404 can be reused to perform multiplication operations on arbitrary inputs when it would otherwise be idle.

[0042] Figure 5 is a block diagram of a hijack control circuit according to one embodiment of the present invention.

[0043] In the illustrated embodiment, register 502 performs the same function as register 402 in Figure 4 and the description thereof is not repeated herein.

[0044] As illustrated, the output of register 502 is transmitted to processing element 404 as described in Figure 4 . Additionally, the register 502 output, which includes the neural network layer output, is also transmitted simultaneously to second register 504 and the first input of comparator circuit 506. Comparator circuit 506 has a second input connected to the output of register 504. This register 504 and comparator circuit 506 constitute hijack control circuit 500.

[0045] As illustrated, comparator circuit 506 effectively compares the current layer output with the previous layer output. In the illustrated embodiment, the previous layer output is stored in register 504 over a predetermined number of clock cycles. Thus, comparator circuit 506 determines whether the layer output has changed within a given interval.

[0046] In the illustrated embodiment, comparator circuit 506 is configured to raise hijack ctrl 508 when inputs A and B are equal and otherwise maintain hijack ctrl 508 in a low state. Thus, when comparator circuit 506 detects that the current layer output and the previous layer output are the same, the comparator circuit detects that processing element 404 should be idle and raises hijack ctrl 508 to switch the input to the processing means of processing element 404 as described above.

[0047] As an example, during a multi-stage neural network, artificial neuron clusters are typically not activated by the activation function. Thus, during a stage of neural network processing, a subset of the artificial neurons is receiving static (zero) inputs. Hijack control circuit 500 detects this condition by using a register to buffer the layer output and detecting an invariant layer output. Upon detecting such idleness, hijack control circuit raises a signal that transfers the input to one or more processing means to an external input.

[0048] Figure 6 is a block diagram illustrating a processing element control circuit according to some embodiments of the present invention.

[0049] In the illustrated embodiment, a processing element control circuit (PECC) 600 is communicatively coupled to a processing element 616. The connection between the PECC 600 and the processing element 616 is described in more detail in Figure 4 and is not repeated herein.

[0050] As illustrated, the interconnect between the PECC 600 and the processing element 616 includes a plurality of tri-buses. Each bus includes a hijack control signal (hijack ctrli ), an input vector data bus (a i , b i ), and an output data bus (y i ). As discussed above, the hijack control signal is generated by the processing element and includes control signals used by the processing unit control logic 610 to identify which processing units 616 are idle and available for processing.

[0051] The processing element control logic 610 is configured to monitor the various buses to detect when the hijack control signal is active, indicating that one or more processing elements 616 are available for processing. In response to detecting an available element, the processing element control logic 610 retrieves two input vectors 604, 606 from a table 602 of stored input vectors. The processing element control logic 610 then transmits the input vectors 604, 606 to the processing element via the bus. In some embodiments, this bus includes the input vector data bus associated with the elevated hijack control signal. In other embodiments, the processing element control logic 610 manages an internal table of available processing elements and simply selects a different processing element input vector data bus.

[0052] After transmitting the input vectors, the processing element control logic 610 records which processing element received the input vectors and waits for the result on the output data bus. Once a change in the value of the output data bus is detected, the processing element control logic 610 records the returned data in the table 602 as the corresponding output result 608.

[0053] In addition to the processing element control logic 610, the PECC 600 further includes cloud interface logic 612. The cloud interface logic 612 serves as a network interface between the PECC 600 and one or more remote data sources 614. These remote data sources 614 may include cloud computing services or may include any other remote computing system. In some embodiments, the cloud interface logic 612 provides an external application programming interface (API) that allows the remote data source 614 to upload input vectors to the table 602. In other embodiments, the cloud interface logic 612 actively extracts input data vectors from the remote data source 614 for insertion into the table 602. As described, there is no limitation on the type of data represented by the input data vectors stored in the table 602.

[0054] In the illustrated embodiment, the cloud interface logic 612 monitors and manages the table 602. In some embodiments, the cloud interface logic 612 determines when a given row contains both input values 604, 606 and output value 608. When all three values are present, the cloud interface logic 612 may identify the computation as complete and upload the output value to the remote data source 614. In some embodiments, the cloud interface logic 612 stores its own internal table mapping values 604, 606, 608 to a specific endpoint in the remote data source, thus enabling the cloud interface logic 612 to return data to the remote data source that provided the input values.

[0055] In the illustrated embodiment, some or all of the elements may be implemented as circuitry. However, in other embodiments, the various components may be implemented as a combination of hardware and / or firmware. For example, the cloud interface logic 612 may include a processor coupled to embedded firmware that provides the functionality of managing the table 602.

[0056] Figure 7 is a flowchart illustrating a method for reusing processing elements in an AI processor according to some embodiments of the present invention.

[0057] In block 702, the method 700 normally operates the processing element. In one embodiment, the processing element includes Figure 4 the processing element described in the description of, the disclosure of which is not repeated herein. In one embodiment, normally operating the processing element includes multiplying values or vectors output by artificial neurons in a neural network layer.

[0058] In block 704, the method 700 determines whether the processing element is powered. Block 704 is mainly illustrated as terminating the method 700 and is not intended to be restrictive.

[0059] In block 706, when the processing element is powered, the method 700 determines whether the processing element is active or idle. As described above, this block 706 is performed by a dedicated hijacking control circuit communicatively coupled to a given processing element. InFigure 5 This circuit is described in the description of Figure 5 , the disclosure of which is not repeated herein. Generally, method 700 determines whether the input value to the processing element changes across at least one clock cycle. If so, then method 700 determines that the processing element is active and continues to process the normal input at block 702. However, if the input value does not change, then method 700 determines that the processing element is idle and proceeds to block 708.

[0060] At block 708, method 700 re-uses the input element with a set of alternative input vectors.

[0061] In the illustrated embodiment, the set of alternative input vectors includes a set of input vectors retrieved from a remote data source. In one embodiment, block 708 is performed by a dedicated processing element control circuit more fully described in the description of Figure 6 , the disclosure of which is not repeated herein. Briefly, at block 708, method 700 detects that the processing element is idle, retrieves a set of input vectors and routes the alternative input vectors to the input of the processing element. Then, the resulting computation is stored, as more fully described in Figure 9 . Figure 6 In the description of Figure 6 , the disclosure of which is not repeated herein. Briefly, at block 708, method 700 detects that the processing element is idle, retrieves a set of input vectors and routes the alternative input vectors to the input of the processing element. Then, the resulting computation is stored, as more fully described in Figure 9 . Figure 9 As more fully described in Figure 9 .

[0062] Figure 8 is a flowchart illustrating a method for detecting that a processing element is idle according to some embodiments of the present invention.

[0063] At block 802, method 800 stores the previous layer output value (A). As described above, method 800 may store this value A in a register connected to the output value bit line. In one embodiment, value A includes the output of a neural network layer. That is, the output value A may include a vector generated by one or more artificial neurons.

[0064] At block 804, method 800 receives the current layer output value (B). In one embodiment, value B includes the value of a neural network layer computed for the current clock cycle. In some embodiments, the output of the neural network layer is directly connected to a comparator (which performs block 806) and the output received at block 804 is received by this comparator.

[0065] At block 806, method 800 compares the values of A and B. If the values are not equal, then method 800 determines that the processing element associated with the layer output is active. If method 800 determines that the values are equal, then method 800 determines that the processing element associated with the layer output is idle.

[0066] At block 808, if method 800 determines that the values of A and B are not equal, then method 800 outputs an active signal. In some embodiments, this includes driving a hijack control signal low (logic 0).

[0067] In block 810, if method 800 determines that the values ​​of A and B are equal, then method 800 outputs an idle signal. In some embodiments, this includes driving the hijack control signal high (logic 1).

[0068] Figure 9 is a flow chart illustrating a method for managing an input vector table according to some embodiments of the present invention.

[0069] In block 902, method 900 determines whether the input vector table is empty, and if so, method 900 ends. Alternatively, method 900 may continue to perform block 912 (described later). As described above, the table may include a table of input vectors (A, B) and output results (Y) generated by one or more processing elements.

[0070] In block 904, method 900 determines whether any processing element (PE) is idle and therefore available. As described above, this block may be performed by detecting whether the hijack control signal for a given processing element is raised (e.g., Figure 8 and elsewhere).

[0071] In block 906, the method 900 selects the next available PE. In one embodiment, multiple PEs may be available and idle. In this embodiment, the method 900 may select a random PE. In other embodiments, the method 900 may select a PE using a least recently used (LRU) or similar algorithm. In some embodiments, the method 900 selects a PE based on a hijack control signal (i.e., by using a PE associated with a received hijack control signal).

[0072] In block 908, method 900 selects one or more inputs from the table. In one embodiment, method 900 may randomly select inputs from the table. In other embodiments, the table may include a stack or queue and select inputs from the top or bottom of the structure, respectively. In some embodiments, each input is associated with a time to live (TTL), an expiration date, or other timing value and method 900 selects the oldest (or closest to expiration) input. In some embodiments, the input may be associated with a priority and method 900 selects the input based on the priority. In some embodiments, this results in method 900 selecting the highest priority input. In some embodiments, method 900 uses a combination of the foregoing methods.

[0073] In block 910, method 900 transmits the input to the processing element selected in block 906. In one embodiment, this block 910 includes transmitting the input value to the input of the processing element and awaiting a calculation result of the processing element.

[0074] In block 912, method 900 records the output values in a table. In one embodiment, method 900 records the output of the PE along with the corresponding input values in the table.

[0075] In block 914, method 900 manages the table. In some embodiments, this block is executed after each write. In other embodiments, it may be executed before each read. Alternatively, or in combination with the foregoing, the block may be executed periodically. Alternatively, or in combination with the foregoing, the block may be executed when the table is full or nearly full.

[0076] In some embodiments, a separate cloud interface logic device ( Figure 6 discussed in) executes block 914. Generally, method 900 may retrieve input vectors from a remote data source and insert these input vectors into the table. Additionally, method 900 monitors the output values recorded in the table to detect when new output values are input into the table. After detecting these output values, method 900 transmits the input vectors and the associated output vectors to the remote data source. Then, method 900 may remove the completed input vectors and output values from the table, clearing space for more input vectors.

[0077] However, the subject matter disclosed above may be embodied in a variety of different forms and thus, the subject matter covered or claimed is intended to be construed as not limited to any of the example embodiments set forth herein; the example embodiments are provided merely for illustrative purposes. Similarly, a reasonably broad scope of the claimed or covered subject matter is desired. Among other things, for example, the subject matter may be embodied as a method, apparatus, component, or system. Thus, an embodiment may, for example, take the form of hardware, software, firmware, or any combination thereof (other than software itself). Accordingly, the following detailed description is not intended to be limiting in nature.

[0078] Throughout the specification and claims, terms may have nuanced meanings suggested or implied by the context beyond the explicit meaning stated. Similarly, as used herein, the phrase "in one embodiment" does not necessarily refer to the same embodiment and the phrase "in another embodiment" does not necessarily refer to a different embodiment. For example, the claimed subject matter is intended to include, in whole or in part, combinations of the example embodiments.

[0079] Generally, terms may be understood, at least in part, in light of their context of use. For example, terms such as "and", "or", or "and / or" as used herein may include a variety of meanings that may depend, at least in part, on the context in which such terms are used. Generally, "or" (if used in an associative list, such as A, B, or C) is intended to mean A, B, and C (used herein in an inclusive sense) as well as A, B, or C (used herein in an exclusive sense). Additionally, the term "one or more", as used herein, may, at least in part, depend on the context, be used to describe any feature, structure, or characteristic in a singular sense, or may be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, terms such as "a", "an", or "the" may again be understood to convey a singular usage or to convey a plural usage, at least in part, depending on the context. Additionally, the term "based on" may be understood to not necessarily intend to convey a set of exclusive factors and may instead, again at least in part, depend on the context and allow for additional factors that may not be explicitly described.

[0080] The present invention is described with reference to block diagrams and operational descriptions of methods and apparatuses. It should be understood that each block in the block diagrams or operational descriptions, and combinations of blocks in the block diagrams or operational descriptions, can be implemented by means of analog or digital hardware and computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer to change its functionality, a special-purpose computer, an ASIC, or other programmable data processing device, such that the instructions executed by the processor of the computer or other programmable data processing device implement the functions / actions specified in the block diagrams or operational blocks or several operational blocks. In some alternative embodiments, the functions / actions indicated in the blocks may not occur in the order indicated in the operational drawing. For example, depending on the functions / actions involved, two consecutively presented blocks may actually be executed substantially simultaneously, or sometimes the blocks may be executed in the reverse order.

[0081] These computer program instructions can be provided to the processor of the following: a general-purpose computer to change its functionality to a special purpose; a special-purpose computer; an ASIC; or other programmable digital data processing device, such that the instructions executed by the processor of the computer or other programmable data processing device implement the functions / actions specified in the block diagrams or operational blocks or several operational blocks, thereby transforming its functionality according to the embodiments herein.

[0082] For purposes of the present invention, a computer-readable medium (or computer-readable storage medium / computer-readable storage media) stores computer data, which may include computer program code (or computer-executable instructions) executable by a computer in machine-readable form. By way of example and not limitation, a computer-readable medium may include a computer-readable storage medium for tangible or fixed storage of data, or a communication medium for transient interpretation of signals carrying code. As used herein, a computer-readable storage medium refers to a physical or tangible storage device (as opposed to a signal) and includes, but is not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technology for tangible storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid state memory technologies, CD-ROM, DVD or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other physical or material medium that can be used to tangibly store the desired information, data or instructions and that can be accessed by a computer or processor.

[0083] For purposes of the present invention, a module is a software, hardware, or firmware (or combination thereof) system, process, or function or component thereof (with or without human interaction or augmentation) that performs or facilitates the processes, functions, and / or features described herein. A module may include sub-modules. The software components of a module may be stored on a computer-readable medium for execution by a processor. A module may be integrated with one or more servers, or may be loaded and executed by one or more servers. One or more modules may be grouped into an engine or an application.

[0084] Those skilled in the art will recognize that the methods and systems of the present invention may be implemented in many ways and are thus not limited by the foregoing exemplary embodiments and examples. In other words, the functional elements and individual functions performed by single or multiple components in various combinations of hardware and software or firmware may be distributed among software applications at the client level or server level or both. In this regard, any number of features of the different embodiments described herein may be combined into a single or multiple embodiments, and alternative embodiments with fewer or more than all of the features described herein are possible.

[0085] The functions may also be distributed in whole or in part among multiple components in a manner now known or later developed. Accordingly, a variety of software / hardware / firmware combinations are possible in implementing the functions, features, interfaces, and preferences described herein. In addition, the scope of the present invention encompasses conventional and known ways of implementing the described features and functions and interfaces, as well as those variations and modifications to the hardware or software or firmware components described herein that will be understood by those skilled in the art now and in the future.

[0086] In addition, embodiments of the methods presented and described in the present invention as flowcharts are provided by way of example to provide a more complete understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are envisioned, where the order of various operations is changed and where sub-operations that are described as part of a larger operation are performed independently.

[0087] Although various embodiments have been described for the purposes of the present invention, such embodiments should not be considered to limit the teachings of the present invention to those embodiments. Various changes and modifications can be made to the elements and operations described above to obtain results that remain within the scope of the systems and processes described in the present invention.

Claims

1. An apparatus, comprising: a processing element including processing means configured to receive a first set of vectors; a hijacking control circuit configured to replace the first set of vectors with a second set of vectors in response to detecting that the processing element is idle; and a processing element control circuit (PECC) that stores a set of values representing the second set of vectors, the set of values being retrieved from a remote data source, wherein detecting that the processing element is idle includes detecting that a previous output of the processing means is the same as a current output of the processing means, the hijacking control circuit includes a register and a comparator circuit, wherein a first register receives the first set of vectors; and the comparator circuit has a first input for receiving the first set of vectors and a second input connected to an output of a second register, and an output of the comparator circuit includes a signal that causes the processing element to replace the first set of vectors with the second set of vectors.

2. The apparatus according to claim 1, wherein the processing means includes a multiplier configured to multiply a set of vectors.

3. The apparatus according to claim 1, wherein the processing element further includes: a first multiplexer having an output coupled to an input of the processing means, the first multiplexer including a first input for receiving the first set of vectors, a second input for receiving the second set of vectors, and a control input connected to the hijacking control circuit; a register connected to an output of the processing means; and a second multiplexer including a first input for receiving an output of the first multiplexer, a second input for receiving an output of the register, and a control input connected to the hijacking control circuit.

4. The apparatus according to claim 1, wherein an output of the processing means is connected to an input of the processing element control circuit.

5. The apparatus according to claim 1, wherein the first set of vectors includes a set of outputs from a neural network layer.

6. The apparatus according to claim 1, wherein the PECC is configured to receive a signal indicating that the processing element is idle from the hijacking control circuit and provide the second set of vectors in response to receiving the signal.

7. The apparatus according to claim 6, wherein the PECC is further configured to select the second set of vectors from a vector table stored by the PECC.

8. The apparatus according to claim 7, wherein the PECC is further configured to insert an output value associated with the second set of vectors into the table and upload the output value to the remote data source.

9. A method, comprising: receiving, at a processing element including processing means, a first set of vectors; storing, by a processing element control circuit (PECC), a set of values representing a second set of vectors, the set of values being retrieved from a remote data source; and replacing, by a hijacking control circuit, the first set of vectors with the second set of vectors in response to detecting that the processing element is idle Wherein detecting that the processing element is idle includes detecting that a previous output of the processing device is the same as a current output of the processing device. Storing a previous set of vectors in a register at the hijacking control circuit; and Comparing, by the hijacking control circuit, the first set of vectors with the previous set of vectors using a comparator, an output of the comparator circuit including a signal that causes the processing element to replace the first set of vectors with a second set of vectors.

10. The method according to claim 9, wherein the processing device includes a multiplier configured to multiply a set of vectors.

11. The method according to claim 9, further comprising: Switching, by a multiplexer of the processing element, between the first set of vectors and the second set of vectors based on a hijacking control signal generated by the hijacking control circuit; Storing, by the processing element, an output of the processing device in a register; and Transmitting, by the processing element, the output of the processing device to a second multiplexer configured to switch between an output of the register and the output of the processing device based on the hijacking control circuit.

12. The method according to claim 9, further comprising transmitting an output of the processing device connected to the PECC.

13. The method according to claim 9, wherein the first set of vectors includes a set of outputs from a neural network layer.

14. The method according to claim 9, further comprising: Receiving, by the PECC, a signal from the hijacking control circuit indicating that the processing element is idle and providing, by the PECC in response to receiving the signal, the second set of vectors.

15. The method according to claim 14, further comprising selecting, by the PECC, the second set of vectors from a vector table stored by the PECC.

16. The method according to claim 15, further comprising: Inserting, by the PECC, an output value associated with the second set of vectors into the table and uploading the output value to the remote data source.

Citation Information

Patent Citations

  • Deep neural network processing on hardware accelerators with stacked memory

    US20160379115A1

  • Control apparatus with improved recovery from power reduction, and storage device therefor

    US5220206A

  • Digital signal processor having enhanced utilization of multiply accumulate (MAC) stage and method

    US6367003B1