CMOS synaptic array with linear weight updateability independent of cell position

By using cross-switch synaptic array units and charge sharing technology of CMOS transistors in the synaptic array, the problem of inconsistent weight update characteristics caused by unit position is solved, and faster operation speed and consistent weight update are achieved.

CN112805783BActive Publication Date: 2025-09-30INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980065342.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-26
Filing Date
2019-10-02
Publication Date
2025-09-30
Estimated Expiration
2039-10-02

AI Technical Summary

Technical Problem

In the prior art, the weight update characteristics in the synaptic array are inconsistent due to the unit position, resulting in different weight update characteristics across the unit array, affecting the operation efficiency and consistency.

Method used

A crossbar synaptic array unit using complementary metal oxide semiconductor (CMOS) transistors uses non-overlapping pulses through charge sharing technology to control the gate voltage of the CMOS transistors to achieve linear update of weights.

Benefits of technology

Almost the same weight update characteristics are maintained across the entire cell array, which improves the operation speed and simplifies the circuit design by eliminating the need for a global bias circuit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112805783B_ABST
    Figure CN112805783B_ABST
Patent Text Reader

Abstract

A neuromorphic circuit (500) includes a crossbar synapse array unit. The crossbar synapse array unit includes a complementary metal oxide semiconductor (CMOS) transistor (T6), the on-resistance of which is controlled by the gate voltage of the CMOS transistor (T6) to update the weight of the crossbar synapse array unit. The neuromorphic circuit (500) also includes a set of row lines, each of which connects the synapse array unit in series with a plurality of presynaptic neurons at a first end of the synapse array unit. The neuromorphic circuit (500) also includes a set of column lines, each of which connects the synapse array unit in series with a plurality of postsynaptic neurons at a second end of the synapse array unit. The gate voltage of the CMOS transistor (T6) is controlled by performing a charge sharing technique, wherein the charge sharing technique uses non-overlapping pulses on a cell control line aligned with the set of row lines and the set of column lines to update the weight of the crossbar synapse array unit.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art Technical Field

[0002] The present invention relates generally to machine learning and, more particularly, to CMOS synaptic arrays with updateable linear weights that are independent of cell position.

[0003] Description of Background Art

[0004] Analog multiply-add accelerators have gained significant interest due to their power efficiency. Various synaptic elements are being developed, such as PCM, RRAM, MRAM, and others. Analog transistors are one option for synaptic elements due to their linearity in the triode region. However, conventional techniques unintentionally generate different pulse shapes at the proximal and distal ends of the synaptic array. This, in turn, results in different weight update characteristics across the cell array. Therefore, a solution is needed to address the variations in weight update characteristics due to cell position in the synaptic array. Summary of the Invention

[0005] According to one aspect of the present invention, a neuromorphic circuit is provided. The neuromorphic circuit includes a crossbar synapse array unit. The crossbar synapse array unit includes a complementary metal oxide semiconductor (CMOS) transistor, the on-resistance of the CMOS transistor being controlled by the gate voltage of the CMOS transistor to update the weight of the crossbar synapse array unit. The neuromorphic circuit also includes a set of row lines, each of which connects the synapse array unit in series with a plurality of presynaptic neurons at a first end of the synapse array unit. The neuromorphic circuit also includes a set of column lines, each of which connects the synapse array unit in series with a plurality of postsynaptic neurons at a second end of the synapse array unit. The gate voltage of the CMOS transistor is controlled by performing a charge sharing technique, wherein the charge sharing technique uses non-overlapping pulses on a cell control line aligned with the set of row lines and the set of column lines to update the weight of the crossbar synapse array unit.

[0006] According to another aspect of the present invention, a neuromorphic chip is provided. The neuromorphic chip includes a synaptic array formed by cross-switch synaptic array cells. Each cross-switch synaptic array cell includes a complementary metal oxide semiconductor (CMOS) transistor having an on-resistance controlled by a gate voltage of the CMOS transistor to update each weight of the cross-switch synaptic array cell. Each cross-switch synaptic array cell also includes a set of row lines that respectively connect the synaptic array cell in series with a plurality of pre-synaptic neurons at a first end of the synaptic array cell. Each cross-switch synaptic array cell also includes a set of column lines that respectively connect the synaptic array cell in series with a plurality of post-synaptic neurons at a second end of the synaptic array cell. The gate voltages of the CMOS transistors are controlled by performing a charge sharing technique that uses non-overlapping pulses on cell control lines aligned with the set of row lines and the set of column lines to update the weights of the cross-switch synaptic array cell.

[0007] According to another aspect of the present invention, a method is provided. The method includes forming a crossbar synapse array cell, the crossbar synapse array cell including complementary metal oxide semiconductor (CMOS) transistors, the CMOS transistors having an on-resistance controlled by a gate voltage of the CMOS transistors, to update a weight of the crossbar synapse array cell. The method also includes forming a set of row lines, the set of row lines respectively connecting the synapse array cell in series with a plurality of pre-synaptic neurons at a first end of the synapse array cell. The method also includes forming a set of column lines, the set of column lines respectively connecting the synapse array cell in series with a plurality of post-synaptic neurons at a second end of the synapse array cell. The gate voltages of the CMOS transistors are controlled by performing a charge sharing technique, the charge sharing technique using non-overlapping pulses on cell control lines aligned with the set of row lines and the set of column lines to update the weight of the crossbar synapse array cell.

[0008] These and other features and advantages will become apparent from the following detailed description of illustrative embodiments thereof, which is to be read in connection with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The following description will provide details of the preferred embodiments with reference to the following drawings, in which:

[0010] Figure 1 is a block diagram illustrating a processing system to which the present invention may be applied;

[0011] Figure 2 is a block diagram illustrating an environment in which the present invention may be applied;

[0012] Figure 3is a block diagram illustrating another environment in which the present invention may be applied;

[0013] Figure 4 is a diagram illustrating waveforms at the proximal and distal ends of a synapse array according to an embodiment of the present invention;

[0014] Figure 5 is a block diagram illustrating a neuromorphic cell circuit according to an embodiment of the present invention;

[0015] Figure 6 is shown with Figure 5 A timing diagram of an exemplary timing sequence associated with a circuit;

[0016] Figure 7 is shown with Figure 5 A timing diagram of another timing sequence related to the circuit;

[0017] Figure 8 It is shown by Figure 5 A block diagram of a neuromorphic array formed by multiple circuits;

[0018] Figure 9 is shown applied to Figure 8 Timing diagram of the pulses of the array;

[0019] Figure 10 is a flow chart illustrating a method for a linearly weighted updateable CMOS synapse array without cell position dependency according to an embodiment of the present invention;

[0020] Figure 11-12 is a graph for Case 1-PA, which increases the gate voltage of transistor T6 and uses a pulse with a width of 0.8 ns at 500 nA;

[0021] Figure 13-14 is a graph for Case 1-PA, which increases the gate voltage of transistor T6 and uses a pulse with a width of 1.6 ns at 250 nA;

[0022] Figure 15-16 is a graph for Case 1-PA, which increases the gate voltage of transistor T6 and uses a pulse with a width of 3.2 ns at 125 nA;

[0023] Figure 17-18 is a graph for Case 1-PA, which increases the gate voltage of transistor T6 and uses a pulse with a width of 6.4 ns at 62.5 nA;

[0024] Figure 19-20 is a graph for Case 1-PA, which increases the gate voltage of transistor T6 and uses a pulse with a width of 12.8 ns at 31.25 nA;

[0025] Figure 21-22 is a graph for Case 1-PA, which increases the gate voltage of transistor T6 and uses a pulse with a width of 25.6 ns at 15.625 nA;

[0026] Figure 23-25 is a graph for case 1-PI, which increases the gate voltage of transistor T6 and uses a pulse with a width of 1.3 ns;

[0027] Figures 26-27 is a graph for Case 1-PA, which reduces the gate voltage of transistor T6 and uses a pulse with a width of 0.8 ns at 500 nA;

[0028] Figures 28-29 is a graph for case 1-PA, which reduces the gate voltage of transistor T6 and uses a pulse with a width of 1.6 ns at 250 nA;

[0029] Figure 30-31 is a graph for case 1-PA, which reduces the gate voltage of transistor T6 and uses a pulse with a width of 3.2 ns at 125 nA;

[0030] Figures 32-33 is a graph for Case 1-PA, which reduces the gate voltage of transistor T6 and uses a pulse with a width of 6.4 ns at 62.5 nA;

[0031] Figures 34-35 is a graph for Case 1-PA, which reduces the gate voltage of transistor T6 and uses a pulse with a width of 12.8 ns at 31.25 nA;

[0032] Figures 36-37 is a graph for Case 1-PA, which reduces the gate voltage of transistor T6 and uses a pulse with a width of 25.6 ns at 15.625 nA;

[0033] Figures 38-40 is a graph for case 1-PI, which reduces the gate voltage of transistor T6 and uses a pulse having a width of 3.1 ns;

[0034] Figure 41 is a block diagram illustrating a cloud computing environment having one or more cloud computing nodes with which a local computing device used by a cloud consumer communicates according to an embodiment of the present invention; and

[0035] Figure 42 is a block diagram illustrating a set of functional abstraction layers provided by a cloud computing environment according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The present invention relates to a CMOS synaptic array with linear weights that are updateable without cell position dependency.

[0037] In one embodiment, charge sharing technique is used for weight updates by using non-overlapping pulses from vertical and horizontal control lines. In one embodiment, the present invention involves adding a pFET and an nFET in the synaptic unit cell, which enables the charge sharing technique and provides the resulting fast overall operation.

[0038] Conventionally, the difference in pulse shape at the near and far ends can be suppressed by using a wider pulse width, which undesirably reduces the operating speed. In contrast, the present invention advantageously increases the operating speed and maintains nearly identical weight update characteristics across the entire cell array.

[0039] Furthermore, the present invention does not require a global bias circuit and its distribution.

[0040] Furthermore, the present invention is easy to implement because maintaining non-overlapping pulses that preserve the pulse shape across the cell array is relatively easy from a circuit design point of view.

[0041] These and other advantages among the many attendant advantages of the present invention may be readily ascertained by those skilled in the art given the teachings of the invention provided herein while maintaining the spirit of the invention.

[0042] Figure 1 1 is a block diagram illustrating an exemplary processing system 100 to which the present invention may be applied. Processing system 100 includes a set of processing units (e.g., CPUs) 101, a set of GPUs 102, a set of memory devices 103, a set of communication devices 104, and a set of peripheral devices 105. CPU 101 may be a single-core or multi-core CPU. GPU 102 may be a single-core or multi-core GPU. One or more memory devices 103 may include cache, RAM, ROM, and other memories (flash memory, optical memory, magnetic memory, etc.). Communication device 104 may include wireless and / or wired communication devices (e.g., network (e.g., WIFI, etc.) adapters, etc.). Peripheral devices 105 may include display devices, user input devices, printers, imaging devices, etc. The elements of processing system 100 are connected via one or more buses or networks (collectively represented by reference numeral 110 in the drawing).

[0043] Of course, as one skilled in the art will readily appreciate, the processing system 100 may also include other elements (not shown), as well as omit certain elements. For example, as one skilled in the art will readily appreciate, various other input devices and / or output devices may be included in the processing system 100, depending on the specific implementation of the processing system 100. For example, different types of wireless and / or wired input and / or output devices may be used. Furthermore, as one skilled in the art will readily appreciate, additional processors, controllers, memories, etc. in different configurations may also be utilized. Further, in another embodiment, a cloud configuration (e.g., see Figure 9-10 These and other variations on processing system 100 will be readily contemplated by one of ordinary skill in the art given the teachings of the invention provided herein.

[0044] Furthermore, it should be understood that the different figures as described below with respect to different elements and steps related to the present invention may be implemented in whole or in part by one or more of the elements of the system 100 .

[0045] A description will now be given of two exemplary environments 200 and 300 to which the present invention may be applied according to various embodiments of the present invention. Figure 2 and 3 Describe environments 200 and 300. In more detail, environment 200 includes a learning-based prediction system operatively coupled to a controlled system, while environment 300 includes a learning-based prediction system as part of the controlled system. In addition, either environment 200 or 300 can be part of a cloud-based environment (e.g., see Figure 7 and 8 ). Given the teachings of the present invention provided herein, those skilled in the art can readily determine these and other environments to which the present invention can be applied while maintaining the spirit of the present invention.

[0046] Figure 2 is a block diagram illustrating an environment 200 in which the present invention may be applied.

[0047] The environment 200 includes a learning-based prediction system 210 and a controlled system 220. The learning-based prediction system 210 and the controlled system 220 are configured to enable communication between them. For example, transceivers and / or other types of communication devices, including wireless, wired, and combinations thereof, may be used. In an embodiment, communication between the learning-based prediction system 210 and the controlled system 220 may be performed over one or more networks (collectively represented by reference numeral 230 in the figure). The communication may include, but is not limited to, multivariate time series data from the controlled system 220, and predictions and action initiation control signals from the learning-based prediction system 210. The controlled system 220 may be any type of processor-based system, such as, but not limited to, a banking system, an access system, a surveillance system, a manufacturing system (e.g., an assembly line), an advanced driver assistance system (ADAS), and the like.

[0048] The controlled system 220 provides data (eg, multivariate time series data) to the learning-based forecasting system 210, which uses the data to make forecasts.

[0049] In one embodiment, to make predictions, the learning-based prediction system 210 may use neuromorphic circuits as described herein.

[0050] The controlled system 220 can be controlled based on the predictions generated by the learning-based prediction system 210. For example, based on a prediction that a machine will fail in x time steps, a corresponding action (e.g., powering off the machine, enabling machine safety measures to prevent injury, etc.) can be performed at t<x to prevent the failure from actually occurring. As another example, based on a predicted trajectory of an intruder, a controlled surveillance system can lock or unlock one or more doors to secure a person in a certain location (holding area) and / or direct them to a safe place (safe room) and / or confine them to a restricted area, etc. Verbal (from a speaker) or displayed (on a display device) instructions can be provided along with the locking and / or unlocking of the doors (or other actions) to direct the person. As another example, a vehicle can be controlled (braking, steering, accelerating, etc.) to avoid an obstacle predicted to be in the vehicle's path in response to the prediction. As yet another example, the present invention can be incorporated into a computer system to predict an impending failure and take action before the failure occurs, such as switching a component that is about to fail with another component, routing it through a different component, processing it through a different component, etc. It should be understood that the aforementioned actions are merely illustrative, and thus, other actions may be performed depending on the embodiment, as readily understood by one of ordinary skill in the art given the teachings of the present invention provided herein, while maintaining the spirit of the present invention.

[0051] In an embodiment, the learning-based prediction system 210 can be implemented as a node in a cloud computing arrangement. In an embodiment, a single learning-based prediction system 210 can be assigned to a single controlled system or multiple controlled systems, for example, different robots in an assembly line, etc. Given the teachings of the present invention provided herein, while maintaining the spirit of the present invention, these and other configurations of the elements of the environment 200 are readily determined by one of ordinary skill in the art.

[0052] Figure 3 is a block diagram illustrating another exemplary environment 300 in which the present invention may be applied, according to an embodiment of the present invention.

[0053] Environment 300 includes a controlled system 320, which in turn includes a learning-based prediction system 310. One or more communication buses and / or other devices can be used to facilitate inter-system and intra-system communication. Controlled system 320 can be any type of processor-based system, such as, but not limited to, a banking system, an access control system, a surveillance system, a manufacturing system (e.g., an assembly line), an advanced driver assistance system (ADAS), etc.

[0054] The operation of these elements in environments 200 and 300 is similar, except that system 310 is included in system 320. Therefore, for the sake of brevity, elements 310 and 320 are not compared to each other. Figure 3 In further detail, the reader is directed to elements 210 and 220 relative to Figure 2 The description of the environment 200 is given, giving the common functions of these elements in the two environments 200 and 300.

[0055] Figure 4 is a diagram illustrating exemplary waveforms 400 at the proximal end 411 and distal end 412 of a synapse array 410 , in accordance with an embodiment of the present invention.

[0056] The waveform 400 changes in shape at the distal end 412 compared to the proximal end 411 of the synapse array 410. The present invention addresses and overcomes this undesirable shape change.

[0057] Figure 5 is a block diagram illustrating a neuromorphic cell circuit 500 according to an embodiment of the present invention. Figure 6 is shown with Figure 5 A timing diagram of a timing sequence 600 associated with the circuit 500 is shown. Figure 7 is shown with Figure 5 FIG. 5 is a timing diagram of another timing sequence 700 related to the circuit 500 .

[0058] refer to Figure 5 , the circuit 500 includes two p-type metal oxide semiconductor field effect transistors (MOSFETs), namely T1 and T2.

[0059] The circuit 500 further includes two n-type MOSFETs, namely T3 and T4.

[0060] The circuit 500 further includes a CMOS transistor T6.

[0061] Circuit 500 further includes three capacitors, namely C1, C2, and C3. In one embodiment, capacitors C1 and C2 may be MOSFET parasitic capacitances. In another embodiment, capacitors C1 and C2 may be "intentional" CMOS capacitors.

[0062] The on-resistance controlled by the gate voltage T6 is used to update each weight of the crossbar synapse array unit.

[0063] A set of row lines respectively connects the synapse array unit in series to a plurality of presynaptic neurons at a first end thereof.

[0064] A set of column lines respectively connects the synapse array unit in series to a plurality of post-synaptic neurons at the second end thereof.

[0065] The gate voltage of the CMOS transistor T6 is updated by performing a charge sharing technique that updates the weight using non-overlapping pulses. Specifically, the charge sharing technique is performed row by row, thereby linearly updating the gate voltage in an increasing and decreasing manner using non-overlapping pulses so as to switch the vertical and horizontal control lines using different clock and capacitor combinations.

[0066] The gate voltage of T6 is updated by using two types of charge sharing. Figure 6 shows a charge sharing (incremental) for updating the gate voltage of T6, while Figure 7 Another type of charge sharing (decreasing) is shown for updating the gate voltage of T6.

[0067] refer to Figure 6 , by using C1 and C3 to update the gate voltage of T6 so that wclk_i and Wud_i do not overlap. That is, the increment line (Wclk_i) of the clock of transistor T1 and the update line (Wud_i) of the transient T2.

[0068] refer to Figure 7 , by using C2 and C3 to update the gate voltage of T6 so that wclk_d and Wud_d do not overlap. That is, the decrement line (Wclk_d) of the clock for transistor T4 and the update line (Wud_d) for transistor T3.

[0069] Thus, a neuromorphic chip having one or more neuromorphic units 500 performs a charge sharing technique such that (i) the increment line (Wclk_i) of the clock is switched and the increment update (Wud_i) is updated in an incremental manner using non-overlapping pulses of the clock, and (ii) the decrement line (Wclk_d) of the clock is switched and the decrement update (Wud_d) is updated in a decremental manner using non-overlapping pulses of the clock.

[0070] Figure 8 is a diagram showing an embodiment of the present invention, Figure 5 FIG. 8 is a block diagram of a neuromorphic array 800 formed by a plurality of neuromorphic unit circuits 500 . Figure 9 is shown applied to Figure 8 The timing diagram of the pulse 900 of the array 800 is shown. Figure 5 As shown, the array 800 is formed of a plurality of neuromorphic cells, where a charge sharing technique is applied such that weight updates use a row-by-row access scheme.

[0071] Figure 10 is a flow chart illustrating a method 1000 for a linear weight-updatable CMOS synapse array without cell position dependency, according to an embodiment of the present invention.

[0072] At block 1010 , a crossbar synapse array cell is formed, the crossbar synapse array cell including a CMOS transistor having an on-resistance controlled by a gate voltage of the CMOS transistor to update a weight of the crossbar synapse array cell.

[0073] At block 1020 , a set of row lines is formed that respectively connect the synapse array unit in series to a plurality of presynaptic neurons at first ends thereof.

[0074] At block 1030 , a set of column lines are formed that respectively connect the synapse array unit in series to a plurality of post-synaptic neurons at second ends thereof.

[0075] At block 1040 , gate voltages of CMOS transistors are controlled by performing a charge sharing technique that updates weights of crossbar synapse array cells using non-overlapping pulses on cell control lines aligned with the set of row lines and the set of column lines.

[0076] In one embodiment, block 1040 includes one or more of blocks 1040A and 1040B.

[0077] At block 1040A, a charge sharing technique is performed to enable updating in an incremental manner using non-overlapping pulses to switch the update delta line and the clock delta line.

[0078] In block 1040B, a charge sharing technique is performed to perform the update in a decrementing manner using non-overlapping pulses to switch the update decrement line and the clock decrement line.

[0079] It should be understood that any known fabrication technology can be used to form neuromorphic circuits and / or chips according to the teachings of the present invention. Therefore, for the sake of brevity, they will not be described in detail here.

[0080] It should be understood that the present invention can be included as part of a prediction system. The prediction system can in turn be part of another system (e.g., ADAS). Moreover, at least a portion of the prediction system can be implemented using a cloud configuration, as described in more detail below.

[0081] Figure 11-40 Graphs are provided showing exemplary experimental results obtained for a case involving a 128-cell load (hereinafter interchangeably referred to as "Case 1"). Some graphs show experimental results according to the prior art (Case 1-PA), while other graphs show experimental results according to the present invention (Case 1-PI). The prior art experimental results involve the use of a memory cell formed of 3T1C cells.

[0082] In particular, Figure 11-25 involves increasing the gate voltage of transistor T6, Figure 26-40 Involving reducing the gate voltage of transistor T6, all cases involve 128 unit loads. Various pulse widths and currents are described below.

[0083] The key point to note here is that if we use shorter pulses, the difference in weight update characteristics on the near and far ends is significant. This can be mitigated by using wider pulses in the prior art. However, the PA case requires a pulse width of at least about 25.6ns to minimize the difference, as shown in Figure 21-22 Instead, using the present invention, it is possible to Figure 23-25 This is achieved with a pulse width of 3.2 ns as shown. Thus, the present invention enables faster operation than the prior art (PA). This improvement is even more pronounced in larger arrays. Figure 26-40 Similar results are shown for the decreasing case.

[0084] Reference Figure 11-12 (Graphs 1100 and 1200, respectively), which are for case 1-PA and using pulses with a width of 0.8 ns at 500 nA.

[0085] Reference Figure 13-14 (Graphs 1300 and 1400, respectively), which are for case 1-PA and using pulses with a width of 1.6 ns at 250 nA.

[0086] Reference Figure 15-16 (Graphs 1500 and 1600, respectively), which are for case 1-PA and using pulses with a width of 3.2 ns at 125 nA.

[0087] Reference Figure 17-18 (Graphs 1700 and 1800, respectively), which are for case 1-PA and using pulses with a width of 6.4 ns at 62.5 nA.

[0088] Reference Figure 19-20 (Graphs 1900 and 2000, respectively), which are for case 1-PA and using pulses with a width of 12.8 ns at 31.25 nA.

[0089] Reference Figure 21-22 (Graph 2100 and graph 2200, respectively), which are for case 1-PA and using a pulse with a width of 25.6 ns at 15.625 nA.

[0090] Reference Figure 23-25 (Graphs 2300, 2400, and 2500, respectively), for case 1-PI and using pulses with a width of 3.7 ns.

[0091] Reference Figures 26-27 (Graph 2600 and graph 2700, respectively), which are for case 1-PA and using pulses with a width of 0.8 ns at 500 nA.

[0092] Reference Figures 28-29 (Graph 2800 and graph 2900, respectively), which are for case 1-PA and using pulses with a width of 1.6 ns at 250 nA.

[0093] Reference Figure 30-31 (Graphs 3000 and 3100, respectively), which are for case 1-PA and using pulses with a width of 3.2 ns at 125 nA.

[0094] Reference Figures 32-33 (Graphs 3200 and 3300, respectively) for Case 1-PA and using a pulse with a width of 6.4 ns at 62.5 nA.

[0095] Reference Figures 34-35 (Graph 3400 and graph 3500, respectively), which are for case 1-PA and using a pulse with a width of 12.8 ns at 31.25 nA.

[0096] Reference Figures 36-37(Graphs 3600 and 3700, respectively), which are for Case 1-PA and using pulses with a width of 25.6 ns at 15.625 nA.

[0097] Reference Figures 38-40 (Graphs 3800, 3900, and 4000, respectively) for case 1-PI and using a pulse of 5.8 ns width, respectively.

[0098] It should be understood that although the present disclosure includes detailed descriptions about cloud computing, the implementation of the teachings listed herein is not limited to cloud computing environments. Instead, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0099] Cloud computing is a service delivery model that provides convenient, on-demand network access to a shared pool of configurable computing resources. Configurable computing resources are those that can be rapidly deployed and released with minimal management overhead or interaction with the service provider. Examples include networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0100] Features include:

[0101] On-demand self-service: Cloud consumers can unilaterally and automatically deploy computing capabilities such as server time and network storage on demand without human interaction with the service provider.

[0102] Broad network access: Computing power can be accessed over the network through standard mechanisms that facilitate the use of the cloud through different types of thin-client or thick-client platforms (e.g., mobile phones, laptops, personal digital assistants (PDAs)).

[0103] Resource pooling: A provider's computing resources are grouped into resource pools and served to multiple consumers through a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated on demand. Generally, consumers cannot control or even know the exact location of the provided resources, but can specify the location at a higher level of abstraction (such as country, state, or data center), thus achieving location independence.

[0104] Rapid elasticity: The ability to quickly and elastically (sometimes automatically) deploy computing power for rapid expansion and quickly release it for rapid reduction. To consumers, the available computing power for deployment often appears to be unlimited, and any amount of computing power can be accessed at any time.

[0105] Measurable services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency for both service providers and consumers.

[0106] The service model is as follows:

[0107] Software as a Service (SaaS): The ability provided to consumers is to use the provider's applications running on the cloud infrastructure. Applications can be accessed from a variety of client devices through a thin client interface such as a web browser (e.g., web-based email). Aside from limited user-specific application configuration settings, consumers neither manage nor control the underlying cloud infrastructure, including networks, servers, operating systems, storage, or even individual application capabilities.

[0108] Platform as a Service (PaaS): The capability provided to consumers is to deploy applications they create or acquire on cloud infrastructure. These applications are built using programming languages ​​and tools supported by the provider. Consumers neither manage nor control the underlying cloud infrastructure, including networks, servers, operating systems, or storage. However, they do have control over the applications they deploy and may also have control over the configuration of the application hosting environment.

[0109] Infrastructure as a Service (IaaS): The capabilities provided to consumers are processing, storage, networking, and other basic computing resources on which they can deploy and run arbitrary software, including operating systems and applications. Consumers neither manage nor control the underlying cloud infrastructure, but do have control over the operating system, storage, and deployed applications, and may have limited control over selected network components (such as host firewalls).

[0110] The deployment model is as follows:

[0111] Private cloud: Cloud infrastructure is run solely for an organization. The cloud infrastructure can be managed by the organization or a third party and can exist inside or outside the organization.

[0112] Community Cloud: A cloud infrastructure is shared by several organizations to support a specific community with common interests (e.g., mission, security requirements, policies, and compliance considerations). A community cloud can be managed by multiple organizations within the community or by a third party and can exist within or outside the community.

[0113] Public cloud: Cloud infrastructure is provided to the public or a large industry group and is owned by the organization that sells the cloud services.

[0114] Hybrid cloud: A cloud infrastructure consisting of two or more clouds (private, community, or public) deployed in different models that remain distinct entities but are bound together by standardized or proprietary technologies (such as cloud bursting for load balancing between clouds) that enable data and application portability.

[0115] The cloud computing environment is service-oriented, with characteristics centered on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is the infrastructure consisting of a network of interconnected nodes.

[0116] Now refer to Figure 41 , depicts an illustrative cloud computing environment 4150. As shown, the cloud computing environment 4150 includes one or more cloud computing nodes 4110 with which a local computing device used by a cloud computing consumer can communicate, such as a personal digital assistant (PDA) or cellular phone 4154A, a desktop computer 4154B, a laptop computer 4154C, an automobile computer system 4154N, and / or an automobile computer system 4154N. The nodes 4110 can communicate with each other. The cloud computing nodes 4110 can be physically or virtually grouped in one or more networks including, but not limited to, private clouds, community clouds, public clouds, or hybrid clouds, or a combination thereof, as described above (not shown). In this way, cloud consumers can request Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and / or Software as a Service (SaaS) provided by the cloud computing environment 4150 without having to maintain resources on local computing devices. It should be understood that Figure 41 The various types of computing devices 4154A-N shown are merely illustrative, and the cloud computing node 4110 and cloud computing environment 4150 can communicate with any type of computing device on any type of network and / or with a network addressable connection (e.g., using a web browser).

[0117] Now refer to Figure 42 , showing the cloud computing environment 4150 ( Figure 41 ) provides a set of functional abstraction layers. First of all, it should be understood that Figure 42 The components, layers, and functions shown are merely illustrative, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0118] The hardware and software layer 4260 includes hardware and software components. Examples of hardware components include: mainframe 4261; server 4262 based on RISC (Reduced Instruction Set Computer) architecture; server 4263; blade server 4264; storage device 4265; in some embodiments, software components include network application server software 4267 and database software 4268.

[0119] The virtualization layer 4270 provides an abstraction layer from which examples of the following virtual entities can be provided: virtual servers 4271; virtual storage 4272; virtual networks 4273 (including virtual private networks); virtual applications and operating systems 4274; and virtual clients 4275.

[0120] In one example, the management layer 4280 may provide the functions described below: Resource provisioning function 4281: Provides dynamic acquisition of computing resources and other resources for performing tasks in a cloud computing environment. Metering and pricing function 4282: Tracks the cost of resource usage within the cloud computing environment and provides bills and invoices for this purpose. In one example, the resources may include application software licenses. Security function: Provides identity authentication for cloud consumers and tasks, and provides protection for data and other resources. User portal function 4283: Provides access to the cloud computing environment for consumers and system administrators. Service level management function 4284: Provides allocation and management of cloud computing resources to meet required service levels. Service level agreement (SLA) planning and fulfillment function 4285: Provides pre-arrangement and provisioning of future demand for cloud computing resources based on SLA forecasts.

[0121] The workload layer 4290 provides examples of functionality that may be implemented in a cloud computing environment. Examples of workloads or functionality that may be implemented in this layer include: mapping and navigation 4291; software development and lifecycle management 4292; virtual classroom instruction 4293; data analytics 4294; transaction processing 4295; and linearly weighted, updateable CMOS synapse arrays 4296 that are independent of cell location.

[0122] At any possible level of technical detail combination, the present invention may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0123] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0124] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0125] The computer program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and procedural programming languages ​​such as "C" language or similar programming languages. The computer readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), is personalized by utilizing the state information of the computer readable program instructions, and the electronic circuit can execute the computer readable program instructions, thereby realizing various aspects of the present invention.

[0126] Various aspects of the present invention are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0127] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0128] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0129] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the function or action of the specification, or can be implemented with a combination of dedicated hardware and computer instructions.

[0130] References in the specification to "one embodiment" or "an embodiment" of the present invention, as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" and any other variations thereof in various places throughout the specification are not necessarily all referring to the same embodiment.

[0131] It should be understood that use of any of the following " / ", "and / or", and "at least one of", for example, in the case of "A / B", "A and / or B", and "at least one of A and B", is intended to encompass selection of only the first listed option (A), or only the second listed option (B), or selection of both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to encompass selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selection of all three options (A, B, and C). As will be readily apparent to one of ordinary skill in the art, this can be expanded for many of the items listed.

[0132] Having described preferred embodiments of the systems and methods (which are intended to be illustrative and not limiting), it should be noted that modifications and variations may be made by those skilled in the art in light of the above teachings. Therefore, it should be understood that changes may be made in the specific embodiments disclosed that are within the scope of the invention as outlined by the appended claims. Having thus described various aspects of the invention with the details and particularity required by the patent laws, what is claimed and protected by Letters Patent is set forth in the appended claims.

Claims

1. A neuromorphic circuit comprising: a crossbar synapse array unit, the crossbar synapse array unit comprising a complementary metal oxide semiconductor (CMOS) transistor, the on-resistance of the CMOS transistor being controlled by a gate voltage of the CMOS transistor to update a weight of the crossbar synapse array unit; a group of row lines, wherein the group of row lines respectively connect the synapse array unit and a plurality of presynaptic neurons at a first end of the synapse array unit in series; as well as a group of column lines, each of which connects the synapse array unit to a plurality of post-synaptic neurons at a second end of the synapse array unit in series; wherein the gate voltages of the CMOS transistors are controlled by performing a charge sharing technique that updates the weights of the crossbar synapse array cells using non-overlapping pulses on cell control lines aligned with the set of row lines and the set of column lines.

2. The neuromorphic circuit of claim 1 , wherein: The crossbar synapse array unit includes three capacitors C1, C2, and C3 to update the gate voltage, and the charge sharing technique is performed row by row such that the gate voltage is updated in an incremental manner by using the capacitors C1 and C3, and is updated in a decremental manner by using the capacitors C2 and C3.

3. The neuromorphic circuit of claim 2, wherein: The neuromorphic circuit further comprises: a pair of p-type field-effect transistors (pFETs) connected in series; A pair of nFETs in series, One end of the capacitor C1 is connected to a common point between the pair of pFETs connected in series, and one end of the capacitor C2 is connected to a common point between the pair of nFETs connected in series.

4. The neuromorphic circuit of claim 3, wherein: One end of the capacitor C3 is connected to a common point between the pair of pFETs and the pair of nFETs.

5. The neuromorphic circuit of claim 2, wherein: The synapse array unit comprises: a pair of p-type field effect transistors (pFETs) connected in series for regulating the capacitors C1 and C3; and A pair of nFETs connected in series with each other and connected to the pair of pFETs are used to regulate the capacitors C2 and C3.

6. The neuromorphic circuit of claim 5, wherein: The synapse array unit further includes a connection point connecting the pair of pFETs to the pair of nFETs in series and further connected to one end of the capacitor C3 and the gate of the CMOS transistor.

7. The neuromorphic circuit of claim 5, wherein: The neuromorphic circuit performs the charge sharing technique such that the update delta line and the clock delta line are switched using the non-overlapping pulses, thereby performing updates in an incremental manner.

8. The neuromorphic circuit of claim 5, wherein: The neuromorphic circuit performs the charge sharing technique such that the update decrement line and the clock increment line are switched using the non-overlapping pulses, thereby performing updates in a decrementing manner.

9. The neuromorphic circuit of claim 5, wherein: The capacitors C1 and C2 are replaced by parasitic capacitances of the pair of pFETs and the pair of nFETs, respectively. 10 . A neuromorphic chip comprising a synapse array formed by a plurality of neuromorphic circuits, wherein the neuromorphic circuit is the neuromorphic circuit according to claim 1 .

11. A method of forming a neuromorphic circuit, comprising: forming a crossbar synapse array cell, the crossbar synapse array cell comprising a complementary metal oxide semiconductor (CMOS) transistor having an on-resistance controlled by a gate voltage of the CMOS transistor to update a weight of the crossbar synapse array cell; forming a set of row lines, wherein the set of row lines respectively connects the synapse array unit and a plurality of presynaptic neurons at the first end of the synapse array unit in series; as well as forming a set of column lines, wherein the set of column lines respectively connects the synapse array unit to a plurality of post-synaptic neurons at the second end of the synapse array unit in series, wherein the gate voltages of the CMOS transistors are controlled by performing a charge sharing technique that updates the weights of the crossbar synapse array cells using non-overlapping pulses on cell control lines aligned with the set of row lines and the set of column lines.

12. The method of claim 11, wherein: The crossbar synapse array unit is formed to include three capacitors C1, C2, and C3 to update the gate voltage, and the charge sharing technique is performed row by row such that the gate voltage is updated in an incremental manner by using the capacitors C1 and C3, and is updated in a decremental manner by using the capacitors C2 and C3.