Matrix Sketching Using Analog Crossbar Architecture
The analog crossbar architecture for matrix sketching addresses the computational intensity of training large-scale DNNs by using low-rank updates and probabilistic pulses, achieving efficient and energy-effective processing for deep learning workloads.
Patent Information
- Application Number
- JP2022567611
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-15
- Filing Date
- 2021-04-13
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-04-13
AI Technical Summary
Training large-scale deep neural networks (DNNs) is computationally intensive and requires significant resources, while existing hardware solutions for accelerating deep learning workloads face challenges in energy efficiency and throughput.
The method employs an analog crossbar architecture for matrix sketching, which involves updating matrices with low-rank updates, using probabilistic pulses to reset matrices to symmetric points, and transferring the sketched matrix to a digital computer for regression analysis.
This approach enables efficient and energy-effective matrix sketching, facilitating the training of large-scale DNNs by leveraging the analog domain for vector-matrix multiplication, thereby improving system throughput and reducing computational costs.
Smart Images

Figure 0007695012000001 
Figure 0007695012000002 
Figure 0007695012000003
Abstract
Description
Technical Field
[0001] The present invention generally relates to streaming algorithms, and more particularly to matrix sketching using an analog crossbar architecture.
Background Art
[0002] Deep neural networks (DNNs) have made significant progress and in some cases outperform human-level performance, tackling difficult tasks such as speech recognition, natural language processing, image classification, and machine translation. However, training large-scale DNNs is a time-consuming and computationally intensive task that requires data center-scale computing resources composed of state-of-the-art graphics processing units (GPUs). Custom hardware has been attempted to accelerate deep learning workloads beyond GPUs by designing arithmetic with reduced precision to improve the throughput and energy efficiency of the underlying complementary metal-oxide-semiconductor (CMOS) technology. As an alternative to digital approaches, resistive cross-point device arrays have been proposed to further increase the overall system throughput and energy efficiency by performing vector-matrix multiplication in the analog domain.
Summary of the Invention
[0003] According to one embodiment, a method for performing matrix sketching by utilizing an analog crossbar architecture is provided. The method includes updating a first matrix over a first time period, copying the first matrix to a dynamic correction computing device, switching to a second matrix and updating the second matrix over a second time period, supplying a first probabilistic pulse to the first matrix to reset the first matrix back to a symmetric point of the first matrix when the second matrix is updated with a low rank, copying the second matrix to the dynamic correction computing device, switching back to the first matrix and updating the first matrix over a third time period, and supplying a second probabilistic pulse to the second matrix to reset the second matrix back to a symmetric point of the second matrix when the first matrix is updated with a low rank. The update is a low rank update.
[0004] According to another embodiment, a system for performing matrix sketching by utilizing an analog crossbar architecture is provided. The system includes a memory and one or more processors communicating with the memory, and the processors are configured to update a first matrix over a first time period, copy the first matrix to a dynamic correction computing device, switch to a second matrix and update the second matrix over a second time period, supply a first probabilistic pulse to the first matrix to reset the first matrix back to a symmetric point of the first matrix when the second matrix is updated with a low rank, copy the second matrix to the dynamic correction computing device, switch back to the first matrix and update the first matrix over a third time period, and supply a second probabilistic pulse to the second matrix to reset the second matrix back to a symmetric point of the second matrix when the first matrix is updated with a low rank. The update is a low rank update.
[0005] According to yet another embodiment, a non-transitory computer-readable storage medium is presented that includes a computer-readable program for performing matrix sketching by utilizing an analog crossbar architecture. The non-transitory computer-readable storage medium performs the steps of updating a first matrix over a first time period, copying the first matrix to a dynamic correction calculation device, switching to a second matrix and updating the second matrix over a second time period, supplying a first probabilistic pulse to the first matrix to reset the first matrix back to a symmetric point of the first matrix when the second matrix is being updated with a low-rank update, copying the second matrix to the dynamic correction calculation device, switching back to the first matrix and updating the first matrix over a third time period, and supplying a second probabilistic pulse to the second matrix to reset the second matrix back to a symmetric point of the second matrix when the first matrix is being updated with a low-rank update. The update is a low-rank update.
[0006] According to one embodiment, a method for performing matrix sketching by utilizing an analog crossbar architecture is provided. The method includes applying dimensionality reduction to streaming data using outer product updates and transferring the sketched matrix to a digital computer to perform regression analysis when the dimensionality reduction is applied to the entire input. The update is a low-rank update.
[0007] According to another embodiment, a system for performing matrix sketching by utilizing an analog crossbar architecture is provided. The system includes a memory and one or more processors in communication with the memory, and the one or more processors are configured to apply dimensionality reduction to streaming data using outer product updates and transfer the sketched matrix to a digital computer to perform regression analysis when the dimensionality reduction is applied to the entire input. The update is a low-rank update.
[0008] Note that the exemplary embodiments are described with reference to multiple different subjects. In particular, some embodiments are described with reference to method-type claims, while other embodiments are described with reference to apparatus-type claims. However, those skilled in the art will appreciate from the above and following descriptions that, unless otherwise noted, any combination of features belonging to one type of subject, in addition to any combination between features related to different subjects, particularly between the features of method-type claims and the features of apparatus-type claims, is considered to be described herein.
[0009] These and other features and advantages will become apparent from the following detailed description of the exemplary embodiments, which should be read in conjunction with the accompanying drawings.
[0010] The present invention will be presented in detail in the following description of the preferred embodiments with reference to the drawings.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
[0012] Throughout the drawings, the same or similar reference numerals represent the same or similar elements.
[0013] Embodiments according to the present invention provide methods and devices for performing matrix sketching by using an analog crossbar architecture. Due to recent advances in data collection techniques, vast amounts of data are being collected at an extremely rapid pace. Also, these data may be unbounded. A boundless stream of data collected from sensors, devices, and other data sources is referred to as a data stream. Various data mining tasks can be performed on the data stream to search for patterns of interest. Data mining tasks typically involve utilizing data streaming algorithms.
[0014] A streaming algorithm can process extremely large and even unbounded data sets using only a certain amount of random access memory (RAM) and compute some desired output. If the data set is unbounded, it is referred to as a data stream. In such cases, if the processing of the data stream is stopped at some position n, the streaming algorithm has a solution corresponding to the data seen up to that point. Thus, a streaming algorithm is an algorithm for processing a data stream where the input is presented as a sequence of items and can be examined in only a few passes (usually only one pass). In most models, these algorithms can access only limited memory. They can also have a limited processing time per item. These constraints can mean that the algorithm generates an approximate answer based on a summary or "sketch" of the data stream.
[0015] Exemplary embodiments of the present invention utilize streaming algorithm-based matrix sketching by using an analog crossbar architecture. The final sketching matrix is placed within analog crossbar hardware. The outer product of the final sketching matrix is changed by applying probabilistic pulses. When the sketching procedure is applied to the entire input, the final sketched matrix is transferred to a digital computer to perform regression analysis.
[0016] Although the present invention is described with respect to a given exemplary architecture, it is to be understood that other architectures, structures, substrate materials and process features as well as steps / blocks may be varied within the scope of the present invention. Note that not all specific features can be shown in all figures for clarity. This is not intended to be construed as a limitation of any particular embodiment, or illustration, or of the claims.
[0017] Figure 1 is an analog crossbar architecture using a sketching matrix according to an embodiment of the present invention.
[0018] The analog crossbar array 100 includes a plurality of sketching resistive processing units (RPUs) 110 arranged in a matrix configuration. The outer product is changed by applying probabilistic pulses 102, 104. The outer product is implicitly calculated and changed to perform a rank-1 update. At the end of the operation, the final sketching matrix exists in the analog domain. The solution part may be implemented in either the analog or digital domain. The pulses are reduced to a simple coincidence detection 115 that can be recognized by the RPU device 110 for multiplication.
[0019] In this probabilistic update scheme, the numbers encoded from the columns and rows (x i and δ j ) are converted to a probabilistic bit stream using stochastic translators. These stochastic translators adjust the pulse probabilities at the periphery and thus control the total number of pulse coincidences occurring at each crossbar element. In this scheme, these pulses are sent to the crossbar array simultaneously for all rows and all columns, and then for each coincidence event, the corresponding RPU device changes its conductance by a small amount Δg min . However, there are many pulses in the pulse stream, and as a result, the total conductance change Δg total,ij required by the algorithm is implemented as a series of small conductance changes Δg min per pulse coincidence. As a result, the weight update is performed as a series of coincidence events each triggering a conductance increment (or decrement). Δw minis the prediction weight change due to a single matching event. The pulses generated at the periphery are applied to all RPU devices in a column (or row), and thus, when the pulse probability is calculated to bring about the desired weight change in each RPU, a single Δg min (or equivalently Δw min ) value can be assumed.
[0020] Figure 2 is an exemplary 3D crossbar array incorporating the sketching matrix of FIG. 1, according to one embodiment of the present invention.
[0021] In various exemplary embodiments, the sketching matrix 110 represents memory cells incorporated between a plurality of bit lines 122 and a plurality of word lines 124. Thus, the array 120 is obtained by orthogonal conductive word lines (rows) 124 and bit lines (columns) 122, and the sketching matrix 110 exists at the intersections between each row and column. The sketching matrix 110 by the resistive memory elements can be accessed for reading and writing by applying biases to the corresponding word lines 124 and bit lines 122.
[0022] Figure 3 is an exemplary system for analog streaming with dynamic calculations, according to one embodiment of the present invention.
[0023] System 130 includes a first matrix 132 and a second matrix 134. The first matrix 132 and the second matrix 134 communicate with a dynamic correction calculation device 136.
[0024] The first matrix 132 is updated over the first time period τ1. After being updated, the first matrix 132 is copied to the dynamic correction calculation device 136. Next, the update function is switched to the second matrix 134 over the second time period τ2. When the second matrix 134 is being updated, the first matrix 132 is supplied with a stochastic pulse to reset the first matrix 132 back to the symmetry point. Next, the second matrix 134 is copied to the dynamic correction calculation device 136. The update is switched back to the first matrix 132 over the third time period τ3. When the first matrix 132 is updated again, the second matrix 134 is supplied with a stochastic pulse to reset the second matrix 134 back to the symmetry point. This process is repeated until the final sketching matrix SXy is obtained.
[0025] By utilizing the second array 134, continuous operation becomes possible in exchange for the addition of architecture. The streaming data (X,y) should be normalized to avoid further enhancing the asymmetric effect. Further, the update should be scaled to adjust the final SXy to operate closer to the symmetry point (0) and to be between (-0.1, 0.1) over the entire range of (-1,1).
[0026] In an alternative embodiment, if the device is completely symmetric, dynamic correction becomes unnecessary. In those cases, a single array can be used without any carry / zeroing operations. Alternatively, assuming that a random sequence guarantees operation in the vicinity of the symmetry point, zeroing can be skipped. Even in such cases, a single device can be used in addition to the digital accumulation unit, and the copy operation is handled in small parts (e.g., column by column). All of these decisions can be made flexibly and interchangeably according to the immediate issues and devices at hand.
[0027] The streaming algorithm is an open-loop integration algorithm. Therefore, the dynamic system approach using the auxiliary array does not guarantee that the first matrix is reset. However, unlike neural networks, these algorithms do not rely on vector-matrix multiplication, and thus the accumulation matrix can exist in the digital domain. Therefore, instead, two arrays or matrices 132, 134 can be used in parallel in a toggling pattern.
[0028] In alternative embodiments, three or more matrices may be utilized. For example, three matrices may be utilized. In another exemplary embodiment, four matrices may be utilized. Of course, one of ordinary skill in the art can contemplate a plurality of matrices that are utilized in cooperation with the dynamic correction calculation device 136. In one example, a first matrix can be updated over a first time period τ1. After the first matrix is updated, it is copied to the dynamic correction calculation device 136. Next, the update function is switched to a second matrix over a second time period τ2. While the second matrix is being updated, the first matrix and the third matrix are supplied with stochastic pulses to reset the first matrix and the third matrix back to their respective symmetry points. Next, the second matrix is copied to the dynamic correction calculation device 136. The update is switched to the third matrix, and the third matrix can be updated over a third time period τ3. After the third matrix is updated, it is copied to the dynamic correction calculation device 136. Next, the update function is switched back to the first matrix over a fourth time period τ4. While the first matrix is being updated, the second matrix and the third matrix are supplied with stochastic pulses to reset the second matrix and the third matrix back to their respective symmetry points. Thus, one matrix can be updated and two other matrices can be supplied with stochastic pulses to reset them to their respective symmetry points. Similarly, one matrix can be updated and three other matrices can be supplied with stochastic pulses to reset them to their respective symmetry points. Thus, a plurality of matrices can be connected to one or more dynamic correction calculation devices 136, and one matrix is updated while all the other matrices are supplied with stochastic pulses to reset them to their respective symmetry points. Thus, the exemplary embodiments are not limited to the number of matrices connected to the dynamic correction calculation device 136.
[0029] FIG. 4 shows an exemplary graph illustrating three different device switching characteristics exemplifying symmetry points according to one embodiment of the present invention.
[0030] Regarding asymmetric updates, the incremental changes obtained when an analog device is updated depend on the weight values that can be represented using a softbound model. This effect causes bias and significantly degrades training performance (generation of SXy). For these devices, there exists a point called the symmetry point, Δw increment (w) = Δw decrement (w). A random sequence of operations where the increments and decrements are estimated to be equal will necessarily drive the device state to this symmetry point. This behavior is observed in streaming algorithms when the streaming algorithm contains a large number of incremental changes. The requirement is that these analog resistance devices must change their conductance symmetrically when subjected to positive or negative voltage pulse stimuli.
[0031] Figure 4 shows the switching characteristics of three different devices. In the ideal device 160, the conductance increments and decrements are equal in magnitude and do not depend on the conductance of the device. In the symmetric device 170, the conductance increments and decrements are equal in magnitude, but both depend on the conductance of the device. In the asymmetric device 180, the conductance increments and decrements are not equal in magnitude, and both have different dependencies on the conductance of the device. However, there exists a single point where the magnitudes of the conductance increments and decrements are equal. This point is called the symmetry point and, for the example shown, coincides with the reference device conductance and thus occurs at w = 0. The symmetry point 162 is shown within the ideal device 160, the symmetry point 172 is shown within the symmetric device 170, and the symmetry point 182 is shown on the asymmetric device 180.
[0032] Note that even for the asymmetric device shown in 180, there exists a single point (conductance value) where the magnitudes of the conductance increments and decrements are equal. This point is called the symmetry point of the updated device, and it can correspond to any weight value (not necessarily zero as shown in 180) due to device-to-device variations.
[0033] Figure 5 is a block / flow diagram of an exemplary method for two arrays utilized in parallel in an open-loop integration scheme, according to one embodiment of the present invention.
[0034] In block 202, update the first matrix to a low rank over a first time period.
[0035] In block 204, copy the first matrix to a dynamic correction calculation device.
[0036] In block 206, switch to the second matrix and update the second matrix to a low rank over a second time period.
[0037] In block 208, while updating the second matrix to a low rank (or at the same time, simultaneously), supply a pulse to the first matrix to reset it back to a symmetric point.
[0038] In block 210, copy the second matrix to a dynamic correction calculation device.
[0039] In block 212, switch to the first matrix and update the first matrix to a low rank over a third time period.
[0040] In block 214, while updating the first matrix to a low rank (or at the same time, simultaneously), supply a pulse to the second matrix to reset it back to a symmetric point.
[0041] Figure 6 is a block / flow diagram of an exemplary method for switching between a first matrix and a second matrix, according to one embodiment of the present invention.
[0042] In block 220, the method waits.
[0043] In block 222, it is determined whether a new sample has been received. If the answer is no, the process proceeds to block 230 where a matrix is read and the device is switched. If the answer is yes, the process proceeds to block 224.
[0044] In block 224, a vector “s” is generated.
[0045] In block 226, an analog low-rank update is performed and the process proceeds to block 220.
[0046] In block 232, the matrix is stored in digital form or in a separate analog device.
[0047] FIG. 7 is an exemplary processing system for handling streaming algorithms according to an embodiment of the present invention.
[0048] Referring now to FIG. 7, this figure shows the hardware configuration of a computing system 600 according to an embodiment of the present invention. As can be seen, this hardware configuration includes at least one processor or central processing unit (CPU) 611. The CPU 611 is interconnected via a system bus 612 to a random access memory (RAM) 614, a read only memory (ROM) 616, an input / output (I / O) adapter 618 (for connecting peripheral devices such as a disk unit 621 and a tape drive 640 to the bus 612), a user interface adapter 622 (for connecting a keyboard 624, a mouse 626, a speaker 628, a microphone 632, or other user interface devices or combinations thereof to the bus 612), a communication adapter 634 for connecting the system 600 to a data processing network, the Internet, an intranet, a local area network (LAN), etc., and a display adapter 636 for connecting the bus 612 to a display device 638 or a printer 639 (e.g., a digital printer, etc.) or both.
[0049] FIG. 8 is a block / flow diagram of an exemplary cloud computing environment in accordance with one embodiment of the present invention.
[0050] The present invention includes a detailed description of cloud computing, but it should be understood that the embodiments of the teachings described herein are not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented with any other type of computing environment now known or later developed.
[0051] Cloud computing is a service delivery model for enabling convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or service provider interaction. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0052] The characteristics are as follows.
[0053] On-demand self-service: Cloud consumers can unilaterally provision computing capabilities such as server time and network storage automatically as needed without human interaction with the service provider.
[0054] Broad network access: Capabilities are available over the network and accessed through standard mechanisms that promote use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0055] Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model where various physical and virtual resources are dynamically assigned and re-assigned according to demand. There is a sense of location independence in that consumers generally have no control or knowledge of the exact location of the resources provided, but may be able to specify a location at a higher level of abstraction (e.g., country, state, or data center).
[0056] Rapid scalability: Functions can be supplied quickly and adaptively, sometimes automatically, to scale out rapidly, release quickly, and scale in quickly. To the consumer, the functions available for supply often appear to be unlimited, and any amount can be purchased at any time.
[0057] Measurability of services: The cloud system automatically controls and optimizes resource usage by leveraging a metering function at some level of abstraction appropriate for the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both the provider and consumer of the services utilized.
[0058] The service model is as follows.
[0059] Software as a Service (SaaS): The functionality provided to consumers is to use the provider's application that runs on cloud infrastructure. The application is accessible from various client devices through a thin-client interface such as a web browser (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application features, except for, in some cases, limited user-specific application configuration settings.
[0060] Platform as a Service (PaaS): The functionality provided to consumers is to deploy the applications that consumers create or obtain, which are created using programming languages and tools supported by the provider, onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, or storage, but control the deployed applications and, in some cases, the applications that host the environment settings.
[0061] Infrastructure as a Service (IaaS): The functionality provided to consumers is to supply processing, storage, network, and other basic computing resources when consumers are able to deploy and run any software that can include an operating system and applications. Consumers do not manage or control the underlying cloud infrastructure, but control the operating system, storage, deployed applications, and, in some cases, restrictively control the selection of network connection components (e.g., host firewall).
[0062] The deployment models are as follows.
[0063] Private Cloud: The cloud infrastructure is operated solely for an organization. This infrastructure can be managed by the organization or a third party and can exist either on - premise or off - premise.
[0064] Community Cloud: The cloud infrastructure is shared by several organizations and supports a specific community that shares concerns (e.g., mission, security requirements, policies, and compliance considerations). This infrastructure can be managed by the organization or a third party and can exist either on - premise or off - premise.
[0065] Public Cloud: The cloud infrastructure is made available to the general public or a large industry group and is owned by an organization that sells cloud services.
[0066] Hybrid Cloud: The cloud infrastructure is a composition of two or more clouds (private, community, or public) that retain distinct entities but are bound together by standardized or proprietary technologies that enable portability of data and applications (e.g., cloud bursting for load balancing between clouds).
[0067] Cloud computing environments are service - oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0068] Referring now to FIG. 8, an exemplary cloud computing environment 750 enabling use cases of the present invention is shown. As illustrated, cloud computing environment 750 includes one or more cloud computing nodes 710 that can communicate with local computing devices used by cloud consumers, such as, for example, personal digital assistants (PDAs) or cellular phones 754A, desktop computers 754B, laptop computers 754C, or in-vehicle computer systems 754N, or combinations thereof. Nodes 710 can communicate with one another. These nodes can be physically or virtually grouped in one or more networks, such as private, community, public, or hybrid clouds as described above, or combinations thereof (not shown). Thereby, cloud computing environment 750 can provide infrastructure, platform, software, or combinations thereof as services, so that cloud consumers need not maintain resources on local computing devices. It is understood that the types of computing devices 754A-N shown in FIG. 8 are only intended to be exemplary, and that computing nodes 710 and cloud computing environment 750 can communicate with any type of computerized device via any type of network or network addressable connection (e.g., using a web browser), or both.
[0069] FIG. 9 is a schematic diagram of an exemplary abstraction model layer according to an embodiment of the present invention. It should be understood upfront that the components, layers, and functions shown in FIG. 9 are only intended to be exemplary, and embodiments of the present invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided.
[0070] The hardware and software layer 860 includes hardware and software components. Examples of hardware components include mainframe 861, RISC (Reduced Instruction Set Computer) architecture-based server 862, server 863, blade server 864, memory device 865, as well as network and network connection components 866. In some embodiments, the software components include network application server software 867 and database software 868.
[0071] The virtualization layer 870 provides an abstraction layer that can provide the following examples of virtual entities, namely, virtual server 871, virtual storage 872, virtual network 873 including a virtual private network, virtual applications and operating systems 874, and virtual client 875.
[0072] In one example, the management layer 880 can provide the functions described below. Resource provisioning 881 enables the dynamic procurement of computing resources and other resources used to perform tasks within a cloud computing environment. Metering and pricing 882 enables cost tracking when resources are utilized within a cloud computing environment and billing or charging for the consumption of these resources. In one example, these resources may include software license agreements for applications. Security enables the identification and verification of cloud consumers and tasks, as well as the protection of data and other resources. The user portal 883 enables consumers and system administrators to access the cloud computing environment. Service level management 884 enables the allocation and management of cloud computing resources so that the required service levels are met. Service quality assurance (SLA) planning and fulfillment 885 enables the pre-arrangement and procurement of cloud computing resources that are predicted to be required in the future by the SLA.
[0073] The workload layer 890 provides examples of functions that a cloud computing environment can utilize. Examples of workloads and functions that may be provided from this layer include mapping and navigation 891, software development and lifecycle management 892, delivery of virtual classroom education 893, data analysis processing 894, transaction processing 895, and streaming algorithms 896.
[0074] FIG. 10 is a block / flow diagram of a method for applying an open-loop integration scheme in an Internet of Things (IoT) system / device / infrastructure according to an embodiment of the present invention.
[0075] According to some embodiments of the present invention, a network is implemented using an IoT method. For example, a streaming algorithm 902 can be incorporated into, for example, wearable, embeddable, or ingestible electronic devices and Internet of Things (IoT) sensors. Wearable, embeddable, or ingestible devices may include at least health and wellness monitoring devices and fitness devices. Wearable, embeddable, or ingestible devices may further include at least embeddable devices, smart watches, head-mounted devices, security protection devices, and gaming lifestyle devices. IoT sensors can be incorporated into at least home automation applications, automotive applications, user interface applications, lifestyle or entertainment or both applications, urban or infrastructure or both applications, toys, medical, fitness, retail tags or trackers or both, platforms and components, etc. The streaming algorithm 902 described herein can be incorporated into any type of electronic device for any type of use or application or operation.
[0076] The IoT system enables users to achieve deeper automation, analysis, and integration within the system. IoT improves the reach and accuracy of these areas. IoT utilizes existing and emerging technologies in sensing, network connectivity, and robotics. The characteristics of IoT include artificial intelligence, connectivity, sensors, active engagement, and the use of small devices. In various embodiments, the streaming algorithm 902 of the present invention can be incorporated into various different devices or systems or both. For example, the streaming algorithm 902 can be incorporated into wearable or portable electronic devices 904. The wearable / portable electronic device 904 may include embeddable devices 940 such as smart clothing 943. The wearable / portable device 904 may further include a smart watch 942, and smart jewelry 945. The wearable / portable device 904 may further include a fitness monitoring device 944, a health wellness monitoring device 946, a head-mounted device 948 (e.g., smart glasses 949), a security protection system 950, a gaming lifestyle device 952, a smart phone / tablet 954, a media player 956, or a computer / computing device 958 or a combination thereof.
[0077] The streaming algorithm 902 of the present invention may be further incorporated into Internet of Things (IoT) sensors 906 for various applications, such as home automation 920, automobiles 922, user interfaces 924, lifestyle or entertainment or both 926, city or infrastructure or both 928, retail 910, tags or trackers or both 912, platforms and components 914, toys 930, or medical 932 or combinations thereof, as well as fitness 934. The IoT sensors 906 can utilize the streaming algorithm 902. Of course, those skilled in the art can contemplate incorporating such a streaming algorithm 902 into any type of electronic device for any type of application, not limited to those described herein.
[0078] FIG. 11 is a block / flow diagram of an exemplary IoT sensor used to collect data / information related to an open-loop integration scheme streaming algorithm, according to an embodiment of the present invention.
[0079] IoT loses its distinction without sensors. IoT sensors act to define devices that convert IoT from a standard passive network of devices to an active system capable of real-world integration.
[0080] The IoT sensor 906 can continuously and real - time transmit information / data to any type of distributed system via the network 908 using the streaming algorithm 902. Exemplary IoT sensors 906 may include, but are not limited to, displacement sensors 1006 such as position / presence / proximity sensors 1002, motion / speed sensors 1004, acceleration / tilt sensors 1007, temperature sensors 1008, humidity / moisture sensors 1010, and flow sensors 1011, acoustic / voice / vibration sensors 1012, chemical / gas sensors 1014, force / load / torque / strain / pressure sensors 1016, or electrical / magnetic sensors 1018 or combinations thereof. Those skilled in the art can contemplate using any combination of such sensors to collect data / information via the streaming algorithm 902 of the distributed system for further processing. Those skilled in the art can contemplate using other types of IoT sensors, including, but not limited to, magnetometers, gyroscopes, image sensors, optical sensors, radio frequency identification (RFID) sensors, or micro - flow sensors or combinations thereof. The IoT sensor can also include an energy module, a power management module, an RF module, and a sensing module. The RF module manages communications through their signal processing, WiFi(R), ZigBee(R), Bluetooth(R), wireless transceivers, duplexers, etc.
[0081] As used herein, the terms "data," "content," "information," and like terms can be used interchangeably to refer to data that can be captured, transmitted, received, displayed, stored, or any combination thereof according to various exemplary embodiments. Accordingly, the use of any such terms should not be construed as limiting the spirit and scope of the present disclosure. Further, where a computing device is described herein as receiving data from another computing device, the data can be received directly from the other computing device or indirectly via one or more intermediate computing devices such as, for example, one or more servers, relays, routers, network access points, base stations, or the like, or any combination thereof.
[0082] To enable interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, such as, for example, a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, a keyboard, and a pointing device, such as, for example, a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to enable interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback, and input received from the user can be in any form, including acoustic, speech, or tactile input.
[0083] The present invention may be a system, a method, or a computer program product, or any combination thereof. The computer program product may include a computer-readable storage medium having computer-readable program instructions for causing a processor to perform aspects of the present invention.
[0084] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as a raised structure in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, should not be construed as being a transient signal per se, such as a radio wave, or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0085] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.
[0086] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk(R), C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may utilize the state information of the computer-readable program instructions to individualize the electronic circuit in order to carry out the operations of the present invention and thereby execute the computer-readable program instructions.
[0087] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0088] These computer-readable program instructions are provided to at least one processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / operations specified in one or more blocks of a flowchart illustration and / or a block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that includes instructions for a computer, programmable data processing apparatus, or other device to function in a particular manner, such that the computer-readable storage medium stores a manufactured article including instructions for implementing the functions / operations specified in one or more blocks or modules of a flowchart and / or a block diagram.
[0089] The computer-readable program instructions may also be loaded onto a computer, other programmable apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / operations specified in one or more blocks or modules of a flowchart and / or a block diagram.
[0090] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible embodiments of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions, including one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions described in the blocks may be performed out of the order noted in the drawings. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, depending on the functions involved, or these blocks may sometimes be executed in the reverse order. It should also be noted that each block of the block diagram or flowchart, or combinations of blocks of the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or a combination of dedicated hardware and computer instructions.
[0091] References herein to "one embodiment" or "an embodiment" of the present principle, and other variations thereof, mean that the particular features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment of the present principle. Thus, the phrases "in one embodiment" or "in an embodiment" that appear in various places throughout this specification, and any other variations, are not necessarily all referring to the same embodiment.
[0092] For example, in the cases of "A / B", "A or B or both", and "at least one of A and B", it should be understood that the use of any of the following, namely " / ", "··· or ··· or both", and "at least one of ···", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, or C or combinations thereof" and "at least one of A, B, and C", such expressions are intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A, B, and C). This can be applied regardless of the number of items listed, as will be readily apparent to those skilled in the art or those skilled in the relevant technology.
[0093] Preferred embodiments of systems and methods for streaming algorithms using an analog crossbar architecture (intended to be illustrative rather than limiting) have been described, but it should be noted that those skilled in the art can make modifications and changes in light of the above teachings. Therefore, it should be understood that changes to the specific embodiments described can be made within the scope of the present invention as defined by the appended claims. Thus, aspects of the present invention have been described, including what is particularly required by the details and patent laws, but what is claimed and desired to be protected by patent is set forth in the appended claims.
Claims
1. A computer-implemented method, executed on a processor, for performing matrix sketching by utilizing an analog crossbar architecture, comprising: updating a first matrix with a low rank over a first time period; copying the first matrix to a dynamic correction calculation device; switching to a second matrix and updating the second matrix with a low rank over a second time period; when the second matrix is updated with a low rank, supplying a first probabilistic pulse to the first matrix to reset the first matrix back to a symmetric point of the first matrix; copying the second matrix to the dynamic correction calculation device; switching back to the first matrix and updating the first matrix with a low rank over a third time period; when the first matrix is updated with a low rank, supplying a second probabilistic pulse to the second matrix to reset the second matrix back to a symmetric point of the second matrix A computer-implemented method.
2. The method according to claim 1, wherein the first matrix and the second matrix are arranged within the analog crossbar architecture.
3. The method according to claim 1, wherein the first matrix and the second matrix include streaming data.
4. The method according to claim 3, wherein the streaming data is normalized to prevent an asymmetric effect.
5. The low rank updates of the first matrix and the second matrix are scaled to adjust a final sketching matrix such that it operates near symmetric points of the first matrix and the second matrix and ranges between (-0.1, 0.1) over the entire range of (-1, 1), according to the method of claim 4.
6. The method according to claim 5, wherein when the matrix sketching is applied to the entire input, the final sketching matrix is transferred to a digital computer for performing a regression analysis.
7. The method according to claim 1, wherein the dynamic correction calculation device corrects the first matrix and the second matrix simultaneously.
8. A system for performing matrix sketching by utilizing an analog crossbar architecture, comprising: a memory; one or more processors communicating with the memory, wherein the one or more processors are configured to: update a first matrix with low rank over a first time period; copy the first matrix to a dynamic correction calculation device; switch to a second matrix and update the second matrix with low rank over a second time period; supply a first probabilistic pulse to the first matrix to reset the first matrix back to a symmetric point of the first matrix when the second matrix is being updated with low rank; copy the second matrix to the dynamic correction calculation device; switch back to the first matrix and update the first matrix with low rank over a third time period; supply a second probabilistic pulse to the second matrix to reset the second matrix back to a symmetric point of the second matrix when the first matrix is being updated with low rank. A system configured to perform the above operations.
9. The system according to claim 8, wherein the first matrix and the second matrix are arranged within the analog crossbar architecture.
10. The system according to claim 8, wherein the first matrix and the second matrix include streaming data. **Claim 11** The system according to claim 10, wherein the streaming data is normalized to prevent an asymmetric effect. **Claim 12** The system according to claim 11, wherein the low-rank update of the first matrix and the second matrix is scaled to adjust a final sketching matrix so as to operate near a symmetric point of the first matrix and a symmetric point of the second matrix and to be between (-0.1, 0.1) over the entire range of (-1, 1). **Claim 13** The system according to claim 12, wherein when the matrix sketching is applied to the entire input, the final sketching matrix is transferred to a digital computer to perform regression analysis. **Claim 14** The system according to claim 8, wherein the dynamic correction calculation device corrects the first matrix and the second matrix simultaneously. **Claim 15** A computer-implemented method, executed on a processor, for performing matrix sketching, comprising: applying dimensionality reduction to streaming data using outer product low-rank updates; when the dimensionality reduction is applied to the entire input, transferring the sketched matrix to a digital computer to perform regression analysis; wherein the sketched matrix is derived from a first matrix and a second matrix that are used in parallel in a toggling manner, the first matrix and the second matrix are disposed within an analog crossbar architecture, and the computer-implemented method further comprises: performing a low-rank update of the first matrix over a first time period; copying the first matrix to a dynamic correction calculation device; switching to the second matrix and performing a low-rank update of the second matrix over a second time period; A method further comprising
16. The computer-implemented method Supplying a first probabilistic pulse to the first matrix to reset the first matrix to a symmetric point of the first matrix when the second matrix is being low-rank updated; Copying the second matrix to the dynamic correction calculation device; Switching back to the first matrix and low-rank updating the first matrix over a third time period; Supplying a second probabilistic pulse to the second matrix to reset the second matrix to a symmetric point of the second matrix when the first matrix is being low-rank updated The method according to claim 15, further comprising
17. A system for performing matrix sketching, A memory, One or more processors communicating with the memory Comprising, the one or more processors Applying dimensionality reduction to streaming data using outer product low-rank updates; Transferring the sketched matrix to a digital computer to perform regression analysis when the dimensionality reduction is applied to the entire input; Configured to perform, the sketched matrix being derived from a first matrix and a second matrix used in parallel in a toggling manner, the first matrix and the second matrix being disposed within an analog crossbar architecture, the one or more processors Low-rank updating the first matrix over a first time period; Copying the first matrix to a dynamic correction calculation device; Switching to the second matrix and low-rank updating the second matrix over a second time period A system further configured to perform.
18. The one or more processors are configured to: supply a first stochastic pulse to the first matrix to reset the first matrix to its symmetric point when the second matrix is updated with a low rank; copy the second matrix to the dynamic correction calculation device; switch back to the first matrix and update the first matrix with a low rank over a third time period; supply a second stochastic pulse to the second matrix to reset the second matrix to its symmetric point when the first matrix is updated with a low rank; The system according to claim 17, further configured to perform the above.
19. A computer program for causing a computer to execute the method according to any one of claims 1 to 7 or 15 to 16.
20. A non-transitory computer-readable storage medium recording the computer program according to claim 19.
Citation Information
Patent Citations
Extensible execution unit interface architecture
US20140229713A1
Systems and methods for low-rank matrix approximation
US20160055124A1
Acceleration of Convolutional Neural Networks on Analog Arrays
US20190354847A1
Instruction cache in a multi-threaded processor
US20200210192A1
Systems and methods for analysis and design of radiating and scattering objects
US7844407B1