Method and system for performing matrix sketching using an analog crossbar architecture
By performing matrix sketching in the analog domain using a simulated cross-switch architecture, and by utilizing dynamic correction computing devices and random pulses to reset the matrix, the computationally intensive problem of deep neural network training is solved, improving computational and energy efficiency and increasing system throughput.
Patent Information
- Application Number
- CN202180035612.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-15
- Filing Date
- 2021-04-13
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-04-13
AI Technical Summary
Existing deep neural network training processes are computationally intensive and time-consuming, making it difficult to perform efficiently using existing data center-scale computing resources composed of graphics processing units (GPUs). Furthermore, the computational and energy efficiency of digital methods needs to be improved.
A simulated cross-switch architecture is adopted. The matrix sketching process is realized by performing matrix sketching in the simulated domain, utilizing dynamic correction computing equipment and random pulse to reset the matrix, and combining low-rank update and outer product update.
This improves the computational and energy efficiency of the matrix sketching process, reduces the computational requirements for training deep neural networks, and increases the system throughput.
Smart Images

Figure CN115605872B_ABST
Abstract
Description
Background Technology
[0001] This invention generally relates to streaming algorithms, and more specifically, to matrix sketching using an analog crossbar architecture.
[0002] Deep learning neural networks (DNNs) have made tremendous progress, surpassing human-level performance in some cases and solving challenging problems such as speech recognition, natural language processing, image classification, and machine translation. However, training large DNNs is a time-consuming and computationally intensive task, requiring data center-scale computing resources comprised of graphics processing units (GPUs) at current technological levels. Attempts have been made to accelerate deep learning workloads beyond GPU levels by leveraging precision-reduced arithmetic to design custom hardware that improves the throughput and energy efficiency of the underlying complementary metal-oxide-semiconductor (CMOS) technology. As an alternative to digital methods, resistive crosspoint device arrays have been proposed to further increase the overall system throughput and energy efficiency by performing vector-matrix multiplication in the analog domain. Summary of the Invention
[0003] According to an embodiment, a method for performing matrix sketching by employing an analog cross-switch architecture is provided. The method includes updating a first matrix in a first time period, copying the first matrix to a dynamic correction computing device, switching to a second matrix to update the second matrix in a second time period, feeding the first matrix with a first random pulse to reset the first matrix back to its symmetry point when the second matrix is updated to a low rank, copying the second matrix to the dynamic correction computing device, switching back to the first matrix to update the first matrix in a third time period, and feeding the second matrix with a second random pulse to reset the second matrix back to its symmetry point when the first matrix is updated to a low rank. The update is a low-rank update.
[0004] According to another embodiment, a system is provided for performing matrix sketching by employing an analog cross-switch architecture. The system includes a memory and one or more processors in communication with the memory, the processors being configured to update a first matrix in a first time period, copy the first matrix to a dynamic correction computing device, switch to a second matrix to update the second matrix in a second time period, feed a first random pulse to the first matrix to reset the first matrix back to its symmetry point when the second matrix is updated to a low rank, copy the second matrix to the dynamic correction computing device, switch back to the first matrix to update the first matrix in a third time period, and feed a second random pulse to the second matrix to reset the second matrix back to its symmetry point when the first matrix is updated to a low rank. The updates are low-rank updates.
[0005] According to another embodiment, a non-transitory computer-readable storage medium is proposed, comprising a computer-readable program for performing matrix sketching by employing an analog cross-switch architecture. The non-transitory computer-readable storage medium performs the following steps: updating a first matrix in a first time period; copying the first matrix to a dynamic correction computing device; switching to a second matrix to update the second matrix in a second time period; when the second matrix is updated to a low rank, feeding a first random pulse to the first matrix to reset the first matrix back to its symmetry point; copying the second matrix to the dynamic correction computing device; switching back to the first matrix to update the first matrix in a third time period; and when the first matrix is updated to a low rank, feeding a second random pulse to the second matrix to reset the second matrix back to its symmetry point. The update is a low-rank update.
[0006] According to an embodiment, a method is provided for performing matrix sketching by employing an analog cross-switch architecture. The method includes applying dimensionality reduction to the streaming data using an outer product update, and once the dimensionality reduction is applied to the entire input, moving the sketched matrix to a digital computer to perform regression analysis. The update is a low-rank update.
[0007] According to another embodiment, a system is provided for performing matrix sketching by employing an analog cross-switch architecture. The system includes memory and one or more processors in communication with the memory, the processors being configured to apply dimensionality reduction to the streaming data using an external product update, and once the dimensionality reduction has been applied to the entire input, to move the sketched matrix to a digital computer to perform regression analysis. The update is a low-rank update.
[0008] It should be noted that exemplary embodiments have been described with reference to different subjects. In particular, some embodiments are described with reference to method-type claims, while others are described with reference to apparatus-type claims. However, those skilled in the art will understand from the above and below description that, unless otherwise indicated, any combination of features related to different subjects, in particular any combination of features between features of method-type claims and features of apparatus-type claims, is also considered to be described herein, except for any combination of features belonging to one type of subject matter.
[0009] These and other features and advantages will become apparent from the following detailed description of illustrative embodiments of the invention, which is read in conjunction with the accompanying drawings. Attached Figure Description
[0010] The present invention will be provided in detail in the following description of preferred embodiments with reference to the accompanying drawings, wherein:
[0011] Figure 1 This is an exemplary analog cross switch architecture employing a sketch matrix according to an embodiment of the present invention;
[0012] Figure 2 It is a combination of the embodiments of the present invention. Figure 1 An exemplary 3D cross switch array of a sketch matrix;
[0013] Figure 3 This is an exemplary system for analog streaming with dynamic calculation according to an embodiment of the present invention;
[0014] Figure 4 An exemplary graph illustrating the switching characteristics of three different devices at a symmetrical point, according to an embodiment of the present invention, is shown.
[0015] Figure 5 This is a block diagram / flowchart of an exemplary method for using two arrays in parallel in an open-loop integration scheme according to an embodiment of the present invention;
[0016] Figure 6 This is a block diagram / flowchart of an exemplary method for switching between first and second matrices according to an embodiment of the present invention;
[0017] Figure 7 This is an exemplary processing system according to an embodiment of the present invention;
[0018] Figure 8 This is a block diagram / flowchart of an exemplary cloud computing environment according to embodiments of the present invention;
[0019] Figure 9 This is a schematic diagram of an exemplary abstract model layer according to an embodiment of the present invention;
[0020] Figure 10 This is a block diagram / flowchart of a method for applying an open-loop integration scheme in an Internet of Things (IoT) system / device / infrastructure according to embodiments of the present invention; and
[0021] Figure 11 This is a block diagram / flowchart of an exemplary IoT sensor for collecting data / information related to an open-loop synthesis streaming algorithm, according to an embodiment of the present invention.
[0022] In all the accompanying drawings, the same or similar reference numerals denote the same or similar elements. Detailed Implementation
[0023] Embodiments of the present invention provide a method and apparatus for performing matrix sketching using an analog crossbar architecture. Due to recent advances in data collection technology, massive amounts of data are being collected at extremely high speeds. Moreover, this data may be unbounded. Streams of unbounded data collected from sensors, devices, and other data sources are referred to as data streams. Various data mining tasks can be performed on data streams when searching for patterns of interest. Data mining tasks typically involve employing data streaming algorithms.
[0024] Streaming algorithms are capable of handling very large, even unbounded, datasets and compute a desired output using only a constant amount of random access memory (RAM). If the dataset is unbounded, it is called a data stream. In this case, if processing of the data stream stops at some position n, the streaming algorithm has a solution corresponding to the data seen up to that point. Therefore, streaming algorithms are algorithms used to process data streams where the input is presented as a sequence of items and can be checked only a few times (usually only once). In most models, these algorithms have access to limited memory. They may also have finite processing time per item. These constraints can mean that the algorithm produces an approximate answer based on a summary or "sketch" of the data stream.
[0025] An exemplary embodiment of the present invention employs a matrix sketching based on a streaming algorithm using an analog cross-switch architecture. The final sketch matrix is placed in the analog cross-switch hardware. The outer product of the final sketch matrix changes with the application of random pulses. Once the sketching process has been applied to the entire input, the final sketch matrix is moved to a digital computer to perform regression analysis.
[0026] It should be understood that the invention will be described based on the given illustrative architecture; however, other structures, substrate materials, and process features and steps / blocks may be varied within the scope of the invention. It should be noted that for practical purposes, certain features may not be shown in all the figures. This should not be construed as limiting the scope of any particular embodiment or illustration or claim.
[0027] Figure 1 This is an exemplary analog cross switch architecture employing a sketch matrix according to an embodiment of the present invention.
[0028] The analog cross switch array 100 includes multiple sketch resistor processing units (RPUs) 110 arranged in a matrix configuration. The outer product is altered by applying random pulses 102, 104. This alteration implicitly computes and updates the rank-one update. At the end of the operation, the final sketch matrix exists in the analog domain. The solution process can be performed in either the analog or digital domain. The pulses reduce the multiplication to a simple coincidence detection 115, which can be implemented by the RPU devices 110.
[0029] In a random update scheme, a random transformer is used to transform the columns and rows (x) i and δ j The encoded digital signals are converted into a random bit stream. These random converters adjust the peripheral pulse probabilities, thus controlling the total number of pulse overlaps occurring at each cross-switch element. In this scheme, these pulses are simultaneously sent to the cross-switch array for all rows and all columns, and then for each overlap event, the corresponding RPU device slightly changes its conductance by Δg. min However, the presence of numerous pulses in the pulsed current causes a change in the total conductance required by the algorithm, Δg. total,ij The change in conductance Δg is realized as a series of pulse coincidences. min As a result, the weight update occurs as a series of coincidence events, each triggering an increase (or decrease) in conductance. ΔW min This is the expected weight change caused by a single coincidence event. Note that the pulse generated on the periphery is applied to all RPU devices in the column (or row), therefore, when calculating the pulse probability to produce the desired weight change at each RPU, the random converter can assume a single Δg across the entire array. min (or equivalently Δw) min )value.
[0030] Figure 2 It is a combination of the embodiments of the present invention. Figure 1 An exemplary 3D cross switch array of a sketch matrix.
[0031] In various example embodiments, the sketch matrix represents memory cells 110 contained between a plurality of bit lines 122 and a plurality of word lines 124. Thus, the array 120 is obtained via vertically conductive word lines (rows) 124 and bit lines (columns) 122, with a sketch matrix present at the intersection between each row and column. The sketch matrix, containing resistive storage elements, can be accessed for reading and writing by biasing the respective word lines 124 and bit lines 122.
[0032] Figure 3 This is an exemplary system for analog streaming with dynamic computation according to an embodiment of the present invention.
[0033] System 130 includes a first matrix 132 and a second matrix 134. The first matrix 132 and the second matrix 134 communicate with the dynamic correction computing device 136.
[0034] The first matrix 132 is updated in the first time period τ1. After updating the first matrix 132, it is copied to the dynamic correction computing device 136. Then, in the second time period τ2, the update function is switched to the second matrix 134. When the second matrix 134 is updated, the first matrix 132 is fed random pulses to reset the first matrix 132 back to its symmetry point. Then, the second matrix 134 is copied to the dynamic correction computing device 136. In the third time period τ3, the update is switched back to the first matrix 132. When the first matrix 132 is updated again, the second matrix 134 is fed random pulses to reset the second matrix 134 back to its symmetry point. This process is repeated until the final sketch matrix is obtained. .
[0035] The second array 134 is utilized to achieve continuous operation at the cost of additional architecture. The streaming data (X, y) should be normalized to avoid further amplifying the asymmetric effect. In addition, the updates should be scaled to target the full range of (-1, 1) to finally adjust Sky to be between (-0.1, 0.1) in order to be closer to the symmetric point (0) operation.
[0036] In alternative embodiments, if the device is fully symmetrical, the need for dynamic correction is eliminated. In these cases, a single array can be used without any carry / zeroing operations. Alternatively, assuming the random sequence ensures operation near symmetry points, zeroing can be skipped. In this case, a single device can be used in addition to the digital accumulation unit, where copy operations are handled in small increments (e.g., column-by-column). All these decisions can be made flexibly and interchangeably depending on the problem at hand and the available equipment.
[0037] Streaming algorithms are open-loop ensemble algorithms. Therefore, dynamic systems approaches with secondary arrays cannot guarantee the resetting of the first matrix. However, unlike neural networks, these algorithms do not rely on vector-matrix multiplication, so the accumulated matrix can reside in the digital domain. Therefore, instead, two arrays or matrices 132 and 134 can be used in parallel in a flipped manner.
[0038] In alternative embodiments, more than two matrices may be used. For example, three matrices may be used. In another example embodiment, four matrices may be used. Of course, those skilled in the art can envision multiple matrices used in conjunction with the dynamic correction computing device 136. In one instance, the first matrix may be updated within a first time period τ1. After updating the first matrix, it is copied to the dynamic correction computing device 136. Then, during a second time period τ2, the update function is switched to the second matrix. When the second matrix is updated, the first and third matrices are fed random pulses to reset the first and third matrices back to their respective symmetry points S. The second matrix is then copied to the dynamic correction computing device 136. The update is switched to the third matrix, which may be updated within a third time period τ3. After updating the third matrix, it is copied to the dynamic correction computing device 136. Then, during a fourth time period τ4, the update function is switched back to the first matrix. When the first matrix is updated, the second and third matrices are fed random pulses to reset the second and third matrices back to their respective symmetry points, thus one matrix can be updated while the other two matrices can be fed random pulses to reset to their respective symmetry points. Similarly, one matrix can be updated, and three other matrices can be fed random pulses to reset them to their respective symmetric points. Thus, multiple matrices can be connected to one or more dynamic correction computing devices 136, where one matrix is updated while all other matrices are fed random pulses to reset them to their respective symmetric points. Therefore, the exemplary embodiment is not limited to the number of matrices connected to the dynamic correction computing device 136.
[0039] Figure 4 An exemplary graph illustrating the switching characteristics of three different devices at symmetrical points is shown in the description of an embodiment of the present invention.
[0040] Regarding asymmetric updates, the incremental change obtained by the simulated device during an update depends on the weight values, which can be represented by a soft-constraint model. This effect introduces bias and severely impairs training performance. (The generation of these devices). For these devices, there exist points called symmetric points. A random sequence of increment and decrement operations with equal probability will inevitably drive the device state to this symmetric point. This behavior is observed in streaming algorithms because they involve a large number of incremental changes. This requires these analog resistive devices to change their conductance symmetrically when excited by positive or negative voltage pulses.
[0041] Figure 4Three distinct device switching characteristics are described. In ideal device 160, the conductance increment and decrement are equal in magnitude and independent of the device conductance. In symmetric device 170, the conductance increment and decrement are equal in strength, but both depend on the device conductance. In asymmetric device 180, the conductance increment and decrement are not equal in strength, and both have different dependencies on the device conductance. However, there exists a single point where the conductance increment and decrement are equal in strength. This point is called the symmetric point, and for the example shown, this point matches the reference device conductance, thus occurring at 𝑤=0. Symmetric point 162 is shown in ideal device 160, symmetric point 172 is shown in symmetric device 170, and symmetric point 182 is shown in asymmetric device 180.
[0042] Note that even for the asymmetric device shown in 180, there exists a single point (conductance value) where the magnitude of the conductance increment and conductance decrease are equal. This point is called the symmetric point of the updated device, and due to device-to-device variations, it can correspond to any weight value (not necessarily zero as shown in 180).
[0043] Figure 5 This is a block diagram / flowchart of an exemplary method for using two arrays in parallel in an open-loop integration scheme according to an embodiment of the present invention.
[0044] In box 202, the first matrix is updated in low rank during the first time period.
[0045] In box 204, the first matrix is copied to the dynamic correction computing device.
[0046] In box 206, switch to the second matrix and update the second matrix in low rank during the second time period.
[0047] In box 208, while updating the second matrix in low rank (or simultaneously, concurrently), pulses are fed into the first matrix to reset it back to the symmetric point.
[0048] At box 210, copy the second matrix into the dynamic correction computing device.
[0049] In box 212, switch to the first matrix and update the first matrix in low rank during the third time period.
[0050] In box 214, while updating the first matrix in low rank (or simultaneously, concurrently), pulses are fed to the second matrix to reset it back to the symmetric point.
[0051] Figure 6 This is a block diagram / flowchart of an exemplary method for switching between first and second matrices according to an embodiment of the present invention.
[0052] In box 220, the method waits.
[0053] In box 222, determine if a new sample has been received. If not, the process proceeds to box 230, where the matrix is read and the device is switched. If yes, the process proceeds to box 224.
[0054] In box 224, generate vector "s".
[0055] In box 226, a simulated low-rank update is performed, and the process proceeds to box 220.
[0056] In box 232, the matrix is stored in digital form or in a separate analog device.
[0057] Figure 7 This is an exemplary processing system for processing streaming algorithms according to embodiments of the present invention.
[0058] Now for reference Figure 7 The figure illustrates a hardware configuration of a computing system 600 according to an embodiment of the present invention. As shown, the hardware configuration has at least one processor or central processing unit (CPU) 611. The CPU 611 is interconnected via a system bus 612 to random access memory (RAM) 614, read-only memory (ROM) 616, input / output (I / O) adapter 618 (for connecting peripheral devices such as disk unit 621 and peripheral devices with driver 640 to bus 612), user interface adapter 622 (for connecting keyboard 624, mouse 626, speaker 628, microphone 632 and / or other user interface devices to bus 612), communication adapter 634 for connecting the system 600 to a data processing network, Internet, intranet, local area network (LAN), etc., and display adapter 636 for connecting bus 612 to display device 638 and / or printer 639 (e.g., digital printer, etc.).
[0059] Figure 8 This is a block diagram / flowchart of an exemplary cloud computing environment according to an embodiment of the present invention.
[0060] Figure 8 This is a block diagram / flowchart of an exemplary cloud computing environment according to an embodiment of the present invention.
[0061] It should be understood that although this invention includes a detailed description of cloud computing, the implementation of the teachings described herein is not limited to a cloud computing environment. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed in the future.
[0062] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0063] The characteristics are as follows:
[0064] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without requiring manual interaction with the service provider.
[0065] Wide Area Network (WAN) Access: Capabilities are available on the network and accessed through standard mechanisms that facilitate the use of heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0066] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. Location independence has significance because consumers typically do not control or know the exact location of the resources provided, but can specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0067] Rapid Flexibility: In some cases, the ability to scale outwards and inwards quickly and flexibly can be provided. For consumers, the available capacity often appears unlimited and can be purchased in any quantity at any time.
[0068] Measurement services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both the providers and consumers of the services being utilized.
[0069] The service model is as follows:
[0070] Software as a Service (SaaS): The capability offered to consumers is the ability to use the provider's applications running on cloud infrastructure. Applications can be accessed from various client devices through thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.
[0071] Platform as a Service (PaaS): This provides consumers with the ability to deploy consumer-created or acquired applications onto cloud infrastructure using programming languages and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and the configuration of any application hosting environments.
[0072] Infrastructure as a Service (IaaS): This provides consumers with the capability to deliver processing, storage, networking, and other basic computing resources that enable them to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).
[0073] The deployment model is as follows:
[0074] Private cloud: Cloud infrastructure operated solely by an organization. It can be managed by the organization or a third party and can exist on-site or off-site.
[0075] Community cloud: Cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.
[0076] Public cloud: Cloud infrastructure available to the general public or large industrial groups and owned by organizations that sell cloud services.
[0077] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported together (e.g., cloud bursting for load balancing between clouds).
[0078] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure of a network of interconnected nodes.
[0079] Now for reference Figure 8This paper depicts an illustrative cloud computing environment 750 for enabling use cases of the present invention. As shown, the cloud computing environment 750 includes one or more cloud computing nodes 710 to which local computing devices used by cloud consumers can communicate, such as personal digital assistants (PDAs) or cellular phones 754A, desktop computers 754B, laptop computers 754C, and / or automotive computer systems 754N. The nodes 710 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds or combinations thereof described above. This allows the cloud computing environment 750 to provide infrastructure, platform, and / or software as a service, without requiring cloud consumers to maintain resources on their local computing devices. It should be understood that... Figure 8 The types of computing devices 754A-N shown are illustrative only, and computing node 710 and cloud computing environment 750 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).
[0080] Figure 9 This is a schematic diagram of an exemplary abstract model layer according to an embodiment of the present invention. It should be understood beforehand that... Figure 9 The components, layers, and functions shown are for illustrative purposes only, and embodiments of the invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:
[0081] The hardware and software layer 860 includes hardware and software components. Examples of hardware components include: a host 861; a server 862 based on a RISC (Reduced Instruction Set Computer) architecture; a server 863; a blade server 864; a storage device 865; and a network and network components 866. In some embodiments, the software components include network application server software 867 and database software 868.
[0082] The virtualization layer 870 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 871; virtual storage 872; virtual network 873, including virtual private network; virtual application and operating system 874; and virtual client 875.
[0083] In one example, management layer 880 may provide the functionality described below. Resource provisioning 881 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 882 provides cost tracking when utilizing resources within the cloud computing environment, as well as accounting or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, and protection for data and other resources. User portal 883 provides access to the cloud computing environment for consumers and system administrators. Service level management 884 provides cloud resource allocation and management to ensure the required service level is met. Service level agreement (SLA) planning and fulfillment 885 provides pre-scheduling and procurement of cloud resources, where future needs are anticipated according to the SLA.
[0084] The workload layer 890 provides examples of functionalities that can be leveraged in a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 891; software development and lifecycle management 892; virtual classroom education delivery 893; data analytics and processing 894; transaction processing 895; and streaming algorithms 896.
[0085] Figure 10 This is a block diagram / flowchart of a method for applying an open-loop integration scheme in an Internet of Things (IoT) system / device / infrastructure according to an embodiment of the present invention.
[0086] According to some embodiments of the present invention, a network is implemented using IoT methods. For example, streaming algorithm 902 can be incorporated into, for example, wearable, implantable, or ingestible electronic devices and Internet of Things (IoT) sensors. Wearable, implantable, or ingestible devices may include at least health and wellness monitoring devices and fitness devices. Wearable, implantable, or ingestible devices may also include at least implantable devices, smartwatches, head-mounted devices, security and protection devices, and gaming and lifestyle devices. IoT sensors can be incorporated into at least home automation applications, automotive applications, user interface applications, lifestyle and / or entertainment applications, urban and / or infrastructure applications, toys, health, fitness, retail tags and / or trackers, platforms and components, etc. The streaming algorithm 902 described herein can be incorporated into any type of electronic device for any type of use, application, or operation.
[0087] IoT systems allow users to achieve deeper automation, analytics, and integration within systems. IoT improves the scope and accuracy of these areas. IoT leverages existing and emerging technologies for sensing, networking, and robotics. IoT features include artificial intelligence, connectivity, sensors, proactive engagement, and small device usage. In various embodiments, the streaming algorithm 902 of this invention can be incorporated into a variety of different devices and / or systems. For example, the streaming algorithm 902 can be incorporated into a wearable or portable electronic device 904. The wearable / portable electronic device 904 may include implantable devices 940, such as smart clothing 943. The wearable / portable device 904 may include smartwatches 942 and smart jewelry 945. The wearable / portable device 904 may also include fitness monitoring devices 944, health and wellness monitoring devices 946, head-mounted devices 948 (e.g., smart glasses 949), security and protection systems 950, gaming and lifestyle devices 952, smartphones / tablets 954, media players 956, and / or computers / computing devices 958.
[0088] The streaming algorithm 902 of this invention can be further incorporated into Internet of Things (IoT) sensors 906 for various applications such as home automation 920, automobiles 922, user interfaces 924, living and / or entertainment 926, urban and / or infrastructure 928, retail 910, tags and / or trackers 912, platforms and components 914, toys 930 and / or healthcare 932, and fitness 934. The IoT sensor 906 may employ the streaming algorithm 902. Of course, those skilled in the art will envision incorporating this streaming algorithm 902 into any type of electronic device for any type of application, and not limited to those described herein.
[0089] Figure 11 This is a block diagram / flowchart of an exemplary IoT sensor for collecting data / information related to an open-loop synthesis streaming algorithm, according to an embodiment of the present invention.
[0090] Without sensors, the Internet of Things (IoT) loses its distinctiveness. IoT sensors act as instruments that define the transformation of IoT from a standard passive network of devices into an active system capable of real-world integration.
[0091] The IoT sensor 906 can employ streaming algorithm 902 to continuously and in real-time transmit information / data to any type of distributed system via network 908. Exemplary IoT sensors 906 may include, but are not limited to, location / presence / proximity sensors 1002, motion / velocity sensors 1004, displacement sensors 1006 such as acceleration / tilt sensors 1007, temperature sensors 1008, humidity / moisture sensors 1010, flow sensors 1011, acoustic / sound / vibration sensors 1012, chemical / gas sensors 1014, force / load / torque / deformation / pressure sensors 1016, and / or electro / magnetic sensors 1018. Those skilled in the art can envision using any combination of these sensors to collect data / information via streaming algorithm 902 of the distributed system for further processing. Those skilled in the art can envision using other types of IoT sensors, such as, but not limited to, magnetometers, gyroscopes, image sensors, light sensors, radio frequency identification (RFID) sensors, and / or microfluidic sensors. IoT sensors may also include energy modules, power management modules, RF modules, and sensing modules. RF modules manage communication through their signal processing, WiFi, Bluetooth, radio transceivers, duplexers, etc.
[0092] As used herein, the terms “data,” “content,” “information,” and similar terms are used interchangeably to refer to data that can be captured, transmitted, received, displayed, and / or stored according to various example embodiments. Therefore, the use of any such terms should not be considered as limiting the spirit and scope of this disclosure. Furthermore, where a computing device is described herein as receiving data from another computing device, the data may be received directly from that other computing device or indirectly via one or more intermediate computing devices, such as one or more servers, repeaters, routers, network access points, base stations, etc.
[0093] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) that the user can use to provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input.
[0094] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.
[0095] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or recessed structures with instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0096] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the respective computing / processing device.
[0097] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of this invention, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions to personalize the electronic circuits by utilizing state information from the computer-readable program instructions.
[0098] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0099] These computer-readable program instructions may be provided to at least one processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks or modules of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other apparatus to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks or modules of a flowchart and / or block diagram.
[0100] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operation blocks / steps to be executed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus or other device implement the functions / actions specified in one or more blocks or modules of a flowchart and / or block diagram.
[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the blocks may occur in a non-consecutive order as shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0102] References to the principles of this specification as "an embodiment" or "an embodiment" and other variations mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment of the principles. Therefore, the phrases "in an embodiment" or "in an embodiment" appearing in various places throughout the specification, as well as any other variations, do not necessarily refer to the same embodiment.
[0103] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one of” is intended to cover the selection of only the first listed option (A), or only the selection of only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). This can be extended to many of the listed items, as will be apparent to those skilled in the art and related fields.
[0104] Preferred embodiments of systems and methods for streaming algorithms using analog crossbar switch structures have been described (these are intended to be illustrative and not limiting). It should be noted that modifications and variations can be made by those skilled in the art based on the foregoing teachings. Therefore, it should be understood that changes can be made to the specific embodiments described, and these changes are within the scope of the invention as summarized by the appended claims. Thus, aspects of the invention have been described in the details and features required by patent law, and the claimed and patent-protected aspects are set forth in the appended claims.
Claims
1. A computer-implemented method, executed on a processor, for performing matrix sketching by employing an analog cross-switch architecture, the method comprising: The first matrix and the second matrix are placed in an analog cross switch architecture, wherein the analog cross switch architecture includes multiple resistive devices arranged in a matrix configuration at the intersections between rows and columns; The first matrix is updated in low rank during the first time period; Copy the first matrix into the dynamic correction computing device; Switch to the second matrix to update the second matrix in a low-rank manner during the second time period; When the second matrix is updated to low rank, a first random pulse is fed to the first matrix to reset the first matrix back to the symmetric point of the first matrix, where the symmetric point is a single point where the magnitude of the conductance increment and conductance decrement of the resistive device are equal. Copy the second matrix into the dynamic correction computing device; Switch back to the first matrix to update the first matrix in low rank during the third time period; and When the first matrix is updated to low rank, a second random pulse is fed to the second matrix to reset the second matrix back to the symmetric point of the second matrix.
2. The method according to claim 1, wherein the first matrix and the second matrix comprise streaming data.
3. The method according to claim 2, wherein, The streaming data is normalized to prevent asymmetric effects.
4. The method according to claim 3, wherein, The low-rank updates of the first and second matrices are scaled to adjust the final sketch matrix to between (-0.1, 0.1) for the full range of (-1, 1), in order to operate near the symmetric points of the first and second matrices.
5. The method according to claim 4, wherein, Once the matrix sketch is applied to the entire input, the final sketch matrix is moved to a digital computer to perform regression analysis.
6. The method of claim 1, wherein the dynamic correction computing device simultaneously corrects the first matrix and the second matrix.
7. A computer program product comprising a computer-readable program, executable on a processor of a data processing system, for performing matrix sketching by employing an analog crossbar architecture, wherein, When executed on the processor, the computer-readable program causes the computer to perform the following steps: The first matrix and the second matrix are placed in an analog cross switch architecture, wherein the analog cross switch architecture includes multiple resistive devices arranged in a matrix configuration at the intersections between rows and columns; The first matrix is updated in low rank during the first time period; Copy the first matrix into the dynamic correction computing device; Switch to the second matrix to update the second matrix in a low-rank manner during the second time period; When the second matrix is updated to low rank, a first random pulse is fed to the first matrix to reset the first matrix back to the symmetric point of the first matrix, where the symmetric point is a single point where the magnitude of the conductance increment and conductance decrement of the resistive device are equal. Copy the second matrix into the dynamic correction computing device; Switch back to the first matrix to update the first matrix in low rank during the third time period; and When the first matrix is updated to low rank, a second random pulse is fed to the second matrix to reset the second matrix back to the symmetric point of the second matrix.
8. The computer program product of claim 7, wherein the first matrix and the second matrix comprise streaming data.
9. The computer program product according to claim 8, wherein, The streaming data is normalized to prevent asymmetric effects.
10. The computer program product according to claim 9, wherein, The low-rank updates of the first and second matrices are scaled to adjust the final sketch matrix to between (-0.1, 0.1) for the full range of (-1, 1), in order to operate near the symmetric points of the first and second matrices.
11. The computer program product according to claim 10, wherein, Once the matrix sketch is applied to the entire input, the final sketch matrix is moved to a digital computer to perform regression analysis.
12. The computer program product according to claim 7, wherein, The dynamic correction calculation device simultaneously corrects the first matrix and the second matrix.
13. A system for performing matrix sketching by employing an analog cross-switch architecture, the system comprising: Memory; as well as One or more processors communicating with the memory, the one or more processors being configured to: The first matrix and the second matrix are placed in an analog cross switch architecture, wherein the analog cross switch architecture includes multiple resistive devices arranged in a matrix configuration at the intersections between rows and columns; The first matrix is updated in low rank during the first time period; Copy the first matrix into the dynamic correction computing device; Switch to the second matrix to update the second matrix in a low-rank manner during the second time period; When the second matrix is updated to low rank, a first random pulse is fed to the first matrix to reset the first matrix back to the symmetric point of the first matrix, where the symmetric point is a single point where the magnitude of the conductance increment and conductance decrement of the resistive device are equal. Copy the second matrix into the dynamic correction computing device; Switch back to the first matrix to update the first matrix in low rank during the third time period; and When the first matrix is updated to low rank, a second random pulse is fed to the second matrix to reset the second matrix back to the symmetric point of the second matrix.
14. The system according to claim 13, wherein, The first matrix and the second matrix include streaming data.
15. The system according to claim 14, wherein, The streaming data is normalized to prevent asymmetric effects.
16. The system according to claim 15, wherein, The low-rank updates of the first and second matrices are scaled to adjust the final sketch matrix to between (-0.1, 0.1) for the full range of (-1, 1), in order to operate near the symmetric points of the first and second matrices.
17. The system according to claim 16, wherein, Once the matrix sketch is applied to the entire input, the final sketch matrix is moved to a digital computer to perform regression analysis.
18. The system of claim 13, wherein the dynamic correction computing device simultaneously corrects the first matrix and the second matrix.
19. A computer-implemented method, executed on a processor, for performing matrix sketching by employing an analog cross-switch architecture, the method comprising: Dimension reduction is applied to streaming data using outer product low-rank update; Once dimensionality reduction is applied to the entire input, the sketch matrix is moved to a digital computer to perform regression analysis, wherein the sketch matrix is derived from a first matrix and a second matrix used in parallel in a flipped manner, wherein the first matrix and the second matrix are placed in an analog cross-switch architecture, wherein the analog cross-switch architecture includes multiple resistive devices arranged in a matrix configuration at the intersections between rows and columns; The first matrix is updated in low rank during the first time period; Copy the first matrix into the dynamic correction computing device; Switch to the second matrix to update the second matrix in a low-rank manner during the second time period; When the second matrix is updated to low rank, a first random pulse is fed to the first matrix to reset the first matrix back to the symmetric point of the first matrix, where the symmetric point is a single point where the magnitude of the conductance increment and conductance decrement of the resistive device are equal. Copy the second matrix into the dynamic correction computing device; Switch back to the first matrix to update the first matrix in low rank during the third time period; and When the first matrix is updated to low rank, a second random pulse is fed to the second matrix to reset the second matrix back to the symmetric point of the second matrix.
20. A system for performing matrix sketching by employing an analog cross-switch architecture, the system comprising: Memory; as well as One or more processors communicating with the memory, the one or more processors being configured to: Dimension reduction is applied to streaming data using outer product low-rank update; Once dimensionality reduction is applied to the entire input, the sketch matrix is moved to a digital computer to perform regression analysis, wherein the sketch matrix is derived from a first matrix and a second matrix used in parallel in a flipped manner, wherein the first matrix and the second matrix are placed in an analog cross-switch architecture, wherein the analog cross-switch architecture includes multiple resistive devices arranged in a matrix configuration at the intersections between rows and columns; The first matrix is updated in low rank during the first time period; Copy the first matrix into the dynamic correction computing device; Switch to the second matrix to update the second matrix in a low-rank manner during the second time period; When the second matrix is updated to low rank, a first random pulse is fed to the first matrix to reset the first matrix back to the symmetric point of the first matrix, where the symmetric point is a single point where the magnitude of the conductance increment and conductance decrement of the resistive device are equal. Copy the second matrix into the dynamic correction computing device; Switch back to the first matrix to update the first matrix in low rank during the third time period; and When the first matrix is updated to low rank, a second random pulse is fed to the second matrix to reset the second matrix back to the symmetric point of the second matrix.
Citation Information
Patent Citations
Data processing device and computing equipment for convolution calculation
CN108665061A
Dimensionality Reduction of Computer Programs
US20190138721A1