Hardware acceleration for calculating eigenvalues ​​of matrices

The neuromorphic computing system with RPUs accelerates eigenpair calculations by performing analog matrix operations, addressing inefficiencies in existing methods and enhancing computational speed for large matrices.

JP7849129B2Active Publication Date: 2026-04-21INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2022-04-12
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently calculating eigenvalues and eigenvectors of large matrices, particularly due to the high computational cost and complexity of iterative numerical methods, which are not effectively addressed by current hardware solutions.

Method used

A neuromorphic computing system utilizing resistive processing units (RPUs) with tunable resistive memory devices performs hardware-accelerated matrix eigenpair calculations through analog matrix-vector multiplication and iterative processes like power iteration and matrix deflation, leveraging RPU arrays for in-memory computations.

Benefits of technology

This approach significantly speeds up the calculation of eigenpairs by reducing computational time and resource requirements, making it more efficient for large matrices, especially symmetric and SPD matrices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007849129000027
    Figure 0007849129000027
  • Figure 0007849129000028
    Figure 0007849129000028
  • Figure 0007849129000029
    Figure 0007849129000029
Patent Text Reader

Abstract

Techniques are provided for implementing hardware accelerated computing of eigenpairs of a matrix. For example, a system includes a processor and a resistive processing unit coupled to the processor. The resistive processing unit includes an array of cells, each of the cells including a resistive device, at least some of the resistive devices being adjustable to encode values ​​of a predefined matrix stored in the array of cells. When the predefined matrix is ​​stored in the array of cells, the processor is configured to determine eigenvectors of the stored matrix by performing a process including performing an analog matrix-vector multiplication operation on the stored matrix to converge an initial vector to an estimate of the eigenvector of the stored matrix.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention generally relates to an analog resistance processing system for neuromorphic computing, and to a technique for performing hardware-accelerated numerical computing tasks using the analog resistance processing system. [Background technology]

[0002] Information processing systems such as neuromorphic computing systems and artificial neural network systems are used in a variety of applications, including cognitive recognition, machine learning, and inference processing for computing. Such systems are generally hardware-based systems that include a large number of highly interconnected processing elements (called "artificial neurons") that operate in parallel to perform various types of computations. Artificial neurons (e.g., presynaptic and postsynaptic neurons) are connected using artificial synaptic devices that provide synaptic weights representing the strength of connections between the artificial neurons. Synaptic weights can be implemented using an array of resistive processing unit (RPU) cells with tunable resistive memory devices (e.g., tunable conductance), where the conductance state of the RPU cells is encoded or mapped to the synaptic weights. [Overview of the Initiative]

[0003] Aspects of the present invention include systems, computer program products, and methods for implementing hardware-accelerated computing of matrix eigenvectors, as defined in the appended claims. In an exemplary embodiment, the system comprises a processor and a resistive processing unit coupled to the processor. The resistive processing unit includes an array of cells, each cell including a resistive device, at least a portion of which are tunable to encode values ​​of a given matrix stored in the array of cells. Given that a given matrix is ​​stored in the array of cells, the processor is configured to determine the eigenvectors of the stored matrix by performing a process that includes performing analog matrix-vector multiplication on the stored matrix to converge an initial vector to an estimate of the eigenvectors of the stored matrix.

[0004] Other embodiments will be described in the following detailed description of exemplary embodiments, which should be read in conjunction with the accompanying figures. [Brief explanation of the drawing]

[0005] [Figure 1] An exemplary embodiment of this disclosure schematically illustrates a computing system that implements hardware-accelerated computing of matrix eigenpairs. [Figure 2] An exemplary embodiment of the present disclosure schematically illustrates an RPU computing system that may be implemented in the system shown in Figure 1 to provide hardware-accelerated computing of matrix eigenpairs. [Figure 3] This figure shows a method for performing specific-to-computation processing according to an exemplary embodiment of the present disclosure. [Figure 4A] This figure shows a method for performing specific-to-computation processing according to an exemplary embodiment of the present disclosure. [Figure 4B] This figure shows a method for performing specific-to-computation processing according to an exemplary embodiment of the present disclosure. [Figure 5A]An exemplary embodiment of this disclosure schematically illustrates a method for performing analog matrix-vector multiplication operations using an RPU computing system that includes an array of RPU cells for implementing hardware-accelerated computing of matrix eigenpairs. [Figure 5B] An exemplary embodiment of this disclosure schematically illustrates a method for performing an analog cross product operation using an RPU computing system that includes an array of RPU cells for implementing hardware-accelerated computing of eigenpairs of matrices. [Figure 6A] An exemplary embodiment of this disclosure schematically illustrates a method for configuring an RPU computing system that includes an array of RPU cells for performing matrix-vector operations to implement hardware-accelerated computing of eigenpairs of matrices. [Figure 6B] This figure schematically illustrates a method for configuring an RPU computing system, which includes an array of RPU cells that perform analog cross product operations to implement hardware-accelerated computing of matrix eigenpairs, according to an exemplary embodiment of the present disclosure. [Figure 7] This figure schematically illustrates a method for configuring an RPU computing system, which includes multiple arrays of RPU cells, to perform matrix-vector operations for unique pair computing processing using signed matrix values, according to an exemplary embodiment of the present disclosure. [Figure 8] This figure schematically illustrates an exemplary architecture of a computing node capable of hosting a system configured to perform specific computing operations, according to exemplary embodiments of the present disclosure. [Figure 9] This figure shows a cloud computing environment according to an exemplary embodiment of the present disclosure. [Figure 10] This figure shows an abstraction model layer according to an exemplary embodiment of the present disclosure. [Modes for carrying out the invention]

[0006] Embodiments of the present invention will now be described in further detail with respect to systems and methods for providing hardware-accelerated computing of matrix eigenpairs. It should be understood that the various features shown in the accompanying drawings are schematic diagrams and not drawn to scale. Furthermore, identical or similar reference numerals are used throughout the drawings to indicate identical or similar features, elements, or structures, and therefore, detailed descriptions of identical or similar features, elements, or structures are not repeated for each of the drawings. Furthermore, the term “exemplary” is used herein to mean “serving as an example, illustration, or representation.” Any embodiment or design described herein as “exemplary” should not necessarily be construed as being preferable or advantageous over other embodiments or designs.

[0007] Furthermore, the expression "configured to perform one or more functions or otherwise provide some functionality," used in conjunction with circuits, structures, elements, components, etc., is intended to encompass embodiments in which circuits, structures, elements, components, etc. are implemented in hardware, software, and / or combinations thereof, and in hardware-based implementations, the hardware may comprise discrete circuit elements (e.g., transistors, inverters, etc.), programmable elements (ASICs, FPGAs, etc.), processing devices (CPUs, GPUs, etc.), one or more integrated circuits, and / or combinations thereof. Therefore, although this is merely illustrative, when a circuit, structure, element, component, etc. is defined as being configured to provide a specific function, this is intended to cover embodiments comprising elements, processing devices, or integrated circuits, or combinations thereof, that enable the circuit, structure, element, component, etc. to perform a specific function when it is in an operational state (e.g., connected to or otherwise deployed in a system, powered on, receiving input, or generating output, or a combination thereof), as well as covering embodiments for when the circuit, structure, element, component, etc. is in a non-operational state (e.g., not connected, not deployed in a system, not powered on, not receiving input, or not generating output, or a combination thereof) or partially operational.

[0008] Figure 1 schematically illustrates a computing system that implements hardware-accelerated computing of matrix eigenpairs according to an exemplary embodiment of the present disclosure. In particular, Figure 1 schematically illustrates a computing system 100 comprising an application 110, a digital processing system 120, and a neuromorphic computing system 130. The digital processing system 120 comprises a plurality of processor cores 122. The neuromorphic computing system 130 comprises a plurality of neural cores 132. In some embodiments, the neuromorphic computing system 130 comprises a resistive processing unit (RPU) system in which each neural core 132 comprises one or more analog resistive processing unit arrays (e.g., analog RPU crossbar array hardware). The neural cores 132 are configured to support hardware acceleration for matrix decomposition operations, such as the calculation of matrix eigenpairs, by performing multiply-accumulate (MAC) operations in the analog domain in order to support hardware acceleration of numerical operations such as vector-matrix multiplication, matrix-vector multiplication, vector-vector multiplication, or matrix multiplication operations performed on the RPU array 134.

[0009] In some embodiments, the digital processing system 120 controls the execution of matrix decomposition processes, such as an eigenpair calculation process 140, which is performed to compute eigenpairs of a given matrix A provided by the application 110. The eigenpair calculation process 140 implements various processing and optimization solver techniques to enable the computation of eigenpairs of the matrix. For example, in some embodiments, the eigenpair calculation process 140 implements an eigenpair decision control process 142 and a matrix inversion process 144, which are utilized by the eigenpair calculation process 140 to compute one or more eigenpairs of a given matrix A. In some embodiments, the eigenpair decision control process 142 and the matrix inversion process 144 are software modules executed by the processor core 122 of the digital processing system 120 to perform the eigenpair calculation process 140. As will be described in more detail below, the eigenpair calculation process 140 utilizes the neuromorphic computing system 130 to compute a given matrix (e.g., the original matrix A or an approximate (estimated) inverse matrix A) stored in one or more RPU arrays 134, as schematically shown in Figure 1. -1 For this purpose, hardware-accelerated multiply-accumulate (MAC) operations (e.g., cross product) are performed in the analog domain through various in-memory calculations such as matrix-vector operations and vector-vector product operations.

[0010] Application 110 may include any type of computational application that performs numerical operations, solves linear equations, and utilizes matrices and inverse matrices as computational targets for other calculations (e.g., scientific computing applications, engineering applications, graphics rendering applications, signal processing applications, facial recognition applications, matrix diagonalization applications, MIMO (multiple input, multiple output) systems for wireless communication, encryption, etc.). Mathematically, a linear system (or system of linear equations) is a collection of one or more linear equations with the same set of variables, and the solution to a linear system involves assigning values ​​to the variables such that all linear equations (of the linear system) are satisfied simultaneously. Mathematically, the theory of linear systems is the foundation and basic part of numerical linear algebra. Numerical linear algebra implements computer algorithms that solve linear systems in an efficient and accurate way by utilizing the properties of vectors and matrices and related vector / matrix operations. Common problems in numerical linear algebra include obtaining matrix decompositions such as singular value decomposition, eigenvector decomposition, and matrix diagonalization, which can be used to determine the eigenvalues ​​and eigenvectors of a given matrix, thereby solving common problems in linear algebra, such as systems of linear equations.

[0011] For example, a system of linear equations with constant coefficients can be solved using an algebraic method that represents a linear transformation using matrix A and determines the eigenvalues ​​and eigenvectors of the matrix using iterative processing. The eigenvalues ​​and eigenvectors of matrix A represent the solutions to the linear equations. More specifically, a system of linear equations can be expressed in the form of the following eigenvalue equations. Ax = λx (Equation 1) Here, A is an n×n matrix of real or complex numbers (or more precisely, a diagonalizable matrix), where λ is a scalar number (real or complex) and is an eigenvalue of matrix A, and x is a non-zero vector of dimension n (possibly complex) and is an eigenvector of matrix A. More specifically, in equation (1), the n×n matrix A can represent a linear transformation, and the eigenvectors are n×1 matrices. The eigenvalue equation (equation 1) means that under the action of the linear transformation A, the vector x is transformed into the collinear vector λx. A vector having this property is considered an eigenvector of the linear transformation A, and the associated scalar λ is considered an eigenvalue. As used herein, the term “eigenpair” refers to a mathematical pair of an eigenvector and its associated eigenvalue. In many applications, the eigenpair (λ) of a given matrix A is used. i ,x i It is desirable to find ), where λ i x is an eigenvalue, x i These are the corresponding eigenvectors.

[0012] In essence, the eigenvectors x of a linear transformation of matrix A are non-zero vectors whose direction does not change when A is applied to them. Applying a linear transformation A to eigenvectors x scales them only by their eigenvalues ​​λ. Eigenvectors indicate the direction of the action of a linear transformation, whether it simply stretches, shrinks, or reverses direction, while eigenvalues ​​indicate the magnitude of the change in that direction. In other words, eigenvectors are vectors that a linear transformation A simply stretches, shrinks, or reverses direction, and the amount by which an eigenvector stretches / shrinks / reverses direction is based on its eigenvalue. In this regard, the eigenvalue λ can be any scalar value (or zero or a complex number), and the eigenvalue λ can be negative, in which case the eigenvector reverses direction as part of the scaling.

[0013] A square n×n matrix A is given by A = XΛX -1 A matrix is ​​diagonalizable if there exist inverted matrix X and diagonal matrix Λ such that (on the other hand, symmetric matrices are diagonalizable by definition). In some embodiments, matrix A is diagonalized, i.e., A = XΛX-1 The process is performed by determining matrices X and Λ that satisfy the condition. A given matrix A is diagonalizable if matrix A has n linearly independent eigenvectors. In other words, if a real n×n matrix A has n eigenvalue pairs (λ1,x1),(λ2,x2),...,(λ n ,x n ), and the eigenvectors x1,x2,...,x n are linearly independent, then A = XΛX -1 , where X is an invertible n×n matrix whose columns are the vectors x1,x2,...,x n that is, X=(x1,x2,...,x n ), and Λ is an n×n matrix, and the following holds. TIFF0007849129000001.tif31156

[0014] In this example, the column vectors of X form a basis of the eigenvectors of matrix A, the i-th column of X is the eigenvector x i of A, Λ is a diagonal matrix, and the i-th diagonal element λ i ​​​​​​​​​​​​​​​​​​​​​​Furthermore, if the columns of X are a symmetric n×n matrix of n linearly independent eigenvectors of the symmetric matrix A, then X is an orthogonal matrix whose transpose X is equal to the inverse of X, i.e., X T =X -1 Therefore, the eigenvalue equation (Equation 1) can be written as AX = XΛ, where Λ is a diagonal n × n matrix, and its elements correspond sequentially to the eigenvalues ​​of A and the columns of X. T =X T-1 Based on this, the equation AX=XΛ becomes A=XΛX T It may be rewritten as, and furthermore, X T AX can be rewritten as Λ.

[0016] In some embodiments, the eigenvalue pair calculation process 140 implements an eigendecomposition process for determining the eigenvalues ​​and eigenvectors of a given diagonalizable matrix, and the eigendecomposition is based, for example, on an iterative numerical method. In general, for a small n × n matrix A, the eigenvalues ​​of matrix A can be determined symbolically using the characteristic polynomial. Specifically, from the eigenvalue equation Ax = λx (Equation 1), multiplying the right-hand side by an n × n identity matrix I gives Ax = λIx, which can be rewritten as Ax - λIx = 0 or (A - λI)x = 0. Assuming that x is not 0, the eigenvalues ​​can be calculated using the characteristic polynomial equation p(λ) = det(A - λI) = 0, and based on the calculated eigenvalues, the corresponding eigenvectors x can be calculated using the equation (A - λI)x = 0. For an n × n matrix A, the polynomial p(λ) is the nth degree in λ, and the n square roots λ1, λ2, ..., λ n The square roots of these values ​​are eigenvalues ​​of matrix A. i For this, the corresponding non-zero solution x (Equation 1) i Such methods exist. In practice, the eigenvalues ​​of large matrices are not calculated using characteristic polynomials because the computation is costly and calculating the exact (symbolic) square root of a higher-order polynomial can be difficult. Therefore, iterative numerical computation algorithms are used to calculate eigenvectors and estimates of eigenvalues.

[0017] In some embodiments, the eigenpair calculation process 140 performs the eigendecomposition process using an iterative numerical method known as the “power iteration method”. More specifically, in some embodiments, the eigenpair decision control process 142 implements a method configured to perform a power iteration on a symmetric matrix A in order to calculate one or more eigenpairs of matrix A. Given a symmetric (and therefore diagonalizable) matrix A, the power iteration is performed to determine the largest eigenvalue λ of matrix A (e.g., the dominant eigenvalue λ in absolute value) and the corresponding dominant (non-zero) eigenvector v corresponding to the dominant eigenvalue λ that satisfies the eigenvalue equation: Ax = λx. An exemplary power iteration is described in further detail below with reference to the exemplary flowcharts in Figures 4A and 4B.

[0018] Generally, the power iterative method begins with an initial vector, which may be an approximation of the dominant eigenvector or a random vector. The initial vector is used in an iterative process that computes a sequence of normalized vectors (or unit vectors) representing the current estimates of the dominant eigenvectors after each iteration, and it is expected that the sequence will converge to an eigenvector corresponding to the dominant eigenvalue or principal eigenvalue. In each iteration, the resulting vector x is multiplied by matrix A. Assuming that matrix A has an eigenvalue that is exactly larger in absolute magnitude than the other eigenvalues ​​of matrix A, and that the initial vector has a non-zero component in the direction of the eigenvector associated with the dominant eigenvalue, the computed sequence vector will converge to an eigenvector associated with the dominant eigenvalue. According to exemplary embodiments of this disclosure, hardware-accelerated computing is used between power iterative processes (e.g., via an RPU computing system) to perform the multiplication operation (Ax) for each iteration of the power iterative process. The power iterative process computes estimates of the dominant eigenvectors, and the corresponding dominant eigenvalues ​​may be determined, for example, by the Rayleigh quotient of the eigenvectors.

[0019] TIFF0007849129000002.tif80168

[0020] Furthermore, in some embodiments, the eigenpair determination control process 142 performs a matrix deflation process in conjunction with a power iteration process to compute additional eigenpairs of a given diagonalizable matrix A. As described above, the power iteration process is used to compute an estimate of the dominant eigenvector x1, thereby enabling the eigenpair determination control process 142 to compute an estimate of the corresponding dominant eigenvalue λ1 through any appropriate process such as the Rayleigh quotient. If there is no need to determine other eigenpairs of matrix A, the power iteration process terminates with the computation of the dominant eigenpair (λ1,x1) of matrix A. On the other hand, for one or more additional eigenpairs of an n×n matrix A (e.g., (λ2,x2),(λ3,x3),...,(λ n ,x n When determining )), the eigenpair determination control process 142 performs a matrix deflation process to calculate the deflation matrix, enabling the calculation of the dominant eigenpair of the deflation matrix, and uses a power iteration process to correspond to the next dominant eigenpair of matrix A (e.g., (λ2, x2)). An exemplary matrix deflation process is discussed in more detail below with reference to exemplary flowchart 4B.

[0021] Assume that the dominant eigenvector x1 (e.g., normalized eigenvector x1) and the corresponding dominant eigenvalue λ1 of matrix A are determined by a power-exponentiation process. In some embodiments, the matrix deflation process is performed on the matrix λ1x1x1 T We calculate this and subtract it from matrix A to obtain the deflation matrix D, i.e., D = A - λ1x1x1 T This is done by generating a deflation matrix D. The deflation matrix D has the same eigenvectors and eigenvalues ​​as matrix A, but the difference is that the dominant eigenvalue of A, which was calculated earlier, is mapped to zero in the deflation matrix D. In this respect, the next lower eigenvalue of matrix A becomes the dominant eigenvalue of the deflation matrix D, which can be determined by another iteration of the power iteration method.

[0022] TIFF0007849129000003.tif62168

[0023] TIFF0007849129000004.tif57168

[0024] TIFF0007849129000005.tif68167

[0025] TIFF0007849129000006.tif62168

[0026] As schematically shown in Figure 1, the matrix inversion process 144 performs the estimated inverse matrix A of matrix A provided by application 110. -1 The inverse matrix A was calculated and estimated. -1 A method is implemented which is configured to store in the RPU array 134. The eigenpair determination control process 142 calculates the smallest eigenpair of matrix A in the first iteration of the inverse power iteration process by using the estimated inverse matrix A -1 In some embodiments, the matrix inversion operation is performed on the inverse matrix A of matrix A. -1 To calculate the estimate of, it is performed in the digital domain using any appropriate processing. For example, in some embodiments, the matrix inversion process 144 is performed on the inverse matrix A -1 To calculate an approximate value of , implementations are made using von Neumann series processing, Newton iteration, or both, and exemplary methods thereof are known to those skilled in the art, and their details are not necessary to understand the exemplary eigenpair calculation techniques described herein.

[0027] In some embodiments, the matrix inversion process 144 implements a hardware-accelerated computing technique such as that disclosed in U.S. Patent Application No. 17 / 134,814, title: Matrix Inversion Using Analog Resistive Crossbar Array hardware, filed December 29, 2020, which is commonly assigned and incorporated entirely by reference herein. U.S. Patent Application No. 17 / 134,814 discloses a technique for performing a matrix inversion process that includes, for example, (i) storing a first estimated inverse of a given matrix A in one or more RPU arrays 134, and (ii) performing a first iteration on the first estimated inverse stored in an array of RPU cells to converge the first estimated inverse to a second estimated inverse of a given matrix. In some embodiments, the first iteration includes a stochastic gradient descent optimization process element that includes using row vectors of a given matrix A as training data to train a first estimated inverse matrix stored in an array of RPU cells, and updating the matrix values ​​of the first estimated inverse matrix stored in the array of RPU cells by using error vectors determined based on the matrix values ​​of the same matrix. Further details of the matrix inverse processing flow are described in U.S. Patent Application Serial No. 17 / 134,814, which is incorporated herein by reference.

[0028] As schematically shown in Figure 1, during the execution of application 110, application 110 can call an eigenpair calculation process 140 to calculate one or more eigenpairs of a given matrix A provided to a matrix inversion process 140. In some embodiments, matrix A includes a symmetric matrix or a symmetric positive definite (SPD) matrix. A symmetric matrix is ​​a square matrix equal to its transpose. An SPD matrix is ​​a square symmetric matrix having an orthonormal set of n eigenvectors, where all of the matrix's eigenvalues ​​are real and positive, and the eigenvectors corresponding to different eigenvalues ​​of the symmetric matrix are orthogonal to each other. SPD matrices arise in many physical and mathematical contexts where eigendecomposition of an SPD matrix is ​​required to perform calculations. For illustrative purposes, exemplary embodiments of this disclosure are described herein in the context of symmetric and SPD matrices, but exemplary techniques discussed herein can readily be applied to asymmetric matrices.

[0029] In some embodiments, the digital processing system 120 controls the execution of the eigenpair calculation process 140. As an initial step, upon receiving matrix A from application 110 with a request to compute eigenpairs of a given matrix A, the eigenpair calculation process 140 configures one or more neural cores 132 and associated RPU arrays 134 of the neuromorphic computing system 130 to provide hardware acceleration support for the eigenpair calculation process. Furthermore, the eigenpair calculation process 140 will communicate with the neuromorphic computing system 130 to store matrix A in one or more RPU arrays 134 of the neural cores 132. In some embodiments, the inverse matrix A -1 If necessary for the calculation, the unique pair calculation process 140 calls the matrix inversion process 144 to approximate the inverse matrix A of the received matrix A. -1 Calculate the inverse matrix A -1The matrix A is stored in one or more RPU arrays 134 of one or more neural cores 132 configured to support eigenpair calculation processing. The eigenpair determination control process 142 then performs numerical iteration, including, as necessary, power iteration, inverse power iteration, matrix deflation, or a combination thereof, to calculate one or more eigenpairs of matrix A. Exemplary embodiments of the eigenpair calculation method are described in further detail below.

[0030] Figure 2 schematically illustrates an RPU computing system that may be implemented in the system of Figure 1 to provide hardware-accelerated computing of matrix eigenpairs according to exemplary embodiments of the present disclosure. For example, Figure 2 schematically shows an exemplary embodiment of the neural core 132 and associated RPU array 134 of the neuromorphic computing system 130 of Figure 1. More specifically, Figure 2 schematically shows an RPU system 200 (e.g., a neuromorphic computing system) comprising a two-dimensional (2D) crossbar array of RPU cells 210 arranged in multiple rows R1, R2, R3, ..., Rm and multiple columns C1, C2, C3, ..., Cn. The RPU cells 210 of each row R1, R2, R3, ..., Rm are commonly connected to their respective row control lines RL1, RL2, RL3, ..., RLm (collectively, row control line RL). Each RPU cell 210 in columns C1, C2, C3, ..., Cn is connected in common to its respective column control lines CL1, CL2, CL3, ..., CLn (collectively referred to as column control line CL). Each RPU cell 210 is connected at (and between) one crossing point (or intersection) of the row control lines and the column control lines, respectively. In exemplary embodiments, the number of rows (m) and the number of columns (n) are the same (i.e., n = m). For example, in some embodiments, the computing system 200 comprises a 4,096 × 4,096 array of RPU cells 210.

[0031] The computing system 200 further comprises peripheral circuits 220 connected to row control lines RL1, RL2, RL3, ..., RLm, and peripheral circuits 230 connected to column control lines CL1, CL2, CL3, ..., CLn. Furthermore, peripheral circuits 220 are connected to a data input / output (I / O) interface block 225, and peripheral circuits 230 are connected to a data I / O interface block 235. The computing system 200 further comprises a control signal circuit 240 comprising various types of circuit blocks, such as power, clock, bias, and timing circuits, for power distribution and control and clocking signals for the operation of peripheral circuits 220 and 230 of the computing system 200.

[0032] In some embodiments, each RPU cell 210 of the computing system 200 includes a resistive element having an adjustable conductance value. During operation, some or all of the RPU cells 210 of the computing system 200 receive a predetermined matrix A or an approximate inverse matrix A stored in the array of RPU cells 210. -1This includes the respective conductance values ​​mapped to each numerical matrix value. In some embodiments, the resistive elements of the RPU cell 210 are implemented using resistive elements such as resistive switching devices (interface or filament switching devices), ReRAM, memristor devices, phase-change memory (PCM) devices, and other types of devices having variable conductance (or variable resistance level) that can be programmatically adjusted within a range of several different conductance levels to adjust the weight of the RPU cell 210. In some embodiments, the variable conductance elements of the RPU cell 210 can be implemented using ferroelectric devices such as ferroelectric field-effect transistor devices. Furthermore, in some embodiments, the RPU cell 210 can be implemented using an analog CMOS-based framework in which each RPU cell 210 comprises a capacitor and a readout transistor. In the analog CMOS-based framework, the capacitor functions as a memory element in the RPU cell 210, storing weight values ​​in the form of a capacitor voltage. The capacitor voltage is applied to the gate terminal of the read transistor to modulate the channel resistance of the read transistor based on the level of the capacitor voltage. The channel resistance of the read transistor represents the conductance of the RPU cell and correlates with the level of read current generated based on the channel resistance.

[0033] Although row control lines RL and column control lines CL are shown as single lines in Figure 2 for ease of illustration, it will be understood that each row and column control line may include two or more control lines connected to the RPU cell 210 in each row and column, depending on the implementation and the specific architecture of the RPU cell 210. For example, in some embodiments, each row control line RL may include a complementary pair of word lines to a given RPU cell 210. Furthermore, each column control line CL may comprise multiple control lines, for example, including one or more source lines (SL) and one or more bit lines (BL).

[0034] Peripheral circuits 220 and 230 are connected to the respective rows and columns in the 2D array of RPU cells 210 and comprise various circuit blocks configured to perform various analog, in-memory computational operations, such as vector-matrix multiplication functions, matrix-vector multiplication functions, and cross product update operations, in order to provide hardware-accelerated computing of a given matrix A intrinsic pair, according to exemplary embodiments of this disclosure. For example, in some embodiments, to support read / sensing operations of an RPU cell (e.g., reading weight values ​​of a given RPU cell 210), peripheral circuits 220 and 230 comprise pulse-width modulation (PWM) circuits and read pulse drive circuits, which are configured to generate and apply PWM read pulses to the RPU cell 210 in response to digital input vector values ​​(read input values) received during different operations. More specifically, in some embodiments, peripheral circuits 220 and 230 comprise digital-to-analog (D / A) converters configured to receive a digital input vector (applied to a row or column) and convert the elements of the digital input vector into analog input vector values ​​represented by pulse-width-varying input voltages. In some embodiments, a time coding scheme is used when the input vector is represented by a fixed-amplitude Vin=1V pulse with an adjustable duration (e.g., the pulse duration is a multiple of 1 ns and proportional to the value of the input vector). The input voltage applied to the row (or column) generates an output vector value represented by the output current, and the stored weights / values ​​in the RPU cell 210 are substantially read out by measuring the output current.

[0035] Peripheral circuits 220 and 230 receive the read current (I) output from and stored in the connected RPU cell 210. READThe system further includes a current integrator and an analog-to-digital (A / D) converter that integrate the current and convert the integrated current into a digital value (readout output value) for subsequent calculations. Specifically, the current generated in the RPU cell 210 is summed up column by column (or row by row), and the summed current is integrated over a measurement time tmeas by the current readout circuits of the peripheral circuits 220 and 230. The current readout circuit comprises a current integrator and an analog-to-digital (A / D) converter. In some embodiments, each current integrator comprises an operational amplifier that integrates the current output from a given column (or row) on the capacitor (or the differential current from a pair of RPU cells implementing negative and positive weights), and the analog-to-digital (A / D) converter converts the integrated current (e.g., an analog value) into a digital value.

[0036] The data I / O interfaces 225 and 235 are configured to interface with a digital processing core, which is configured to process inputs / outputs to the RPU system 200 (e.g., neural cores) and route data between different RPU arrays. The data I / O interfaces 225 and 235 are configured to (i) receive external control signals and data from the digital processing core and provide the received control signals and data to the peripheral circuits 220 and 230, and (ii) receive digital read output values ​​from the peripheral circuits 220 and 230 and send the digital read output values ​​to the digital processing core for processing.

[0037] Figures 3, 4A, and 4B illustrate a method for performing eigenpair calculation processing according to exemplary embodiments of the present disclosure. In particular, Figure 3 illustrates a high-level processing flow for calculating eigenpairs, which is implemented in some embodiments by the computing system 100 of Figure 1. During the runtime execution of a given application, the application may need to perform some calculations that require, for example, eigendecomposition of a given matrix A. The computing system 100 receives a request from the given application to determine one or more eigenpairs of a given matrix A (block 300). The request will include values ​​of matrix A. In some embodiments, matrix A includes a symmetric matrix, for example, an n × n matrix having n rows and n columns, where n can be relatively large (e.g., 100 or more). In some embodiments, the symmetric matrix A comprises an SPD matrix. The computing system 100 invokes eigenpair calculation processing (e.g., processing 140 in Figure 1) to calculate eigenpairs of a given matrix A.

[0038] In some embodiments, the call to the eigenpair computation process includes initial processing to configure the neuromorphic computing system 130 (e.g., an RPU system) to perform hardware-accelerated computing operations required to perform the eigenpair computation process (block 301). For example, in some embodiments, the digital signal processing system 120 communicates with the programming interface of the neuromorphic computing system 130 to configure one or more neurons and the routing system of the neuromorphic computing system 130, as will be discussed in more detail below, (i) a given matrix A or an estimated inverse matrix A -1 (ii) Implement one or more interconnected RPU arrays to store the matrix values ​​of (ii) the stored matrix A or the estimated inverse matrix A -1 Use this to allocate and configure one or more neural cores to perform in-memory calculations (matrix-vector calculations, cross product calculations, etc.).

[0039] In some embodiments, the number of RPU arrays assigned and interconnected varies depending on the size of matrix A and the size of the RPU arrays. For example, if each RPU array has a size of 4096 × 4096, one RPU array is a given n × n matrix A, or the estimated inverse matrix A of a given matrix A. -1 It can be configured to store values ​​(where n is 4096 or less). In some embodiments, when a given n×n matrix A is smaller than the physical RPU on which the n×n matrix A is stored, any unused RPU cell can be set to zero, or unused inputs to the RPU array can be padded with a "zero" voltage, or both. In some embodiments, when the size of a given n×n matrix A is larger than the size of a single RPU array, multiple RPU arrays can be operationally interconnected to store the value of a given n×n matrix A, or the estimated inverse matrix A of a given n×n matrix A. -1 A sufficiently large RPU array can be formed to store it.

[0040] Next, the inverse matrix A of a given matrix A. -1 A determination is made as to whether or not it is necessary for the eigenpair calculation process (block 302). For example, as described above, when performing an eigenpair calculation process (e.g., an exponentiation iteration process) to determine the dominant eigenpair of a given matrix A, the eigenpair calculation process will be performed using the given matrix A. On the other hand, when performing an eigenpair calculation process (e.g., an inverse exponentiation iteration process) to determine the smallest eigenpair of a given matrix A, the eigenpair calculation process will use the inverse matrix A of the given matrix A. -1 It is preferable to use the inverse matrix A of a given matrix A. -1 If it is determined that it is not necessary for the unique pair calculation process (negative determination in block 302), the digital signal processing system 120 proceeds to communicate with the neuromorphic computing system 130 and stores a predetermined matrix A in the RPU array assigned to the configured neural core (block 303).

[0041] On the other hand, the inverse matrix A of a given matrix A -1If it is determined that is necessary for the eigenpair calculation process (positive determination in block 302), in some embodiments, the digital signal processing system 120 calculates the approximate inverse matrix A of a given matrix A. -1 To determine this, we proceed with matrix inversion (block 304) and approximate inverse matrix A -1 This will be stored in the RPU array of the configured neural cores (block 303). As mentioned above, the approximate inverse matrix A -1 This is determined by the operation of the matrix inversion process 144 (Figure 1). In some embodiments, the matrix inversion process 144 is performed on the inverse matrix A -1 To estimate this, a von Neumann series inversion process is performed in the digital domain. In some embodiments, the matrix inversion process 144 utilizes hardware-accelerated matrix inversion computing techniques, such as those disclosed in U.S. Patent Application No. 17 / 134,814.

[0042] Matrix A or estimated inverse matrix A -1 Once the matrix is ​​stored in the RPU array, an eigenpair calculation process is performed to compute the eigenpair of the stored matrix (block 305). This eigenpair calculation process utilizes the RPU system to perform various hardware calculations on the stored matrix, such as matrix-vector calculations to compute eigenvectors and cross product calculations to update the values ​​of the stored matrix. In some embodiments, the eigenpair calculation process (block 305) is implemented using exemplary processing flows and calculations, such as those discussed in more detail below in relation to Figures 4A, 4B, 5A, 5B, 6A, 6B, and 7. Once the eigenpair calculation process is complete, the computing system returns the determined eigenpair to the requesting application (block 306).

[0043] Referring here to Figures 4A and 4B, an exemplary eigenpair calculation process is shown in exemplary embodiments of the present disclosure, in which a power iteration process and a matrix deflation process are used to calculate eigenpairs of a matrix. In particular, Figure 4A shows a power iteration process for calculating the dominant eigenpair of a given matrix A, and Figure 4B shows a matrix deflation process used in conjunction with the power iteration process of Figure 4A to calculate a deflation matrix, thereby enabling the calculation of the dominant eigenpair of the deflation matrix. For illustrative purposes, the processing flow in Figures 4A and 4B shows the inverse matrix A stored in the RPU array. 1 In contrast to the inverse exponentiation iteration that applies to the given matrix A, this will be discussed in the context of performing exponentiation iteration under the condition that a given matrix A is stored in the RPU array.

[0044] TIFF0007849129000007.tif93169

[0045] Next, the initial column vector x (0) The RPU system receives the input (block 401) and performs an iterative eigenvector calculation process to estimate the dominant eigenvectors corresponding to the dominant eigenvalues. In the first iteration, the RPU system generates the resulting vector (which is an n × 1 column vector) as a product of matrix-vector multiplication operations using the initial vector x (0) The analog matrix-vector multiplication operation is performed by multiplying by matrix A stored in the RPU array (block 402). In particular, in the j-th iteration, the matrix-vector operation (in block 402) is calculated as follows: x (j+1) =Ax (j) For the initial iterations, the iteration index j is set to be equal to zero (i.e., j=0), and for the initial iterations, the result vector x (1) is, x (1) =Ax (0) It is calculated as follows: x (j+1) This represents an approximation of the dominant eigenvector x1.

[0046] TIFF0007849129000008.tif75167

[0047] Next, a determination is made as to whether the system should check for convergence to a dominant eigenvector (block 404). In some embodiments, the determination is based on the number of iterations of the exponentiation process performed following a previous convergence check operation. For example, in some embodiments, the convergence check operation may be performed every p iterations of the exponentiation process, where p may be 1, 2, 3, 4, 5, etc. In some embodiments, an initial convergence check operation may be performed following a first p iteration (e.g., p=5), and the convergence check operation is then performed after each iteration of the exponentiation process until the convergence criterion is met.

[0048] If it is determined that a convergence check operation should not be performed for a given iteration (negative determination in block 404), the processing flow continues to the next iteration by inputting the current normalized vector into the RPU system (block 401) and performing analog matrix-vector multiplication by multiplying the current normalized vector by matrix A stored in the RPU array to generate the resulting updated vector as a matrix-vector multiplication product (block 402). For example, the resulting vector x (1) Following the initial iteration (j=0) in which x was calculated, the next iteration (j=j+1=1) involves the matrix-vector operation x (j+1) =Ax (j) (Block 402), x (2) Calculate x as follows: (2) =Ax (1) Result vector x (2) The output is then sent from the RPU system and normalized in the digital domain (block 403). The iteration continues again until the convergence criterion is met.

[0049] TIFF0007849129000009.tif49168

[0050] TIFF0007849129000010.tif38168

[0051] The determined error value (err) is compared to an error threshold ε to determine whether the error exceeds the error threshold ε (to determine whether err ≤ ε). In some embodiments, the error threshold ε is 1 × 10⁻¹⁰ -4 Or it can be set to a value of an order of magnitude smaller than that. The error threshold ε can be selected to any desired value depending on the application. For example, the error threshold ε can be set to a value such that the current estimate of the dominant eigenvector is accurate to three, four, five, or six decimal places of the actual dominant eigenvector.

[0052] TIFF0007849129000011.tif56167

[0053] TIFF0007849129000012.tif80167

[0054] The final step of the exponentiation iteration (blocks 407 and 408) yields an estimate of the dominant eigenpair (λ1,x1) of matrix A. If it is not necessary to determine any other eigenpairs of matrix A (negative determination in block 409), the exponentiation iteration can terminate and the dominant eigenpair (λ1,x1) of matrix A can be returned to the requesting application (e.g., block 306 in Figure 3). On the other hand, one or more additional eigenpairs of n×n matrix A (e.g., (λ2,x2),(λ3,x3),...,(λ n ,x n If it is determined that (positive determination in block 409), the processing flow proceeds to block 410 in Figure 4B. As described above, Figure 4B shows a matrix deflation process used in combination with the power iteration process in Figure 4A to compute a deflation matrix that enables the computation of the dominant eigenpair of the deflation matrix corresponding to the next dominant eigenpair of matrix A (for example, (λ2, x2)).

[0055] TIFF0007849129000013.tif74167

[0056] TIFF0007849129000014.tif79168

[0057] TIFF0007849129000015.tif38167

[0058] TIFF0007849129000016.tif55168

[0059] Figures 5A, 5B, 6A, and 6B schematically illustrate analog matrix-vector multiplication and cross product (update) calculations performed by the RPU system using matrices stored in an array of RPU cells to perform the processing of blocks 401 and 402 (Figure 4A) and blocks 412 and 413 (Figure 4B). More specifically, Figure 5A schematically illustrates a method for performing analog matrix-vector multiplication operations by an RPU computing system including an array of RPU cells to implement hardware-accelerated computing of eigenpairs of matrices, according to exemplary embodiments of the present disclosure. In some embodiments, Figure 5A shows matrix-vector multiplication operations (e.g., Ax) performed on matrix values ​​of matrices stored on an array of RPU cells 210 of the computing system 200 in Figure 2. (j) This is a schematic representation. The conductance value of RPU cell 210 is mapped to each matrix element 212 of a matrix (e.g., matrix A) stored in the array of RPU cell 210, and the matrix elements 212 stored in RPU cell 210 are encoded by the respective conductance values ​​of RPU cell 210.

[0060] TIFF0007849129000017.tif93167

[0061] TIFF0007849129000018.tif56167

[0062] TIFF0007849129000019.tif105167

[0063] In some embodiments, to determine the product of first and second vectors U and V for incremental update processing, probabilistic translator circuits within peripheral circuits 220 and 230 may be used to generate probabilistic bitstreams representing the input vectors U and V. The probabilistic bitstreams for vectors U and V are applied to the rows and columns of a 2D crossbar array of RPU cell 210, and the conductance value (and thus the corresponding matrix value) of a given RPU cell 210 will change depending on the match of the U and V probabilistic pulse streams input to the given RPU cell 210. The vector cross product operation for the update operation is performed based on the known concept that the matching of probabilistic streams representing real numbers (using AND logic gate operations) is equivalent to the multiplication operation.

[0064] Figure 6A schematically illustrates a method for configuring an RPU computing system, comprising an array of RPU cells for performing matrix-vector operations to perform hardware-accelerated computing of eigenpairs of matrices, according to an exemplary embodiment of the present disclosure. In particular, Figure 6A schematically illustrates an RPU computing system 600 comprising a crossbar array (or RPU array 605) of RPU cells 605, where each RPU cell 610 of the RPU array 605 includes an analog non-volatile resistor element (represented as a variable resistor with adjustable conductance G) at the intersection of each row (R1, R2, ..., Rn) and column (C1, C2, ..., Cn). As depicted in Figure 6A, the array of RPU cells 605 comprises a matrix A or an estimated inverse matrix A encoded by the conductance value Gij of each RPU cell 610 (where i is the row index and j is the column index). -1 This provides a matrix of conductance values ​​Gij mapped to matrix values. In an exemplary embodiment, matrix A is stored in RPU array 605, where the i-th row of the RPU cell represents the i-th row of matrix A, and the j-th column of the RPU cell represents the j-th column of matrix A.

[0065] To perform matrix-vector multiplication for exponentiation iteration (e.g., block 402 in Figure 4A), a multiplexer in the peripheral circuitry of the computing system 600 is activated to selectively connect a column driver circuit 620 to column lines C1, C2, ..., Cn. The column driver circuit 620 includes multiple digital-to-analog (DAC) circuit blocks 622-1, 622-2, ..., 622-n (collectively, DAC circuit block 622) connected to each of the column lines C1, C2, ..., Cn. Furthermore, a multiplexer in the peripheral circuitry of the computing system 600 is activated to selectively connect a readout circuit 630 to row lines R1, R2, ..., Rn. The readout circuit 630 includes multiple readout circuit blocks 630-1, 630-2, ..., 630-n connected to each of the row lines R1, R2, ..., Rn. Each readout circuit block 630-1, 630-2, ..., 630-n includes a current integrator 632-1, 632-2, ..., 632-n and an analog-to-digital (ADC) circuit 634-1, 634-2, ..., 634-n. Each current integrator includes a current integrator block, and each current integrator includes an operational transconductance amplifier (OTA) with negative capacitive feedback that converts the input current (aggregate current) into an output voltage at the output node of the current integrator. At the end of the integration period, each ADC circuit latches the output voltage generated at the output node of its respective current integrator and quantizes the output voltage to generate a digital output signal.

[0066] TIFF0007849129000020.tif92167

[0067] More specifically, in some embodiments, DAC circuit blocks 622-1, 622-2, ..., 622-n are configured to perform digital-to-analog conversion using a time coding scheme in which the input vector is represented by a fixed-amplitude pulse (e.g., V=1V) with an adjustable duration, the pulse duration being a multiple of a predetermined time period (e.g., 1 nanosecond) and proportional to the value of the input vector. For example, a given digital input value of 0.5 can be represented by a 4ns voltage pulse, and a digital input value of 1 can be represented by an 80ns voltage pulse (for example, digital input value 1 is represented by an integral time T meas (This can be encoded into analog voltage pulses having pulse durations equal to the given values.) As shown in Figure 6A, the resulting analog input voltages V1, V2, ..., V n (For example, a readout pulse) is applied to the array of RPU cells 605 via column lines C1, C2, ..., Cn.

[0068] TIFF0007849129000021.tif87167

[0069] TIFF0007849129000022.tif66167

[0070] The exemplary embodiment in Figure 6A schematically illustrates a process that performs a matrix-vector multiplication (Ax), where (i) the i-th row of the RPU cell represents the i-th row of matrix A and the j-th column of the RPU cell represents the j-th column of matrix A, (ii) a vector x is input to the column, and the output of column (iii) a result vector is generated. In other embodiments, the same matrix-vector multiplication (Ax) is performed where (i) the i-th row of matrix A represents the transpose matrix A T The transpose of matrix A is stored in RPU array 605 as the j-th column of matrix A. T This can be done by storing it in RPU array 605, (ii) applying the input vector x to the row, and (iii) reading the resulting vector from the column output.

[0071] Figure 6B schematically illustrates a method for configuring an RPU computing system, including an array of RPU cells, to perform analog cross product operations for performing hardware-accelerated computing of matrix eigenpairs, according to an exemplary embodiment of the present disclosure. More specifically, Figure 6B schematically illustrates a method for configuring an RPU computing system 600 to perform vector-vector cross product update operations (e.g., blocks 410-413 in Figure 4B) and generate deflation matrices stored in RPU array 605. Figure 6B schematically illustrates the configuration of the RPU computing system 600, in which a multiplexer in the peripheral circuitry of the computing system 600 is activated to selectively connect row driver circuits 640 to rows R1, R2, ..., Rn. Row driver circuit 6 2 0 represents multiple DAC circuit blocks 6 connected to each row R1, R2, ..., Rn. 2 2-1,6 2 2-2,...,6 2 2-n (collectively referred to as DAC circuit block 6) 2 2) is included. Furthermore, as shown in Figure 6B, for the update operation, DAC circuit block 6 4 2-1,6 4 2-2,...,6 4 Lines 2-n are connected to their respective column lines C1, C2, ..., Cn. DAC circuit block 642 performs the same function as DAC circuit block 622 described above.

[0072] TIFF0007849129000023.tif68167

[0073] TIFF0007849129000024.tif49167

[0074] The outer product update process is performed on the RPU array 605 by simultaneously applying voltage pulses representing vectors U and V to rows and columns, performing local multiplication operations and incremental weight updates at each cross point (RPU cell 610), thereby generating a deflation matrix on the RPU array 605. Again, various methods known to those skilled in the art can be used to generate analog voltage pulses (e.g., probabilistic pulses) and perform vector-vector outer product update processing in the analog domain.

[0075] FIG. 6A schematically shows an exemplary method for generating an aggregated row current for matrix-vector multiplication operations performed in the analog domain, but other techniques can be implemented to generate the aggregated current using differential current techniques that enable "signed matrix values". For example, FIG. 7 schematically shows a method for constructing an RPU computing system including multiple arrays of RPU cells for performing matrix-vector operations of eigenpair calculation processing using signed matrix values according to an exemplary embodiment of the present disclosure. In particular, FIG. 7 shows corresponding columns C1 + and Cl - from two separate RPU arrays 610 and 710 to generate a different column current I1 + and I1 - to generate an aggregated column current I COL1 , where the conductance is determined as (G + -G - ). For purposes of illustration, FIG. 7 shows a scheme for performing a matrix-vector operation Ax as described above, assuming that the transposed matrix A T is stored in the RPU array, applying the input vector x to the rows of the RPU array, and outputting the result vector from the columns. FIG. 7 schematically shows a differential readout scheme in which the column current I COL1 input to the readout circuit block 630-1 is determined as I COL1 =I1 + -I1 - . In this differential scheme, the magnitude of I COL1 corresponds to a predetermined matrix value, and the sign of the matrix value is such that I1 is I1 -It depends on whether it is greater than, equal to, or less than. Positive sign (I COL1 >0) is I1>I1 - It is obtained when (I COL1 =0) means I1=I1 - It is obtained when (I COL1 <0) is I1 <I1 - It is obtained at that time.

[0076] More specifically, in the exemplary embodiment shown in Figure 7, each RPU cell 610 of the computing system 600 in Figure 6A has a respective conductance value G ij + and G ij - It includes two unit RPU cells 610-1 and 610-2 having a conductance value of a given RPU cell 610, and the conductance value of a given RPU cell 610 is the difference between the respective conductance values, i.e., G ij =G ij + -G ij - This is determined, and i and j become indices in the RPU array 605. In this way, negative and positive weights can be easily encoded using only positive conductance values. In other words, since the conductance values ​​of the resistive devices of the RPU cell can be only positive, the difference method in Figure 7 is positive (G ij + ) and negative (G ij - To encode the matrix values ​​of a given RPU cell (G), a pair of identical RPU device arrays is implemented, and the matrix values ​​of a given RPU cell (G ij ) consists of two corresponding devices (G) located at the same position in a pair of RPU arrays 610 and 710. ij + -G ij - ) is proportional to the difference between two conductance values ​​stored in (where the two RPU arrays 610 and 710 can be stacked on top of each other in the metallization structure of the chip's backend). In this example, a single RPU tile is considered a pair of RPU arrays with peripheral circuits that support the parallel operation of the arrays in all three cycles.

[0077] As shown in Figure 7, positive voltage pulses (V1, V2, ..., V) are applied to RPU cells 610-1 and 610-2 of the corresponding rows of the same RPU array 610 and 710 used to encode the positive and negative inverse matrix values. n ) and the corresponding negative voltage pulses (-V1, -V2, ..., -V n ) are supplied individually. The corresponding first column C1 of each RPU array 610 and 710 + and C1 - The aggregated column current I1 output from + and I1 - These are combined to form the differential aggregated current I COL1 Generates the corresponding first column C1 + and C1 - This is input to the connected read circuit block 630-1.

[0078] TIFF0007849129000025.tif43167

[0079] For a given calculation, we need matrix A and its transpose (for example, A T It should be further understood that other eigendecomposition operations or linear system calculations requiring both of the above can be easily implemented using hardware acceleration methods such as those discussed herein. For example, assuming that a given matrix A is stored in an array of RPU cells as described above, a matrix-vector operation Ax can be performed by applying a vector x to the column of the RPU array and reading the resulting output vector from the row of the RPU array. Simultaneously, a matrix-vector operation A can be performed by using the matrix A stored in the RPU array, applying a vector x to the row of the RPU array, and reading the output vector from the column of the RPU array. T x can be performed. In this respect, matrix A and its transpose matrix A T There is no need to store this in different RPU arrays.

[0080] In other embodiments, the singular values ​​σ1, σ2, ..., σ of a predetermined symmetric matrix A are defined as σ1, σ2, ..., σ nTo determine this, the same or similar processing flow as shown and described above in relation to Figures 4A and 4B is used to perform singular value decomposition (SVD) processing. Generally, the process for calculating the SVD of a given matrix A is AA T and A T This includes determining the eigenvalues ​​and eigenvectors of A. Assuming that a given matrix A is a symmetric n×n matrix (but not an SPD matrix), the singular values ​​are σ1, σ2, ..., σ n This is matrix AA T or matrix A T This is determined by calculating the square root of the eigenvalues ​​of A. On the other hand, if a given matrix A is an SPD matrix, the singular values ​​σ1, σ2, ..., σ n This is matrix AA T or matrix A T It is equal to the eigenvalue of A.

[0081] TIFF0007849129000026.tif57168

[0082] In another embodiment, the SVD process on a predetermined symmetric n×n matrix A is performed on an n×n matrix B=AA T Or B=A T Eigenvalues ​​of A: λ1, λ2, ..., λ n To determine this, matrix A is stored in the RPU array, and then an iterative process is performed that includes a modified version of the processing flow in Figures 4A and 4B. In such embodiments, the RPU array is, of course, not only matrix A but also its transpose A. T This is based on the fact that it also stores. For example, referring to the exemplary embodiment in Figure 6A, suppose that the RPU array 605 stores an n×n matrix A, where the rows R1, R2, ..., Rn of the RPU array 605 represent the rows of the n×n matrix A, and the columns C1, C2, ..., Cn of the RPU array 605 represent the columns of the n×n matrix A. At the same time, from the perspective of viewing the columns of the RPU array 605 as rows and the perspective of viewing the rows of the RPU array 605 as columns, the columns of the RPU array 605 are the transpose of matrix A (A T ) represents the row, and row 605 of the RPU array is the transpose of matrix A (A TIt can be seen that this represents a column of (A). In other words, the i-th row of the n×n matrix A stored in RPU array 605 is essentially the n×n transpose of matrix A stored in RPU array 605 (A T It can be seen that this is the i-th column of ).

[0083] From the above, an exemplary process for calculating the SVD of a symmetric n×n matrix A stored in the RPU array is, for example, AA T The process is based on an iterative process of calculating the eigenvalues, and the exemplary process uses a variation of the processing flow in blocks 401, 402, and 403 of Figure 4A, AA T x (j) =A(A T x (j) )=Ay (j) =x (j+1) This includes calculating (j=0,1,...). In particular, in block 400, the initial vector x (0) The (n×1 column vector) is generated as described above. Next, the initial digital vector x (0) This is input to the RPU system, and the initial vector x (0) and the transpose A of matrix A stored in the RPU array T By multiplying by (in block 402) analog matrix-vector multiplication (i.e., A T x (0) =y (0) ) execute, and the result vector y (0) This generates matrix-vector multiplication A. In some embodiments, this is done by matrix-vector multiplication A. T x (0) is the initial vector x (0) This is performed by inputting into the row of RPU array 605 (Figure 6A), in which case the input to the row line is the DAC circuit (for example, DAC circuit 6 in Figure 6B). 2 Selectively connected to 0), resulting vector y (0) The output is from a column of the RPU array, in which case the column line is selectively connected to a readout circuit (e.g., readout circuit 630 in Figure 6A), resulting in a digital vector y (0) Generate and output the following.

[0084] Next, the resulting vector y (0)This is then re-inputted into RPU array 605, and analog matrix-vector multiplication Ay (0) Execute, and by doing so, Ay (0) =x (1) This process involves calculating the vector y. (0) This is input to column 605 of the RPU array, in which case the input to the column line is selectively connected to a DAC circuit (e.g., DAC circuit 620 in Figure 6A), and as a result, vector x (1) The output is from a row of RPU array 605, in which case the row line is selectively connected to a readout circuit (e.g., readout circuit 630 in Figure 6A), resulting in the digital vector x (1) It will generate and output the following.

[0085] Result digital vector x (1) This will then be output to a digital computing system and normalized (e.g., block 403 in Figure 4A). Next, the normalized vector x (1) The data is then re-entered into the RPU system and, as described above, performs two matrix multiplication operations (for the next iteration, j=1)AA T x (1) =A(A T x (1) )=Ay (1) =x (2) This will result in the calculation of vector x. Similar to the processing flow in Figures 4A and 4B, the iteration continues until the convergence criterion is met (e.g., block 406 in Figure 4A), in which case, following the j-th iteration, vector x (j+1) The matrix B = AA T This represents an approximation of the dominant eigenvector x1 for [the specified value].

[0086] For the example SVD process, the eigenvectors x1, x2, ..., x n and n×n matrix B=AA T The corresponding eigenvalues ​​λ1, λ2, ..., λ nTo compute some or all of (via block 408 in Figure 4A), the iterative processing flows in Figures 4A and 4B (with modifications to blocks 401, 402, and 403 as described above) may be performed. Eigenvalues ​​of matrix B λ1, λ2, ..., λ n After the calculation, the singular values ​​of the SVD process are the calculated eigenvalues ​​λ1, λ2, ..., λ of matrix B. n This is equal to (assuming matrix A is an SPD matrix). Otherwise, in the digital domain, the singular values ​​of the SVD process are the respective eigenvalues ​​λ1, λ2, ..., λ of the n×n matrix B. n It is calculated by taking the square root of . For a symmetric matrix A, in an alternative embodiment, first Ax (j) To calculate this, the rows of the RPU array are x (j) The following is entered: A T y (j) =x (j+1) The result vector y is calculated. (j) If B=A is entered in the column, the processing flow in Figures 4A and 4B is B=A T Note that this can be done to calculate the eigenvalues ​​of A.

[0087] Exemplary embodiments of the present invention may be systems, methods, or computer program products or combinations thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0088] A computer-readable storage medium can be a tangible device capable of holding and storing instructions used by an instruction execution device. A computer-readable storage medium may, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), SRAM, CD-ROM, DVD, memory stick, floppy disk, punch cards, or grooved raised structures, and mechanically encoded devices on which instructions are recorded, and suitable combinations thereof. The computer-readable storage medium as used herein should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted through wires.

[0089] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). The network consists of copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. The network adapter card or network interface of each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on the computer-readable storage medium within each computing / processing device.

[0090] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++ and procedural programming languages ​​such as the C programming language or similar programming languages. The computer-readable program instructions are executable as a standalone software package, either entirely on the user's computer or partially on the user's computer. Alternatively, they may be executable partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by personalizing them using state information of computer-readable program instructions in order to perform aspects of the present invention.

[0091] Aspects of the present invention are described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and any combination of blocks in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.

[0092] These computer-readable program instructions can be provided to a computer processor or other programmable data processing device to generate a machine, such that instructions executed via the processor of the computer or other programmable data processing device generate means for implementing functions / operations specified in one or more blocks of a flowchart or block diagram or both. These computer-readable program instructions can also be stored in a computer-readable storage medium that can be connected to a computer, a programmable data processing device, or other device or combination of devices that function in a particular way, such that the computer-readable storage medium on which the instructions are stored constitutes one of the outputs containing instructions that implement the modes of functions / operations specified in one or more blocks of a flowchart or block diagram or both.

[0093] Computer-readable program instructions, like instructions that perform a function / action specified in one or more blocks of a flowchart or block diagram or both on a computer, other programmable device, or other device, can also be loaded into a computer, other programmable data processing device, or other device and perform a series of operational steps on the computer, other programmable device, or other device to produce a computer-implemented process.

[0094] The flowcharts and block diagrams in the figures illustrate the configuration, function, and operation of executable implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or part of an instruction, which constitutes one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions shown in the blocks may differ from the order shown in the figures. For example, two blocks shown consecutively may actually be achieved as a single step, executed simultaneously, substantially simultaneously, partially or entirely in overlapping time, or the blocks may be executed in reverse order depending on the functions involved. It should also be noted that each block in a block diagram or flowchart diagram, or both, and any combination of blocks in a block diagram or flowchart diagram, or both, can be implemented by a special-purpose hardware-based system that performs a specified function or operation, or a combination of special-purpose hardware and computer instructions.

[0095] These concepts are illustrated with reference to Figure 8, which schematically illustrates an exemplary architecture of a computing node capable of hosting a system configured to perform specific pair computing operations, according to exemplary embodiments of this disclosure. Figure 8 illustrates a computing node 800 comprising a computer system / server 812, capable of operating in a number of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations or combinations thereof that may be suitable for use with the computer system / server 812 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable home appliances, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments, including any of the above systems or devices.

[0096] A computer system / server 812 may be described in the general context of a computer system executing computer system executable instructions, such as program modules. Generally, a program module can include routines, programs, objects, components, logic, and data structures that perform a specific task or implement a specific abstract data type. A computer system / server 812 may be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked over a communication network. In a distributed cloud computing environment, program modules may reside on both local and remote computer system storage media, including memory storage devices.

[0097] In Figure 8, the computer system / server 812 of the computing node 800 is shown in the form of a general-purpose computing device. The components of the computer system / server 812 may include, but are not limited to, one or more processors or processing units 816, system memory 828, and a bus 818 that connects various system components, including the system memory 828, to the processor 816.

[0098] Bus 818 represents one or more of several types of bus structures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, and processor or local buses using one of various bus architectures. Examples of such architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, Microchannel Architecture (MCA) bus, Expansion ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0099] The computer system / server 812 typically includes various computer system-readable media. Such media may be any available media accessible by the computer system / server 812, and may include both volatile and non-volatile media, as well as removable and non-removable media.

[0100] The system memory 828 may include computer system-readable media as volatile memory, such as random access memory (RAM) 830 or cache memory 832 or both. The computer system / server 812 may further include other removable / non-removable computer system-readable media and volatile / non-volatile computer system-readable media. As an example, the storage system 834 may be provided for reading and writing to a non-removable non-volatile magnetic medium (not shown; commonly referred to as a “hard drive”). Also, although not shown, a magnetic disk drive for reading and writing to removable non-volatile magnetic disks (e.g., floppy disks) and an optical disk drive for reading and writing to removable non-volatile optical disks (e.g., CD-ROMs, DVD-ROMs, or other optical media) may be provided. In these examples, each may be connected to the bus 818 by one or more data medium interfaces. As further illustrated and described below, the memory 828 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the present invention.

[0101] As an example, a program / utility 840 having a set (at least one) of program modules 842 can be stored in memory 828, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or any combination thereof, may include an implementation of a network environment. Program modules 842 generally perform functions or methods, or both, of the embodiments of the present disclosure described herein.

[0102] The computer system / server 812 can communicate with one or more external devices 814 such as a keyboard, pointing device, or display 824, one or more devices that enable interaction between the user and the computer system / server 812, or any device that enables communication between the computer system / server 812 and one or more other computer devices (e.g., a network card or modem), or a combination thereof. Such communication can be performed via the input / output (I / O) interface 822. Furthermore, the computer system 812 can communicate with one or more networks (such as a local area network (LAN), a general-purpose wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof) via the network adapter 820. As shown in the diagram, the network adapter 820 can communicate with other components of the computer system / server 812 via the bus 818. Note that other hardware components, software components, or both can be used in conjunction with the computer system / server 812, although these are not shown in the diagram. Examples of these include microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, SSD drives, and data archive storage systems.

[0103] This disclosure includes a detailed description of cloud computing, but the implementations of the teachings described herein are not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in any other type of computer environment that is currently known or may be developed in the future.

[0104] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with service providers. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.

[0105] The characteristics are as follows:

[0106] On-demand self-service: Cloud consumers can unilaterally prepare computing power, such as server time and network storage, automatically as needed, without requiring human interaction with service providers.

[0107] Broad network access: Computing power is available over the network and accessible through standard mechanisms. This facilitates utilization by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, PDAs).

[0108] Resource pooling: A provider's computing resources are pooled and delivered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated as needed. Generally, consumers have a sense of location independence because they do not manage or know the exact location of the resources provided. However, consumers may be able to identify the location at a higher level of abstraction (e.g., country, state, data center).

[0109] Rapid Elasticity: Computing power can be prepared quickly and flexibly, allowing it to scale out automatically and immediately, and to be quickly released and scale in immediately. To consumers, the computing power available for preparation often appears unlimited and can be purchased in any quantity at any time.

[0110] Measured Services: Cloud systems leverage metric capabilities at a certain level of abstraction, appropriate for the type of service (e.g., storage, processing, bandwidth, active user accounts), to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.

[0111] The service model is as follows:

[0112] Software as a Service (SaaS): The functionality offered to consumers is the ability to use the provider's applications running on a cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, except for configuring a limited number of user-specific applications.

[0113] Platform as a Service (PaaS): The functionality offered to consumers is the ability to deploy applications they have created or acquired to cloud infrastructure using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, and storage, but they can control the deployed applications and, in some cases, the configuration of their hosting environment.

[0114] Infrastructure as a Service (IaaS): The functionality provided to consumers is the provision of processors, storage, networking, and other basic computing resources that enable consumers to deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they can control the operating system, storage, and deployed applications, and in some cases, partially control certain network components (e.g., host firewalls).

[0115] The deployment model is as follows:

[0116] Private Cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by that organization or a third party and can reside on-premises or off-premises.

[0117] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common interests (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can reside on-premises or off-premises.

[0118] Public Cloud: This cloud infrastructure is provided to a large number of people or large industry groups and is owned by organizations that sell cloud services.

[0119] Hybrid Cloud: This cloud infrastructure combines two or more cloud models (private, community, or public). While maintaining the unique entities of each model, they are bound together by standards or individual technologies to achieve data and application portability (e.g., cloud bursting for load balancing across clouds).

[0120] Cloud computing environments are service-oriented environments that emphasize statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is the infrastructure, which includes a network of interconnected nodes.

[0121] Referring to Figure 9, an exemplary cloud computing environment 900 is shown. As illustrated, the cloud computing environment 900 includes one or more cloud computing nodes 950. Local computer devices used by cloud consumers (e.g., personal digital assistants (PDAs) or mobile phones 954A, desktop computers 954B, laptop computers 954C, or automotive computer systems 954N, or a combination thereof) can communicate with these nodes. The nodes 950 can communicate with each other. The nodes 950 can be grouped physically or virtually (not shown) in one or more networks, such as the private, community, public, or hybrid clouds or a combination thereof described above. This allows the cloud computing environment 900 to provide infrastructure, platforms, or software as a service, or a combination thereof, without requiring cloud consumers to maintain resources on their local computer devices. Please note that the types of computer devices 954A to N shown in Figure 9 are merely examples, and the computing node 950 and the cloud computing environment 900 can communicate with any type of electronic device via any type of network, a network addressable connection (e.g., using a web browser), or both.

[0122] Referring to Figure 10, a set of functional abstraction layers provided by the cloud computing environment 900 (Figure 9) is shown. It should be understood that the components, layers, and functions shown in Figure 10 are illustrative only, and embodiments of the present invention are not limited to these. As illustrated, the following layers and corresponding functions are provided.

[0123] The hardware and software layer 1060 includes hardware components and software components. Examples of hardware components include a mainframe 1061, a reduced instruction set computer (RISC) architecture-based server 1062, a server 1063, a blade server 1064, storage devices 1065, and a network and network components 1066. In some embodiments, the software components include network application server software 1067 and database software 1068.

[0124] The virtualization layer 1070 provides an abstraction layer. From this layer, for example, the following virtual entities can be provided: virtual servers 1071, virtual storage 1072, virtual networks 1073 including virtual private networks, virtual applications and operating systems 1074, and virtual clients 1075.

[0125] As an example, the management layer 1080 can provide the following functions: Resource preparation 1081 enables the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 1082 enables cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only protection of data and other resources but also identification and verification of cloud consumers and tasks. User portal 1083 provides consumers and system administrators with access to the cloud computing environment. Service level management 1084 enables the allocation and management of cloud computing resources to ensure that requested service levels are met. Service Level Agreement (SLA) planning and execution 1085 enables the pre-arrangement and procurement of cloud computing resources that are expected to be needed in the future in accordance with the SLA.

[0126] Workload layer 1090 provides examples of the capabilities available to the cloud computing environment. Examples of workloads and capabilities that can be provided from this layer include mapping and navigation 1091, software development and lifecycle management 1092, virtual classroom education delivery 1093, data analysis processing 1094, transaction processing 1095, and various capabilities 1096 for performing operations such as calculating eigenpairs of matrices, performing matrix operations, matrix diagonalization, and eigendecomposition operations such as singular value decomposition, using an RPU system having an RPU array, based on the exemplary methods and capabilities described above in relation to Figures 3, 4A and 4B, for example, to provide hardware-accelerated computing and analog-in-memory computing. Furthermore, in some embodiments, the hardware and software layer 1060 may include the computing system 100 of Figure 1 to implement or support various workloads and capabilities 1096 for performing such hardware-accelerated computing and analog-in-memory computing.

[0127] The descriptions of the various embodiments of this disclosure are presented for illustrative purposes only and are not intended to be exhaustive or to limit the disclosed embodiments. It will be apparent to those skilled in the art that many modifications and changes are possible without departing from the scope of the embodiments described. The terminology used herein has been selected to best describe the principles of the embodiments, their practical application to market-based technologies or technical improvements, or to enable those skilled in the art to understand the embodiments described herein.

Claims

1. A system, Processor and The processor comprises a resistor processing unit coupled to the processor, the resistor processing unit comprising an array of cells, each cell comprising a resistor device, the resistor device comprising a resistor adjustable to encode matrix values ​​that can be stored in the array of cells, The aforementioned processor, The matrix is ​​stored in the resistance processing unit by adjusting the resistance of at least some of the resistor devices in the array of cells, and the values ​​of the matrix are encoded in the resistance processing unit. Using the aforementioned resistance processing unit, the process is performed to determine the eigenvectors of the stored matrix by performing an analog matrix-vector multiplication operation on the stored matrix in order to converge the initial vector to an estimated value of the eigenvectors of the stored matrix. It is configured to do the following: When executing the above process, the processor Performing a first iteration, the first iteration being: Inputting the initial vector into the array of cells, In order to generate a first output vector from the array of cells, the resistance processing unit is used to perform a first matrix-vector multiplication operation by multiplying the stored matrix in the array of cells by the initial vector, The execution includes, Performing at least a second iteration, the second iteration being: Inputting the aforementioned first normalized vector into the cell array, The resistive processing unit is used to perform a second matrix-vector multiplication operation by multiplying the stored matrix in the cell array by the first normalized vector in order to generate a second output vector output from the array of cells, The process includes performing a normalization process that normalizes the second output vector and thereby generates a second normalized vector, A system configured to perform the following actions.

2. The system according to claim 1, wherein the initial vector includes one of a random vector and an estimate of the target eigenvector of the stored matrix.

3. Upon completion of the multiple iterations of the above process, the processor: The process involves determining whether the final output vector generated from the last completed iteration among the aforementioned multiple iterations converges to the target eigenvector of the stored matrix, In response to determining that the last output vector has converged to the target eigenvector of the stored matrix, the last output vector is set as the estimated eigenvector of the stored matrix, The system according to claim 1, configured to perform the following:

4. A system, Processor and The processor comprises a resistor processing unit coupled to the processor, the resistor processing unit comprising an array of cells, each cell comprising a resistor device, the resistor device comprising a resistor adjustable to encode matrix values ​​that can be stored in the array of cells, The aforementioned processor, The matrix is ​​stored in the resistance processing unit by adjusting the resistance of at least some of the resistor devices in the array of cells, and the values ​​of the matrix are encoded in the resistance processing unit. Using the aforementioned resistance processing unit, the process is performed to determine the eigenvectors of the stored matrix by performing an analog matrix-vector multiplication operation on the stored matrix in order to converge the initial vector to an estimated value of the eigenvectors of the stored matrix. It is configured to do the following: When executing the above process, the processor The resistive processing unit is used to update the matrix values ​​of the stored matrix in the array of cells, thereby generating an updated matrix in which the estimated eigenvalues ​​of the matrix are set to zero. In order to estimate the second eigenvector of the matrix, the process is repeated on the updated matrix stored in the cell array, To estimate the second eigenvalue associated with the aforementioned estimated second eigenvector, A system configured to perform the following actions.

5. When using the resistive processing unit to update the matrix values ​​of the stored matrix in the array of cells, the processor is configured to perform an cross product operation of a first vector and a second vector on the stored matrix. The first vector includes the estimated eigenvector scaled by the associated estimated eigenvalues, The second vector includes the transpose of the estimated eigenvector. The system according to claim 4.

6. A computer program, Receiving matrices from the application, The storage of the matrix in an array of cells of a resistance processing unit, wherein each cell includes a resistor device, and the resistor device has a resistor that can be adjusted to store the matrix in the resistance processing unit and encode the values ​​of the matrix into the resistance processing unit by adjusting the resistance of at least some of the resistor devices in the array of cells. Using the resistance processing unit, the process of determining the eigenvectors of the matrix is ​​performed, and the execution of the process includes performing an analog matrix-vector multiplication operation on the stored matrix in order to converge the initial vector to the estimated value of the eigenvectors of the stored matrix. Have the computer run it, Executing the aforementioned process means Performing a first iteration, the first iteration being: Inputting the initial vector into the array of cells, To generate a first output vector from the array of cells, a first matrix-vector multiplication operation is performed by multiplying the stored matrix in the array of cells by the initial vector, The execution includes, Performing at least a second iteration, the second iteration being: Inputting the aforementioned first normalized vector into the cell array, In order to generate a second output vector to be output from the array of cells, a second matrix-vector multiplication operation is performed by multiplying the stored matrix in the array of cells by the first normalized vector, The execution includes, , performing the normalization process which normalizes the second output vector and thereby generates a second normalized vector, A computer program that includes [this].

7. Executing the aforementioned process means Upon completion of multiple iterations of the above process, it is determined whether the final output vector generated from the last completed iteration of the multiple iterations converged to the target eigenvector of the stored matrix. In response to determining that the last output vector has converged to the target eigenvector of the stored matrix, the last output vector is set as the estimated eigenvector of the stored matrix, The computer program according to claim 6, further comprising:

8. A computer program, Receiving matrices from the application, The storage of the matrix in an array of cells of a resistance processing unit, wherein each cell includes a resistor device, and the resistor device has a resistor that can be adjusted to store the matrix in the resistance processing unit and encode the values ​​of the matrix into the resistance processing unit by adjusting the resistance of at least some of the resistor devices in the array of cells. Using the resistance processing unit, the process of determining the eigenvectors of the matrix is ​​performed, and the execution of the process includes performing an analog matrix-vector multiplication operation on the stored matrix in order to converge the initial vector to the estimated value of the eigenvectors of the stored matrix. Have the computer run it, Executing the aforementioned process means The resistive processing unit is used to update the matrix values ​​of the stored matrix in the array of cells, thereby generating an updated matrix in which the estimated eigenvalues ​​of the matrix are set to zero. In order to estimate the second eigenvector of the matrix, the process is repeated on the updated matrix stored in the cell array, Determining the second eigenvalue associated with the estimated second eigenvector, A computer program that includes [this].

9. Using the resistor processing unit to update the matrix values ​​of the stored matrix in the array of cells is, This includes performing an cross product operation between a first vector and a second vector on the stored matrix, The first vector includes the estimated eigenvector scaled by the associated eigenvalues, The second vector includes the transpose of the estimated eigenvector. The computer program according to claim 8.

10. A computing system receiving a matrix from an application, The computing system stores the matrix in an array of cells of a resistance processing unit, each cell comprising a resistor device, the resistor device having a resistor that can be adjusted to store the matrix in the resistance processing unit and encode the values ​​of the matrix into the resistance processing unit by adjusting the resistance of at least some of the resistor devices in the array of cells. The computing system performs a process to determine the eigenvectors of the matrix using the resistance processing unit, and the execution of this process includes performing an analog matrix-vector multiplication operation on the stored matrix in order to converge the initial vectors to the estimated values ​​of the eigenvectors of the stored matrix. Includes, Executing the aforementioned process means To estimate the eigenvalues ​​associated with the estimated eigenvectors, The resistive processing unit is used to update the matrix values ​​of the stored matrix in the array of cells, thereby generating an updated matrix in which the estimated eigenvalues ​​of the matrix are set to zero. In order to estimate the second eigenvector of the matrix, the process is repeated on the updated matrix stored in the cell array, This further includes estimating a second eigenvalue related to the aforementioned estimated second eigenvector, Using the resistance processing unit to update the matrix values ​​of the stored matrix in the array of cells includes performing an cross product operation of a first vector and a second vector on the stored matrix, The first vector includes the estimated eigenvector scaled by the associated eigenvalues, The method wherein the second vector includes the transpose of the estimated eigenvector.

Citation Information

Patent Citations

  • Resistive memory accelerator

    US20180068722A1