An electronic device for QR matrix decomposition
The use of memristive crossbar arrays in an electronic device addresses the high computational complexity and inflexibility of traditional QR matrix decomposition methods, achieving efficient, flexible, and low-power processing for QR decomposition.
Patent Information
- Application Number
- PCT/EP2023/086170
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-19
AI Technical Summary
The hardware implementation of QR matrix decomposition faces high computational complexity, leading to increased hardware cost and power/energy consumption, especially for large matrices. Traditional implementations are also inflexible regarding matrix size, parallelism, and input types.
An electronic device comprising memristive crossbar arrays is used for QR matrix decomposition, leveraging the inherent properties of memristive devices to perform computations efficiently and flexibly, supporting both real and complex inputs.
The proposed solution reduces computational complexity to O(1), leading to lower hardware complexity, reduced power consumption, and increased flexibility in handling different matrix sizes and input types, while maintaining high throughput.
Smart Images

Figure EP2023086170_19062025_PF_FP_ABST
Abstract
Description
[0001]AN ELECTRONIC DEVICE FOR QR MATRIX DECOMPOSITIONTECHNICAL FIELD The embodiments herein relate to an electronic device and a method for QRmatrix decomposition. A corresponding computer program and a computer program carrier are also disclosed. BACKGROUND Matrix decomposition or matrix factorization is the process of representing a matrixas a multiplication of two or more matrices, which have different special properties. Thespecial properties of constituent matrices make various matrix-based operations muchsimpler compared to the case that the same operations are done by using the originalmatrix. One of the most popular types of matrix decompositions is the QR decomposition(QRD), which is a decomposition of a matrix ^ into^ = ^^ (1)where ^ is an upper triangular matrix with real entries on the main diagonal and ^ isan orthogonal matrix such that ^^^ = ^, in which ^ is the identity matrix. That is, Q is anorthonormal matrix. Example Use Cases of QRD: There are lots of applications which require to perform complicated matrix-basedoperations. The QRD camayn be employed to reduce the computational complexity ofsuch operations. Some of these applications are: ^QRD is a very useful decomposition scheme in the context of wirelesscommunications. The QRD is used in several functional blocks including multiple- input multiple-output Dectection (MIMO Detection), Precoding (Beamforming),Channel Estimation, etc. ^QRD is often used to solve linear least squares problem^ QRD is the basis for some of the eigenvalue algorithms. It is often used incalculating the eigenvalues and eigenvectors of a matrix.^ QRD is used to solve system of linear equations for various applications.A problem with hardware implementation of QRD is a high computationalcomplexity, which increases hardware cost and power / energy consumption. This is more pronounced in case of large matrices. Also, the traditional implementations of QRD are not flexible in terms of differentsize of the matrices, degree of parallelism, real or complex inputs, etc. Therefore, thecorresponding hardware for QRD either should be designed for the worst-case scenario(e.g., largest input matrix), or it should be designed from scratch for every specificscenario. Both alternatives will increase the hardware cost significantly.The traditional hardware realization of the QRD requires memory instances to storethe intermediate results of computation in different steps of QRD. The memoryrequirement of QRD increases by the size of matrix, which increases the hardware costand energy / power consumption considerably. Due to the data dependency between successive steps / iterations of algorithms forQRD, a level of parallel processing (i.e., degree of parallelism) will be limited. As a result,the latency and throughput of the QRD kernel will be degrading especially for thedecomposition of large matrices. SUMMARY An objective of embodiments herein is to solve at least some of the drawbacksmentioned above. Embodiments herein disclose an electronic device and a method for QRD.Specifically, embodiments herein disclose an electronic device comprising amemristive crossbar array. An analog memristive crossbar array, such as a 2-dimensionalarray, consists of M×N memristive devices, each of which can be programmed torepresent an m-bit binary value. According to a first aspect, the object is achieved by an electronic device for QRmatrix decomposition of a target matrix H comprising a plurality of columns. The electronicdevice comprises one or more first memristive calculation circuits.Each first memristive calculation circuit comprises one or more first memristivecrossbar arrays, each first memristive crossbar array comprising multiple memristivedevices in a single column. Each first memristive calculation circuit further comprises two second memristivecrossbar arrays. A respective output of the one or more first memristive crossbar arrays isoperatively connected to a respective input of the two second memristive crossbar arraysand an input of the one or more first memristive calculation circuits is operativelyconnected to an input of the one or more first memristive crossbar arrays and an output ofthe two second memristive crossbar arrays is operatively connected to an output of theone or more first memristive calculation circuits. The electronic device further comprises a second memristive calculation circuitcomprising a third memristive crossbar array comprising multiple memristive devices inmultiple columns and multiple rows. The electronic device further comprises a third memristive calculation circuitcomprising a fourth memristive crossbar array comprising one row and multiple columnsof memristors, and further comprising a fourth memristive calculation circuit. An output ofthe fourth memristive calculation circuit is operatively connected to an input of the fourthmemristive crossbar array. The electronic device further comprises a fifth memristive calculation circuitcomprising a fifth memristive crossbar array comprising multiple memristive devices inmultiple columns and multiple rows. An output of the one or more first memristive calculation circuits is operativelycoupled to an input of the second memristive calculation circuit and an output of thesecond memristive calculation circuit is operatively coupled to an input of the one or morefirst memristive calculation circuits. An output of the second memristive calculation circuit is further operatively coupledto an input of the third memristive calculation circuit. An output of the third memristive calculation circuit is operatively coupled to an inputof the fifth memristive calculation circuit. According to a second aspect, the object is achieved by a method, performed by anelectronic device for QR decomposition of a target matrix ^. The electronic devicecomprises one or more first memristive calculation circuits.Each first memristive calculation circuit comprises one or more first memristivecrossbar arrays, each first memristive crossbar array comprising multiple memristive devices in a single column. Each first memristive calculation circuit further comprises two second memristive crossbar arrays. A respective output of the one or more first memristive crossbar arrays is operatively connected to a respective input of the two second memristive crossbar arrays and an input of the one or more first memristive calculation circuits is operatively connected to an input of the one or more first memristive crossbar arrays and an output of the two second memristive crossbar arrays is operatively connected to an output of the one or more first memristive calculation circuits. The electronic device further comprises a second memristive calculation circuit comprising a third memristive crossbar array comprising multiple memristive devices in multiple columns and multiple rows. The electronic device further comprises a third memristive calculation circuitcomprising a fourth memristive crossbar array comprising one row and multiple columns of memristors, and further comprising a fourth memristive calculation circuit. An output ofthe fourth memristive calculation circuit is operatively connected to an input of the fourthmemristive crossbar array. The electronic device further comprises a fifth memristive calculation circuit comprising a fifth memristive crossbar array comprising multiple memristive devices in multiple columns and multiple rows. An output of the one or more first memristive calculation circuits is operativelycoupled to an input of the second memristive calculation circuit and an output of the second memristive calculation circuit is operatively coupled to an input of the one or more first memristive calculation circuits. An output of the second memristive calculation circuit is further operatively coupled to an input of the third memristive calculation circuit. An output of the third memristive calculation circuit is operatively coupled to an input of the fifth memristive calculation circuit. The method comprising: calculating projection coefficients defined as〈^ , ^ 〉^^^ ^^=〈^ , ^ 〉, ^^where ^^ is a j:th orthogonal vector of one or more orthogonal vectors, ^ = ^ −1 and ^^ is column k of the target matrix ^, ^ = 2, … , ^,with the one or more first memristive calculation circuits based on the target matrix^, wherein a first orthogonal vector ^^ equals a first column of the target matrix H;calculating, with the second memristive calculation circuit, the one or moreorthogonal vectors based on previously calculated projection coefficients, the target matrix^ to be decomposed and previously calculated lower indexed vectors; andcalculating, with the third memristive calculation circuit, unitary vectors by normalizing the orthogonal vectors; and calculating, with the fifth memristive crossbar array, a respective row of the uppertriangular matrix ^ based on the target matrix H and the respective unitary vectors.According to a further aspect, the object is achieved by a computer program comprising instructions, which when executed by a processor, causes the processor to perform actions according to any of the aspects above. According to a further aspect, the object is achieved by a carrier comprising thecomputer program of the aspect above, wherein the carrier is one of an electronic signal,an optical signal, an electromagnetic signal, a magnetic signal, an electric signal, a radio signal, a microwave signal, or a computer-readable storage medium. Since the electronic device comprises the memristive calculation circuits and sincethe memristive devices consume much less power compared to traditional multiplyaccumulate (MAC) modules, such as digital MAC modules, the proposed PIM-based electronic calculation circuit has the potential of low power / energy consumption, which isa critical demand in many use cases (e.g., IoT devices). The lower power consumptionmay in turn lead to increased battery life for battery-powered devices such as mobilephones. The latency of the memristive-based electronic device is only limited by the readcycle of the crossbar arrays. The memristive-based electronic device is computationally efficient since thecomputational complexity is reduced to O(1) since the computations are done through theinherent features of memristive devices. This results in a significant reduction in hardwarecomplexity. The memristive-based electronic electronic device is fully flexible in terms ofconfiguration. The memristive-based electronic device supports processing of complex-valuedinput samples without any degradation in the throughput and performance. BRIEF DESCRIPTION OF THE DRAWINGS In the figures, features that appear in some embodiments are indicated by dashed lines. The various aspects of embodiments disclosed herein, including particular features and advantages thereof, will be readily understood from the following detailed description and the accompanying drawings, in which: Figure 1 is a block diagram schematically illustrating a crossbar array of memristors, Figure 2a is a block diagram schematically illustrating an electronic device and aflowchart for QR decomposition according to embodiments herein, Figure 2b is a flowchart illustrating embodiments of a method for QR decompositionaccording to embodiments herein, Figure 2c is a flowchart illustrating embodiments of a method for QR decompositionaccording to embodiments herein, Figure 3a is a block diagram schematically illustrating an electronic circuit accordingto some embodiments herein, Figure 3b is a block diagram schematically illustrating an electronic circuit accordingto some further embodiments herein, Figure 4 is a block diagram schematically illustrating an electronic circuit according to some further embodiments herein, Figure 5 is a block diagram schematically illustrating an electronic circuit according to some further embodiments herein, Figure 6 is a block diagram schematically illustrating an electronic circuit according to some further embodiments herein, Figure 7a is a block diagram schematically illustrating an electronic device for calculating a reciprocal value according to some embodiments herein, Figure 7b is a block diagram schematically illustrating an electronic device forcalculating a reciprocal value according to some embodiments herein, Figure 7c is a block diagram schematically illustrating an electronic device for multiplying according to some embodiments herein, Figure 7d is a block diagram schematically illustrating an electronic device for calculating a reciprocal value according to some embodiments herein, Figure 7e is a block diagram schematically illustrating an electronic device for calculating a reciprocal value according to some embodiments herein, Figure 8a is a block diagram schematically illustrating an electronic device for calculating a square root value according to some embodiments herein, Figure 8b is a block diagram schematically illustrating a crossbar array forinitialisation of calculating a square root value according to some further embodiments herein, Figure 8c is a block diagram schematically illustrating an electronic device for calculating a square root value according to some further embodiments herein, Figure 9 is a further flowchart illustrating embodiments of a method according toembodiments herein, Figure 10 is a block diagram schematically illustrating a memristive crossbar arrayaccording to some embodiments herein, Figure 11 is a block diagram schematically illustrating a network node.Figure 12 is a block diagram schematically illustrating a wireless communicationsdevice. Figure 13 is a block diagram schematically illustrating a wireless communicationsystem. DETAILED DESCRIPTION Algorithms, which are used to realize the QRD may be classified into two maincategories, as described below.1) A first category performs a series of right multiplications with upper triangularmatrices to orthonormalize the matrix ^. This category is called triangularorthogonalization category. One of the main schemes in this category is the Gram- Schmidt (GS) orthogonalization. The process of triangular orthogonalization maybe mathematically expressed as, ^= ^^^^^ ⋯ ^^ (2) ⋯ , ^^ are upper triangular matrices.2) A second category is called orthogonal triangularization category, in which thematrix ^ is converted into an upper triangular matrix by a series of leftmultiplications with orthonormal matrices. An important example of such algorithms is the Householder Transform (HHT) and Given’s rotations algorithms.The process of orthogonal triangularization can be mathematically described as,^^ ⋯ ^^^^^ = ^ (3)where ^^, ^^, ⋯ , ^^ are unitary matrices. As mentioned above, due to the data dependency between successivesteps / iterations of algorithms for QRD, a level of parallel processing (i.e., degree ofparallelism) will be limited. As a result, the latency and throughput of the QRD kernel willbe degrading especially for the decomposition of large matrices.Embodiments herein disclose high-throughput, fully-parallel and flexible QRD schemes using processing in memory (PIM). For example, a procedure of QRD may be started by programming a set of memristive devices with the proper values, which aredetermined by a target matrix. The memristive devices are arranged such that theyconstruct a 2-dimensional crossbar array. Then, the corresponding computations will be performed using the inherent properties of memristor devices, following Ohm’s law and Kirchhoff’s current law, to perform different steps of QRD. In order to demonstrate the proposed PIM-based QRD, the Gram-Schmidt (GS)orthogonalization algorithms is used to perform QRD. However, the idea may be similarlyapplied to other sets of algorithms like Householder Transform (HHT) and Given’srotations algorithms, etc. to perform the QRD.Embodiments herein further relate to memristive-based electronic devices for QRD.A memristor may also be referred to as a memristive device. Analog memristive deviceshave emerged as a new technology for storing and processing information in analog domain. These devices make it possible to perform computations in a place where data isstored. This concept is called in-memory computing (or processing in memory), whicheliminates the need for moving data from a memory to a processing unit. There are different types of memristive devices, which are differentiated with respect to the used materials, switching principles, device endurance, retention, etc. The main types of memristive devices include phase change memory (PCM), resistive random-access memory (ReRAM), spin-transfer torque magnetic RAM (STT-MRAM), ferroelectric memristive devices (FeRAM). Memristive devices may support a limited bit precision, attributed to the limited number of conductance levels that may be reliably programmed inthe device. For example, a PCM device may support around 50 conductance levels,meaning that it may represent around 6 bits. A number of memristor devices may be organized to form an analog crossbar array.Figure 1 illustrates an electronic calculation circuit 101 comprising a memristorcrossbar array 110 which computes MVM by calculating a dot-product of the input vectorapplied to crossbar rows (i.e., word lines) and every column of the crossbar (i.e., bit lines).The memristor crossbar array 110 is a two-dimensional array that comprises an M×Narray of memristors 111, 112, 121, 122, each of which may be programmed to representan m-bit binary value. A memristor is a tunable and programmable. The memristor maycomprise a dielectric layer sandwiched by two electrodes. A unique feature of memristorsis that the conductance depends on historical electrical signals, making them capable ofworking as nonvolatile memory. In addition, memristors may store multibit information with continuously tunable conductance, in contrast to binary states “0” and “1” in traditional digital storage systems, equipping them with higher bit density. Thus, the m-bit binary value of the memristor may be set or programmed by applying a current to the memristor.The binary value may depend on the amplitude of the current. Thus, an ^ × ^ matrix ofbinary words, G, may be represented by the memristor crossbar array 110 comprising^ × ^ memristors. The input to the memristor crossbar array 110 is an electronic inputsignal of multiple samples, such as a vector of ^ analog voltages, e.g., V, whichcorrespond to M binary values. Analog crossbar arrays comprise parallel conductors, such as metal lines, termedword lines and bit lines, respectively, as electrodes of the memristors. The word lines andbit lines may be perpendicular to each other. The memristors are formed at theintersections of word and bit lines. In embodiments herein input conductors 131 of theanalog crossbar array 110 corresponds to the word lines and output conductors 132 ofthe analog crossbar array 110 corresponds to the bit lines.The analog crossbar array 110 computes MVM by calculating the dot-product ofthe input vector applied to crossbar rows (i.e., word lines) and every column of the crossbar (i.e., bit lines), all performed in analog domain using Ohm’s law for multiplication and Kirchhoff’s law for accumulation. In Figure 1 the entries of a matrix G (an M×M matrix) are programmed to thememristive devices 111, 112, 121, 122 of the M×M crossbar array 110 while the inputvector V (an M×1 vector) is applied to the crossbar rows. Note that, the vector Vcorresponds to the actual input vector (Input 1, …, Input M), which may be converted toanalog voltages using one or more Digital to Analog Converter (DAC) modules 104illustrated in Figure 1. As a result, the following MVM may be realized using the illustratedcrossbar array 110, ^ where an output vector I is the output current of crossbar columns, which is equal tothe result of matrix-vector multiplication, i.e., I = G‧V. The output vector I may beconverted to the corresponding binary words using one or more Analog to DigitalConverter (ADC) modules 105 as shown in Figure 1. This conversion may be doneeither separately for each crossbar column (i.e., one ADC for each binary word) or in a time-multiplexed fashion and hence reduce ADC overhead (i.e., multiple bit lines mayshare one ADC 105).In this disclosure vectors and matrices are represented using capital boldface letters while their entries are shown using normal letters. Thus, when the electronic input signal is digital then the electronic calculation circuit101 further comprises the one or more DACs 104 adapted to convert the input signal ofmultiple samples to corresponding analog voltages V1, V2, … VN.In other words, when the input signal of the multiple samples is digital, the electroniccalculation circuit 101 may further comprise the DACs 104 configured to convert thedigital input signal of the multiple samples to the analog voltages. There may be one DAC 104 per input sample. In some other embodiments theremay be less than one DAC 104 per input sample as one DAC 104 may be shared amongseveral input samples by multiplexing. For example, two input samples may share the same DAC 104. Output signals will be extracted from the bit lines (columns in Figure 1) of thecrossbar array 110. If digital output values of the crossbar array 110 are needed then theoutputs of the crossbar array 110 may be converted to digital values. Thus, the electroniccalculation circuit 101 may further comprise the one or more ADCs 105 adapted toconvert the output samples, comprising analog output current, to corresponding digitaloutput values. In other words, the electronic calculation circuit 101 may further compriseADCs 105 configured to convert the output from the respective output conductor to adigital signal.If analog signals are needed in a next block in the processing chain, then the ADCs105 in the electronic device 101 may not be needed.Further, if the analog outputs are sent to another crossbar array, then they may beconverted to voltage signals, which may be done by a resistor.Embodiments herein relate to memristive-based electronic devices for mathematicaloperations. The memristive-based electronic devices for mathematical operations providehigh-throughput, high-performance, fully-parallel and flexible realization of several mathematical operations / functions using processing in memory (PIM). To this end, theparallel nature of analog crossbar arrays is employed to realize several mathematicalfunctions which leads to lower hardware cost, higher throughput, and lower latency compared to the traditional implementation approaches. Embodiments herein may comprise programming a set of memristive devices in acrossbar array with values which are determined by a target mathematicaloperation / function. Computations are performed using inherent properties of memristordevices, following Ohm’s law and Krichhoff’s current law, to generate an output of thetarget mathematical operation / function. Figure 2a shows a block diagram of embodiments herein which employ theprocessing in memory (PIM) approach. A PIM-based hardware accelerator enables to implement complicated mathematical operations of QRD efficiently. Methods forperforming QRD according to embodiments herein are presented in relation to a firstflowchart in Figure 2b and a second flowchart in Figure 2c, which are described below.In Figure 2a, the Gram-Schmidt (GS) orthogonalization algorithms is used toperform QRD. In Figure 2a, double arrows are used to show that the generated values will beprogrammed on the memristors of the specified calculation circuit by these arrows while asingle arrow is used to represent input / output signals.Thus, Figure 2a illustrates embodiments of an electronic device 200 for QR matrixdecomposition of a target matrix ^ comprising a plurality of columns.The electronic device may be configured to decompose the matrix H into an uppertriangular matrix ^ with real entries on a main diagonal and a unitary and orthogonalmatrix ^ satisfying ^^^ = ^, based on triangular orthogonalization. ^ is an identity matrix.The electronic device 200 comprises one or more first memristivecalculation circuits 210a, 210b. Each first memristive calculation circuit 210a, 210bcomprises one or more first memristive crossbar arrays 211a, 211b and two secondmemristive crossbar arrays 212, 213. Each first memristive crossbar array 211a, 211b comprises multiple memristivedevices in a single column.A respective output of the one or more first memristive crossbar arrays211a, 211b is operatively connected to a respective input of the two second memristive crossbar arrays 212, 213.An input of the one or more first memristive calculation circuits 210a, 210bis operatively connected to an input of the one or more first memristive crossbar arrays 211a, 211b. An output of the two second memristive crossbar arrays 212, 213 isoperatively connected to an output of the one or more first memristive calculationcircuits 210a, 210b. The electronic device 200 further comprises a second memristive calculation circuit220 comprising a third memristive crossbar array 221 comprising multiple memristive devices in multiple columns and multiple rows. The electronic device 200 further comprises a third memristive calculation circuit 230 comprising a fourth memristive crossbar array 232 comprising one row and multiple columns of memristors, and further comprising a fourth memristive calculation circuit 231, wherein an output of the fourth memristive calculation circuit 231 is operatively connected to an input of the fourth memristive crossbar array 232. The electronic device 200 further comprises a fifth memristive calculation circuit 240 comprising a fifth memristive crossbar array 241 comprising multiple memristive devices in multiple columns and multiple rows. An output of the one or more first memristive calculation circuits 210a, 210b is operatively coupled to an input of the second memristive calculation circuit 220 and wherein an output of the second memristive calculation circuit 220 is operatively coupled to an input of the one or more first memristive calculation circuits 210a, 210b. An output of the second memristive calculation circuit 220 is further operatively coupled to an input of the third memristive calculation circuit 230. An output of the third memristive calculation circuit 230 is operatively coupled to an input of the fifth memristive calculation circuit 240.Each one of the one or more first memristive calculation circuits 210a, 210b may beconfigured to calculate projection coefficients based on the target matrix ^.The second memristive calculation circuit 220 may be configured to calculate one ormore orthogonal vectors based on previously calculated projection coefficients, the targetmatrix ^ to be decomposed and previously calculated lower indexed orthogonal vectors ofthe one or more orthogonal vectors. The third memristive calculation circuit 230 may be configured to calculate unitaryvectors based on the orthogonal vectors. The fifth memristive crossbar array 240 may be configured to calculate a respectiverow of the upper triangular matrix ^ based on the target matrix H and the respectiveunitary vectors. A method for QRD will now be described in relation to Figure 2b. Action 281 Afirst step may be initialization, in which a first column of the target matrix, ^^, isconsidered as the first orthogonal vector, ^^ = ^^. (5)Action 282 In a next step, a second column of the target matrix, ^^, is projected onto the firstorthogonal vector, ^^. The projection of an arbitrary vector ^^onto the orthogonal vector ^^is performed as ^ where 〈^^, ^^〉 is the inner product of vectors ^^ and ^^, which is obtained as〈^^, ^^〉 = ^ ^^ ^^ . (7)Action 283 The projected vector will be subtracted from the original vector, ^^, to generate a second orthogonal vectors as ^^ = ^^ − Proj^^(^^) . (8)Similarly, a third orthogonal vector will be obtained as follows, This process will be repeated for the other columns, ^ = 1, … , ^, of the target matrix,H, where K is the number of columns in H. To this end k-th column of H will be projectedto all of the generated orthogonal columns, i.e., ^^, ^^, ..., ^^^^, and subtracted from theoriginal column, ^^, to produce the k-th orthogonal vector as^^^ The computation in (10) is realized using the second memristive calculation circuit 220 comprising an array of memristive devices, which will also be referred to asCrossbar_U. The architecture of Crossbar_U is described below.Action 284 Next, the generated orthogonal vectors may be normalized such that orthonormalvectors ^^ are obtained as^ where ‖^^‖ is the norm of vector ^^ and may be calculated as the square root ofthe inner product of 〈^^ , ^^〉. In this way, the unitary matrix ^ may be obtained. Thecomputation in equation (11) is realized using the third memristive calculation circuit 230)comprising an array of memristive devices, which will be referred to as Crossbar_Q and itis described below. Action 285 Finally, the matrix ^ in the QRD, ^ = ^^, may be calculated. To this end, a leftmultiplication by ^^is performed for the expression in (1). Having considered the propertyof matrix ^ such that ^^^ = ^, the result of this left multiplication may be simplified as,^ = ^^^ , (12)where ^ is an upper triangular matrix of size KxK. The computations in equation(12) may be expanded and expressed using multiple inner products as follows The computations in (12) and (13) are realized using the fifth memristive calculationcircuit 240 comprising the fifth memristive crossbar array 241 comprising multiplememristive devices in multiple columns and multiple rows which will also be referred to asan array of memristive devices, which is called Crossbar_R. The architecture ofCrossbar_R is described below.Action 286 The index k of the orthogonal vector U is increased. Action 287 If the index k is less than the number of columns K of the target matrix H the actionsabove may be repeated until all K rows of the matrix R has been computed. The method for QRD will now be described in relation to Figure 2c. Crossbar_U,Crossbar_Q, and Crossbar_R are used to realize different computations of the proposedQRD scheme. Although these crossbar arrays realize different operations, they work based on similar concepts. The common steps of these crossbars are mentioned in the second flowchart of Figure 2c and described below. Action 291 First, the memristive devices of the analog crossbar array is programmed with theentries of corresponding matrix. Action 292 Next, the input samples for the target operation (i.e., a certain step in QRD) areconverted to the corresponding analog voltages using one or multiple DACs.Action 293 These analogue signals will be applied to the crossbar rows.Action 294 After a “read cycle”, the generated signals in the crossbar columns become valid. Action 295 Then, the ADC(s) convert the output signals of the crossbar columns to the corresponding digital values. Actions 296-298 This process will be repeated for the next set of input samples. However, if thetarget matrix / vector changes, the crossbar should be programmed with the entries of new target matrix / vector. Embodiments for the Crossbar_U, Crossbar_Q, and Crossbar_R are presentedbelow. Note that, to simplify the figures DACs and ADCs are not shown in the figures describing these embodiments.PIM-based architectures to realize QRD SchemeVector projection will now be described. The projection of vector ^^ onto anorthogonal vector ^^is performed following equation (6), in which the projection coefficient is defined as ^ where ^ = 1, … , ^ − 1 and ^ = ^ while ^ is the size of orthogonal vectors.Since the first column of matrix H is considered as the initial orthogonal vector (^^ = ^^),the index ^ in equation (14) gets values from 2 to ^. The computation in (14) is performedusing the one or more first memristive calculation circuits 210a, 210b also referred to asCoefficient Generator blocks, which are illustrated in Figure 2a and explained below.Coefficient Generator Block (Coeff Generator in Figure 2a) / the one or morefirst memristive calculation circuits 210a, 210b An operation in (14) is the inner product, which may be realized using PIM. Someembodiments of the one or more first memristive calculation circuits 210a, 210b (theCoefficient Generator block) is shown in Figure 3a. In some embodiments herein the oneor more first memristive calculation circuits 210a, 210b include two crossbars to performthe inner products of 〈^^, ^^〉 and 〈^^ , ^^〉. Each of these crossbars may include ^memristive devices. Thus, the two values, which are generated in the crossbar columns are corresponding to the results of two inner products in (14). The output values of these two crossbars will be used to generate the required projection coefficients, through a reciprocal calculation block 211, which may be PIM-based, followed by a multiplication,which may also be PIM-based, as shown in Figure 3a.Figure 2a illustrates that, in order to perform QRD for a ^ × ^ matrix, (^ − 1) blocksof Coefficient Generator are needed.Note that the input to the crossbar rows of “PIM-based Inner Product1” in Figure 3aare the entries of previously generated orthogonal vector ^^ . These values are alreadyconverted from current to voltage signals using the current-to-voltage-convertor blocks, similar to the ones shown in Figure 3a. The current-to-voltage-convertor block receives a current signal and generates the corresponding voltage signal to be used as the input of the next memristive-based module. One case of realization of such a block is to use a resistor. Having considered the computations in equation (10), to generate k-th orthogonalvector, ^^, ( ^ − 1) projection coefficientsare needed, which will be calculated by following equation (14). These coefficients will be included in the Vector of ProjectionCoefficients (^^), which is used in the next step for generating orthogonal vectors.Thus, in some embodiments herein the target matrix ^ comprises a plurality K ofcolumns. As indicated above, a respective memristive calculation circuit of the one ormore first memristive calculation circuits 210a, 210b may be configured to calculateprojection coefficients, defined as where ^^ is a j:th orthogonal vector of the one or more orthogonal vectors, ^ =1, … , ^ − 1 and ^^ is column k of the target matrix, ^ = 2, … , ^.A first orthogonal vector ^^ may equal a first column of the target matrix H.Then each first memristive crossbar array 211a, 211b of the one or more firstmemristive calculation circuits 210a, 210b may comprise K memristive devices and eachfirst memristive crossbar array 211a, 211b of a j:th first memristive calculation circuit210a, 210b may be configured to calculate either a first inner product of a correspondingj:th orthogonal vector ^^ with itself, 〈^^, ^^〉 or a second inner product of the correspondingj:th orthogonal vector ^^ with a k-th column of the target matrix ^^, 〈^^ , ^^〉.The two second memristive crossbar arrays 212, 213 may be configured tocalculate a division of the second inner product by the first inner product. To do so a firstof the two second memristive crossbar arrays 212, 213 may be configured to calculate areciprocal value of the first inner product of the respective j:th orthogonal vector ^^ withitself, 〈^^, ^^〉. The first of the two second memristive crossbar arrays 212, 213 may bereferred to as a reciprocal calculation circuit 212. A second of the two secondmemristive crossbar arrays 212, 213 may be configured to perform multiplication of the second inner product with the reciprocal of the first inner product. The second of the twosecond memristive crossbar arrays 212, 213 may be referred to as a multiplier circuit213. The multiplier circuit 213 may comprise a single memristor configured to beprogrammed by the reciprocal of the first inner product and configured to receive the second inner product as an input value to the memristor. Figure 3b illustrates some other embodiments wherein the one or more firstmemristive crossbar arrays only comprises a single crossbar array 211c which isconfigured to calculate both the first inner product of the j:th orthogonal vector ^^ withitself, 〈^^, ^^〉 and the second inner product of the j:th orthogonal vector ^^ with the k-thcolumn of the target matrix ^^, 〈^^, ^^〉. For both calculations the memristors of the singlecrossbar array 211c are configured with the same values. For the first inner product theinput to the memristors of the single crossbar array 211c is the j:th orthogonal vector ^^. For the second inner product the input to the memristors of the single crossbar array 211cis the k-th column of the target matrix ^^. A multiplexer 214 may be operatively arrangedbetween the single memristive crossbar array 211c and the two second memristivecrossbar arrays 212, 213. A first output of the multiplexer 214 is arranged between themultiplexer 214 and the reciprocal calculation circuit 212. A second output of the multiplexer 214 is arranged between the multiplexer 214 and the multiplier circuit 213. Thefirst inner product is fed to the reciprocal calculation circuit 212. The second inner productis fed to the multiplier circuit 213.Structure of Vector of Projection Coefficients (^^),The Vector of Projection Coefficients (^^) may be used to generate the k-thorthogonal vector, where ^ = 2, ... , ^. As mentioned before, the first column of matrix Hmay be considered as the first orthogonal vector (^^ = ^^), meaning that there is no needto generate ^^for k=1. The Vector of Projection Coefficients (^^) may have (2^ -2) entries. The first entrymay always equal ^^^, which represents the coefficient for projection of the k-th column ofmatrix H onto the first orthogonal vector. All the next ^ − 1 entries of ^^ may be zeroexcept the k-th entry, which may equal one. This leads to include the k-th column of matrix H in the computation of the k-th orthogonal vector. More specifically, to generate ^^, the k-th entry of corresponding coefficient vector is 1 to include ^^, which is impliedby equation (10).So far, the first ^ entries of ^^ have been specified. The next (^ − 2) entries of ^^represent the required coefficients to project the k-th column of matrix H onto theorthogonal vectors ^^, ^^, …, ^^^^. These coefficients are ^^^, ^^^, ..., ^^^^^. Note that thek-th column of matrix H is already projected on ^^, as described above whichcorresponds to ^^^. The remaining (^ − ^) entries of ^^ may be set to zero.According to equation (10), to generate the k-th orthogonal column ^^, the k-th column of matrix H should be projected into the ^^, ^^, …, ^^^^, using the corresponding coefficients ^^^, ^^^, ..., respectively, which are included in the Vector of Projection Coefficients (^ ^^). As a result, while ^ = 1, ... , ^ − 1 are generated using the Coefficient Generator blocks 210a, 210b in Figure 2a, the Vector of ProjectionCoefficients (^^) may be created without any extra computation.To clarify this concept, let’s consider a 4x4 matrix ^ and define the required Vectorsof Projection Coefficients. Having considered equation (10), theCoefficients to create orthogonal columns, ^^, ^^, and ^^ are ^ respectively. Here, ^^^, ^^^, and ^^^are the required coefficients to project ^^, ^^, and ^^ onto ^^. Similarly, ^^^and ^^^are the required coefficients to project ^^and ^^onto ^^. Finally, ^^^is the required coefficient to project ^^onto ^^. As describedbefore, in the proposed scheme, the first column of the target matrix may be consideredas the first orthogonal vector, i.e., ^^ = ^^.Note that, the negation signs for^ which are implied by equation (10), may be realized via different approaches: 1) It may be included in the PIM-based multiplication in Figure 3a such that the^ ^ memristor should be programmed by −〈^^,^^〉instead of〈^^,^^〉.2) It may be included in Crossbar_U such that the memristors are programmed by−^^instead of ^^. 3) It may be realized by a separate PIM-based multiplication by -1, which should beplaced after each Coefficient Generator block in Figure 2a. Generate the Orthogonal Vectors (^^) The Vector of Projection Coefficients (^^) will be used to generate the orthogonalvectors following equation (10). Figure 4 illustrates the third memristive crossbar array 221 comprising multiplememristive devices in multiple columns and multiple rows which may be used to realizecomputation in equation (10). Thus, the second memristive calculation circuit 220 may be configured to calculate a k:th orthogonal vector ^^of the one or more orthogonal vectors by being configured toreceive the previously calculated projection coefficients, the target matrix ^ to bedecomposed and previously calculated lower indexed orthogonal vectors and apply thepreviously calculated projection coefficients, the target matrix ^ to be decomposed andpreviously calculated lower indexed orthogonal vectors to the third memristive crossbararray 221 which is configured to calculate the one or more orthogonal vectors ^^ as^^^ The the third memristive crossbar array 221 may comprise (2^ − 2) rows and ^columns. As also mentioned above, the third memristive crossbar array 221) is also referredto as Crossbar_U. Thus, the Crossbar_U may have (2^ − 2) rows and ^ columns, i.e.,(2^ − 2) × (^) in total. The first ^ rows of Crossbar_U are programmed by the entries oftarget matrix, ^, such that the k-th column of ^ is programmed into the memristivedevices of the k-th row of Crossbar_U. Also, the last (^ − 2) rows of Crossbar_U areprogrammed with the unitary vectors ^^, where ^ = 2, … , ^ − 1.In other words, a first ^ rows of the third memristive crossbar array 221 may beconfigured to be programmed by entries of the target matrix ^, such that the k-th columnof ^ is programmed into the memristive devices of the k-th row of the third memristivecrossbar array 221, and a last (^ − 2) rows of the third memristive crossbar array 221 areconfigured to be programmed with the orthogonal vectors ^^, where ^ = 2, … , ^ − 1. Note that the unitary vectors ^^will be generated one by one. Therefore, once aunitary vector is generated, it will be programmed to Crossbar_U as well. As mentionedabove, Crossbar_U receives the projection coefficients and generates the correspondingorthogonal vectors following equation (10). Generate Matrix Q In order to generate the columns of matrix Q, the orthogonal vectors ^^should be normalized. Thus, j-th orthonormal vector, ^^, is obtained by normalization of j-thorthogonal vector, ^^. This may be done using embodiments illustrated in Figure 5.As mentioned above, the third memristive calculation circuit 230 is configured to calculate unitary vectors based on the orthogonal vectors. In other words, the thirdmemristive calculation circuit 230 may be configured to normalize the orthogonal vectorsinto unitary vectors. In some embodiments the fourth memristive crossbar array 232 is configured to beprogrammed with the values of a respective j:th orthogonal vector ^^ and the fourthmemristive calculation circuit 231 is configured to calculate a square root of a reciprocal ofa square of the norm of the respective j:th orthogonal vector ^^. The fourth memristive crossbar array 232 is also referred to as Crossbar_Q. Thesize of Crossbar_Q is 1xK, where K is the size of the orthogonal vectors, i.e., size of thecolumns in the target matrix H. In case of normalization of ^^, the input to Crossbar_Q is the reciprocal of the norm^ value of this vector, The squared value of this scaling factor was already calculated. Generate Matrix R The other component of QRD, ^ = ^^, is the matrix ^. As mentioned in equation(12) the matrix ^ may be calculated by a left multiplication of the target matrix ^ by ^^.This may be performed using the fifth memristive calculation circuit 240 comprising thefifth memristive crossbar array 241, also referred to as Crossbar_R in Figure 6. To thisend, Crossbar_R is programmed with the entries of the target matrix ^. Then, eachcolumn of matrix ^ will be applied to the crossbar rows and the corresponding row ofmatrix ^ will be generated along the crossbar columns.PIM-based QRD for Rectangular MatricesSo far it is assumed that the target matrix, ^, is a square matrix of size of ^ × ^.However, the proposed QRD scheme may be extended to the rectangular matrices ofsize ^ × ^, where ^ > ^. In such a case, the QRD of ^ = ^^ results in the matrix ^ ofsize ^ × ^, which includes ^ orthogonal vectors of unit length, and the matrix ^, which isan upper triangular matrix of size ^ × ^.In general, it is possible to perform the full QRD for a rectangular matrix ofsize ^ × ^ by adding (^ − ^) arbitrary orthonormal columns to the matrix ^, and convert itto a ^ × ^ matrix. In this case, the matrix ^ will be an upper triangular matrix ofsize ^ × ^, in which the lower (^ − ^) rows of matrix ^ consist entirely of zeroes.PIM-based QRD for Complex-Valued Matrices QRD for complex-valued matrices may also be computed using embodiments herein. In such cases, extra crossbars are needed to realize complex multiplications, additions, and subtractions. For example, in order to realize equation (10) for a complex-valued target matrix H, two crossbars similar to the one in Figure 4 are needed. One ofthese crossbars performs the computations in equation (10) for the real parts and the second one performs the same computations for the imaginary parts. Moreover, in case of complex-valued vector / matrix, the inner product is defined as 〈^^, ^^〉 = ^^^^^(15) where ^^^ is the Hermitian of vector ^^, which is basically conjugate transpose of^^. The computations in equation (15) may be realized using two crossbars formultiplications of real and imaginary parts. Also, in case of complex-valued matrices, the matrix ^ can be calculated likeequation (12). The only change is that the Hermitian of matrix ^ will be used, i.e.,^ = ^^^ . (16)Solving Systems of Linear Equations using the PIM-based QRD One of the applications of QR decomposition is solving systems of linear equations.Another important field where QR decomposition is often used is in calculatingthe eigenvalues and eigenvectors of a matrix. Let’s consider to solve the following systemof equations: ^^,^^^ + ^^,^^^ + ⋯ + ^^,^^^ = ^^^^,^^^ + ^^,^^^ + ⋯ + ^^,^^^ = ^^⋮ This system may be rewritten as:^^ = ^ (17)The goal is to determine the vector ^ given the matrix ^ and the vector ^.To this end, by employing QRD, the matrix A is decomposed to an orthogonalmatrix Q and the upper-triangular matrix R such that A = QR. Thus, the equation (17) maybe rewritten as ^^^ = ^ . (18)Now, both sides of the above equation can be multiplied by ^^,^^^^^ = ^^^ . (19)Since matrix Q is an orthogonal matrix, it satisfies the property that ^^^ = ^, where^ is the identity matrix. As a result, the left-hand side of equation (19) may be simplifiedas, ^^ = ^^^ . (20)Since the matrix R is a triangular matrix, the equation in (20) is straightforward tosolve. To clarify this concept, let’s solve the following system of linear equations via the QRD. 4^^ + 3.5^^ − 0.45^^ = 2.3^^ − 0.95^^ + 12.5^^ = −9.52^^ + 1.3^^ − 4.1^^ = 42Having considered the above system of equations, the corresponding matrix and vectors are 43.5 − 0.45^^2.3 ^= ^ 1 − 0.95 12.5 ^ , ^ = ^^^ ^ , ^ = ^−9.5 ^. 21 .3 − 4.1^^42 Performing QRD for the matrix A leads to,−0.8729 0.2911 − 0.3916 −4.5826 − 3.4151 − 0.5455^ = ^ −0.2182 − 0.9507 − 0.2203 ^, ^ = 0 1.7831 − 11.5769 ^ .−0.4364 − 0.1068 0.8934 0 0 − 6.2401By substituting the result of QRD in equation (20), the system of equation is converted to, By starting from the last row, the unknown variables can be calculated as follows,^^ = −6.2, ^^ − 37.36, ^^ = 32.56 .All of the required operations to solve the system of linear equations are matrix-vector multiplications, which may be simply performed using crossbar arrays. As a result,the proposed PIM-based QRD scheme may be used to solve a system of linear equationsefficiently. Mapping into Multiple Crossbars with different Shapes and Sizes Sometimes the size of the matrix H, or in general the matrix to be decomposed, islarge such that it cannot be programmed to a single crossbar array. In these cases, thetargeted matrix H may be divided into several sub-matrices, which may be programmed tomultiple crossbar arrays, which are smaller than the size of the targeted matrix. Embodiments herein are independent of the Shape and Size of the crossbar arrays.As a result, the target matrix H may be programmed into multiple smaller crossbars, whichhave different sizes and shapes. This feature leads to a higher hardware utilization sincethe size of required crossbars, which are used to cover the matrix H, may be chosenaccording to the size of corresponding sub-matrices.PIM-based QRD with Word SlicingIn order to improve the accuracy of the QRD and / or to reduce the hardware cost by using simpler and cheaper memristive devices, it is possible to assign more than one memristor to each matrix entry. In the other words, each value in the targeted matrix willbe divided into multiple parts, which will be programmed to multiple devices of a crossbarrow. For example, let’s consider each entry of matrix H is represented by 8 bits, i.e., =ℎ^^ℎ^^ℎ^^ℎ^^ℎ^^ℎ^^ℎ^^ℎ^^. As explained in the previous sections of this document, the coefficient ^^is programmed into one memristor. However, in the word slicing scenario, the coefficient programmed into more than one memristors. Assuming two memristors per word, ℎ^^ℎ^^ℎ^^ℎ^^is programmed to one memristor and ℎ^^ℎ^^ℎ^^ℎ^^will be programmed to the second one. Similar to the bit-serial input scenario, a “Shift & Add” block may be employed tocalculate the output samples of the QRD by combining the outputs of the corresponding crossbar columns.Bit-Serial Inputs for PIM-based QRD The supported bit resolution of ADCs determines the quantization error; the higher supported bit resolution the lower the quantization error and hence the better accuracy. Therefore, in case that either the ADCs do not support the required resolution or in order to improve the performance, each binary word of the input vector can be sent to the DACs and consequently to the crossbar rows in a bit-serial manner. Figure 10 shows a crossbar array which performs the same operation as the one inFigure 1 while the bit-serial scheme is employed for the input vector. At each timeinstance, which corresponds to the read cycle of the crossbar array, one bit of all inputbinary-words is applied to the corresponding DAC. In Figure 10, the i-th bit of j-th input and output words are shown by ^^^and ^^^ , respectively. Thus, after W read cycles, theMVM is completed where W is the number of bits per input binary-word. In this scheme,the Shift and Add blocks 1040 are used to calculate the final results of each crossbarcolumn. Some embodiments herein may take advantage of performing mathematicaloperations with memristive crossbar arrays. Embodiments of a first electronic device for performing a mathematical operation onan input value to calculate an output value by calculating the output value as a reciprocal of the input value, will be presented in relation to Figures 7a-7e. Embodiments of a second electronic device for performing a mathematicaloperation on an input value to calculate an output value by calculating the output value asa square root of the input value will be presented in relation to Figures 8a-8c. The firstand second electronic devices may also be referred to as a first system of electronic calculation circuits and a second system of electronic calculation circuits. Thus embodiments presented in relation to Figures 7a-7e and 8a-8c are for performing a mathematical operation comprising either a calculation of an output value as a reciprocal of an input value or a calculation of the output value as a square root of theinput value. In other words, the first electronic device calculates the reciprocal of theinput, while second electronic device calculates the square root of the input.Both the first and second electronic device may comprise a memristive crossbararray which comprises at least three memristors operatively arranged in series andconfigured to be programmed by a respective memristor value. The memristive crossbar array further comprises at least two inputs.The inputs to the at least three memristors and values of the at least threememristors may be configured based on a Newton-Raphson method for performing themathematical operation on the input value. In the following figures the memristors are programmed with memristor values, which are written near the memristors in the corresponding figure. PIM-based Reciprocal Calculation The reciprocal operation may be realized using Newton–Raphson’s method. TheNewton–Raphson method is an iterative method which may be used to estimate an outputof a function, ^(^), for a given input, ^^. This process may be mathematically expressedas ^ where ^(^) and ^′(^) are the targeted function and its derivative, ^^ is the inputvalue at ^-th iteration, and ^ is the maximum number of iterations. Depending on thetarget function and the required accuracy / performance, the number of iterations, ^, may be specified. One problem in the hardware implementation of mathematical functions usingNewton–Raphson method is that higher number of iterations reduces the throughput whileit increases the hardware cost and latency significantly. In embodiments herein theparallel nature of a memristive crossbar is employed to realize one or more mathematicalfunctions which leads to lower hardware cost, higher throughput, and lower latency compared to the traditional implementation approaches. Let’s consider the reciprocal operation, i.e.,^ ^. In order to realize this operationusing Newton–Raphson method, the function ^(^) should be considered as where ^ is the input number, which we want to calculate the reciprocal of and ^^ is^ the output of reciprocal function, i.e., ^ , in the ^-th iteration. Having considered the concept of Newton–Raphson method in equation (84), an iterative process of reciprocal calculation may be expressed as ^^^^^ = 2^^ − ^. ^^, ^ = 1, … , ^ (85) ^ where ^^and ^^^^are the output of the reciprocal function, i.e., ^ , in the ^-th and(^ + 1)-th iterations, respectively.In this process ^^ is an initial value, which may be obtained from a lookup table(LUT). It has been shown that a 4-bit LUT, which stores 16 initial values, results in a very good accuracy / performance. However, in order to improve the accuracy, a 6-bit LUT maybe used. In general, in case of an m-bit LUT, the first m bits of binary representation of ^may be used as the address of the LUT to generate the initial value, ^^.Embodiments of the first electronic device 700 for reciprocal calculation usingprocessing in memory (PIM) is illustrated in Figure 7a, which may include four blocks: ^A LUT 701, which may be used to store the initial values of the reciprocalcalculation as described above. ^A DAC 702, which may be used to convert the initial value, ^^, as well as the inputvalue to the corresponding analog values. ^A memristive-based reciprocal calculation circuit 703, which includes multiplememristors to realize the presented equations. ^An ADC 704, which is used to convert a final output to a corresponding digitalvalue. The LUT 701 may be a digital LUT or an analog LUT e.g., using capacitors in analog domain. This means that the ADC / DACs are not mandatory. Some first embodiments of the reciprocal calculation circuit 703 using one iterationis shown in Figure 7b. In Figure 7b the reciprocal calculation circuit 703 comprises amemristive crossbar array 703b. This architecture realizes^ ^^ = 2^^ − ^. ^^, (86)where ^^ is the initial value and ^^ is the final result, i.e., ^^ ≈ 1 / ^.In Figure 7b the memristive crossbar array 703b comprises three memristors711b, 712b, 713b operatively arranged in series. The three memristors 711b, 712b, 713bare configured to be programmed by a respective memristor value. A first memristor 211bmay be configured with 1. A second memristor 712b may also be configured with 1. Athird memristor 713b may be configured with −^^^. The memristive crossbar array 703b ofFigure 7b further comprises three inputs, one per memristor. A first input to the firstmemristor 711b may be ^^. A second input to the second memristor 712b may be ^^ anda third input to the third memristor 713b may be ^.The reciprocal calculation circuit 703 may further comprise DACs at every input tothe memristive crossbar array 703b. Likewise, the reciprocal calculation circuit 703 maycomprise ADCs at the output of the memristive crossbar array 703b. Also for the rest ofthe described embodiments below such DACs and ADCs may be used but are notillustrated nor described in the following. Since the value of ^^ may be stored in the LUT 701, it is also possible to store thesquare of the initial values (^^^) in the LUT 701 as well. However, another solution toobtain ^^^or −^^^is to employ a separate memristive-based multiplier 705 for this purpose, like the one shown in Figure 7c. For the above first embodiments it was assumed that the LUT 701 stores the squareof the initial values. However, it is possible to implement the same operation using adifferent architecture. To this end, the computations in equation (85) may be rewritten as Some second embodiments of the memristive-based reciprocal calculation circuit703 for reciprocal calculation using one iteration is shown in Figure 7d. This architecturefollows equation (87), in which the LUT 701 may store the initial values but do not need tostore the square of the initial values (^^^). In Figure 7d a memristive crossbar array 703d-1 comprises three memristors711d, 712d, 713d operatively arranged in series. The three memristors 711d, 712d, 713dare configured to be programmed by a respective memristor value. A first memristor 711dmay be configured with 2. A second memristor 712d may also be configured with −^^^. Athird memristor 713d may be configured with ^^. The memristive crossbar array 703d ofFigure 7d further comprises two inputs. A first input to the first memristor 211d may be 1.A second input to the second memristor 712d may be ^.Thus, the memristive crossbar array 703b, 703d, 703e of the electronic reciprocalcalculation circuit 703 comprises at least three memristors, such as the three memristors711b, 212b, operatively arranged in series and configured to be programmed by a respective memristor value. The memristive crossbar array 703b, 703d, 703e further comprises at least two inputs. As mentioned above, in some embodiments herein thememristive crossbar array 703b, 703d, 703e comprises two or three inputs.In Figure 7d as well as in coming other figures, horizontal thick lines between someof the memristors represent current-to-voltage conversion with a current-to-voltageconverter 715, which may be implemented using a resistor. Thus, the memristivecrossbar array 703d may comprise multiple single-column crossbar sub-arrays 703d-1, 703d-2. At least two of the multiple single-column crossbar sub-arrays 703d-1, 703d-2 may be operatively connected via current-to-voltage converters. In some embodiments herein the multiple single-column crossbar sub-arrays 703d- 1, 703d-2 are two. There are several options to improve the accuracy of the reciprocal calculation. Oneoption is to increase the number of bits to represent ^, ^^, and ^^. This does not changethe presented architectures in Figure 7b and Figure 7d. Another solution to improve theaccuracy is to increase the number of iterations. According to simulation results, twoiterations, i.e., ^ = 2, will result in a very good accuracy. To this end, the computations inequation (86) may be expanded as follows:^^ = 2^^ − ^. ^^^ (88)^^ = 2^^ − ^. ^^^ (89)By employing the presented architectures in Figure 7b or Figure 7d, the final result(i.e., ^^ ≈ 1 / ^) may be obtained in two crossbar read cycles. By repeating this process fora larger number of iterations, more accurate results may be obtained. However, adrawback of this scenario is that some of the memristors in the crossbar should be programmed twice and also the whole operation takes two crossbar read cycles. In order to solve these issues, the complete computations may be implemented atonce. By substituting equation (88) into equation (89) the final result of the reciprocalcalculation after two iterations, i.e., ^^ ≈ 1 / ^, may be obtained as,^ ^ ^ ^^ = 4^^ − 6^. ^^ + 4^^. ^ ^^ − ^ . ^^ . (90)A memristive crossbar array 703e to realize the computations of equation (90) isshown in Figure 7e. In this architecture, the memristors are programmed only once with the depicted values and the input value, ^, is sent to the crossbar rows as shown inFigure 7e. As a result, the final output, ^^ ≈ 1 / ^, will be generated after only one readcycle of the memristive crossbar array 703e. A higher number of iterations (^ > 2) may berealized by following this idea. To this end, the computations in equations (86) may beexpanded for the desired ^ and then a corresponding structure of the crossbar may bespecified. Note that such a procedure may be done offline. Thus, the electronic calculation circuit 703 may be configured to calculate thereciprocal value of the input value based on two iterations of the Newton-Raphsonmethod. Then the memristive crossbar array 703e may comprise seven inputs and mayfurther comprise seven memristive crossbar sub-arrays operatively connected via six current-to-voltage converters. The seven memristive crossbar sub-arrays may comprise eighteen memristors and wherein a first crossbar sub-array 703e-1 is operatively connected to a seventh crossbar sub-array 703e-7 via a first current-to-voltage converter 721, a second crossbar sub-array 703e-2 is operatively connected to a third crossbar sub- array 703e-3 via a second current-to-voltage converter 722, the third crossbar sub-array 703e-3 is operatively connected to a fourth crossbar sub-array 703e-4 and a fifth crossbarsub-array 703e-5 via a third current-to-voltage converter 723 , the fourth crossbar sub-array 703e-4 is operatively connected to the seventh crossbar sub-array 703e-7 via a fourth current-to-voltage converter 724, the fifth crossbar sub-array 703e-5 is operatively connected to a sixth crossbar sub-array 703e-6 via a fifth current-to-voltage converter 725 and the sixth crossbar sub-array 703e-6 is operatively connected to the seventh crossbarsub-array 703e-7 via a sixth current-to-voltage converter 726.As may be seen in Figure 7e, in some embodiments herein at least a further two of the multiple single-column sub-arrays 703d-1, 703d-2 are arranged in parallel. PIM-based Division Atraditional way to perform division is to use off-the-shelf digital division circuitry,which have a high hardware cost. For high throughput and high numeric precision,embodiments herein are based on a PIM architecture to perform division.In embodiments herein division may be performed by using a reciprocal operationaccording to the previously described embodiments followed by a multiplication using amemristive crossbar array. For example, a dividend Y is to be divided by a divisor X. Thereciprocal of X, i.e., 1 / X, may be calculated according to the previously describedembodiments. The reciprocal of the divisor, 1 / X, is then multiplied with Y.In order to improve the accuracy, as described above, a higher number of bits maybe used to represent X and Y, and a higher number of iterations may be used in thereciprocal calculation. PIM-based Square Root (SQRT)In many applications calculation of a square root of a number is needed. A simpleexample is to solve ^^ = ^. However, there are more complicated use cases in wirelesscommunication systems which require SQRT calculation, e.g., Cholesky decomposition in MIMO detection, Norm Calculation of the received Beam Power, etc. The SQRT operation may be realized using one or more calculation circuits, at least one of them comprising amemristive crossbar array configured based on the Newton–Raphson method. This leadsto a lower hardware cost, higher throughput, and lower latency compared to traditional implementation approaches. In order to calculate√^, the target function is defined as ^(^ ) = ^ ^^ − ^^, (91)where ^ is the input number, of which the square root is to be calculated, and ^^ isthe output of the SQRT function, i.e., √^ , in the ^-th iteration.Having considered the concept of Newton–Raphson method in (83), the iterative process of SQRT calculation may be expressed as ^ where ^^ and ^^^^ are the output of SQRT function, i.e., √^ , in the ^-th and (^ +1)-th iterations, respectively. Some embodiments of the fourth memristive calculation circuit 231 to calculate √^is illustrated in Figure 8a. A first action may be initialization, in which ^^ is an initial value.The initial value may be set to ^^ = ^, which result in a good accuracy / performance. Also,by choosing this initial value, there is no need for a lookup table to save the initial values.As a result, by substituting ^^ = ^ in (92), the output of the initialization step will be equalto ^ The fourth memristive calculation circuit 231 may comprise an initializationcalculation circuit 410. In some embodiments herein the initialization calculation circuit410 comprises a memristive crossbar array 410-1 illustrated in Figure 8b. Thememristive crossbar array 410-1 comprises two memristors which are configured to beprogrammed by a respective memristor value. A first memristor 411b of the memristivecrossbar array 410-1 may be configured with 0,5. A second memristor 412b of thememristive crossbar array 410-1 may also be configured 0,5. The memristive crossbararray 410-1 of Figure 8b further comprises two inputs. A first input 401b to the firstmemristor 411b and a second input 402b to the second memristor 412b.A first input value to the first memristor 411 may be the input number, ^, of whichthe square root is to be calculated. A second input value to the second memristor 212dmay be 1.A reciprocal of the output of the initialization, ^^, may be calculated using the PIM-based reciprocal calculation circuit 700 above. Then an iteration of a SQRT calculation,which follows equation (92), may be calculated using embodiments of a square rootcalculation circuit 430 illustrated in Figure 8c. At the (i+1)-th iteration, a memristivecrossbar array 430-1 of the square root calculation circuit 430 receives the input number^ (^), the output of a previous iteration (^^) and its reciprocal (^^) to calculate the SQRT of ^ in (i+1)-th iteration as √^ = ^^^^ . (94)In some embodiments herein the memristive crossbar array 430-1 comprises twomemristive crossbar sub-arrays, such as a first memristive crossbar sub-array 431 anda second memristive crossbar sub-array 432. The memristive crossbar array 430-1may comprise three memristors in total which are configured to be programmed by arespective memristor value. The first memristive crossbar sub-array 431 may comprise afirst memristor 411c and a second memristor 412c. The first memristor 411c of thememristive crossbar array 430-1 may be configured with a value of 1. The secondmemristor 412c of the memristive crossbar array 410-1 may be configured with the inputnumber, ^, of which the square root is to be calculated.The second memristive crossbar sub-array 432 may comprise a third memristor 413c. The memristive crossbar array 430-1 of Figure 8c further comprises two externalinputs. A first input 401c to the first memristor 411c and a second input 402c to thesecond memristor 412c. Afirst input value to the first memristor 411c may be the output of a previousiteration (^^). A second input value to the second memristor 412c may be the reciprocal ofthe output of the previous iteration (1 / ^^). A third input value to the third memristor 413cmay be the output of the second memristor 412c of the first memristive crossbar sub-array 431. Depending on a required accuracy of the respective application, the square rootcalculation as well as the reciprocal calculation may be repeated similarly for (K-1) times,where K is the maximum number of iterations. For a certain application, an efficient valueof K may be obtained from simulations in advance and then a number of required blocks(i.e., iterative steps) in Figure 8a may be specified.Embodiments herein may be combined to implement an electronic device forperforming a mathematical operation on one or more input values to calculate an outputvalue. The electronic device may also be referred to as a system of electronic calculationcircuits and may besides the memristive crossbar arrays also comprise DACs and ADCs and LUTs and other suitable electronics. The mathematical operation may for example be one of: division, mean square, mean squared error, root mean square, weighted average, arithmetic mean, harmonic mean, vector norm, Frobenius norm of a matrix, or Euclidian distance. Some advantages of PIM-based hardware accelerators for mathematical / arithmetic functions include: ^A fully parallel architecture for multiple arithmetic functions, which may achieve anultra-high throughput. ^The latency of embodiments herein is only limited by the read cycle of thecrossbar arrays, and it is not limited by the complexity of the function, number of inputs, etc. ^Since the memristive devices consume much lower power compared to traditionalmultiply accumulate (MAC) modules, the proposed PIM-based hardware accelerators have the potential of low power / energy consumption, which is acritical demand in many use cases (e.g., IoT devices). ^Embodiments herein support processing of complex-valued inputs in a similar wayas real-valued inputs without any degradation in the throughput and performanceof the electronic device. ^The presented PIM-based hardware accelerators are computationally efficientcompared to the traditional coding schemes since the computational complexity of embodiments herein is reduced to ^(1).^ Embodiments herein are scalable in terms of number of inputs, bit resolution, etc.^ Embodiments herein support serial, partial parallel, and fully parallelimplementation of different arithmetic functions while they enable a tradeoff between the achieved accuracy / performance and hardware cost.Embodiments of a method for QR decomposition of the target matrix ^ will now bedisclosed in relation to a flowchart of Figure 9. The method may be preformed by theelectronic device 200. The method actions may be performed in any suitable order.In some embodiments herein decomposing the target matrix ^ comprisesdecomposing the matrix H into an upper triangular matrix ^ with real entries on a maindiagonal and a unitary and orthogonal matrix ^ satisfying ^^^ = ^, wherein ^ is an identitymatrix, based on triangular orthogonalization. Action 901 The method comprises calculating projection coefficients defined as^ where ^^ is a j:th orthogonal vector of one or more orthogonal vectors, ^ = 1, … , ^ − 1 and^^ is column k of the target matrix ^, ^ = 2, … , ^, with the one or more first memristivecalculation circuits 210a, 210b) based on the target matrix ^. A first orthogonal vector ^^equals a first column of the target matrix H. Action 902 The method further comprises calculating, with the second memristive calculationcircuit 220), the one or more orthogonal vectors based on previously calculated projectioncoefficients, the target matrix ^ to be decomposed and previously calculated lowerindexed vectors. In some embodiments herein calculating a k:th orthogonal vector ^^ of the one ormore orthogonal vectors comprises projecting a k-th column of the target matrix H to all ofthe calculated orthogonal vectors, ^^, ^^, ..., ^^^^, and subtracting the projections fromthe original column, ^^ , of the target matrix H to produce the k-th orthogonal vector as^^^ Proj^^(^^) Action 903 The method further comprises calculating, with the third memristive calculationcircuit 230), unitary vectors by normalizing the orthogonal vectors.Action 904The method further comprises calculating, with the fifth memristive crossbar array240, a respective row of the upper triangular matrix ^ based on the target matrix H andthe respective unitary vectors. Figure 11 illustrates a network node 601 of a wireless communications network170, the network node 601 comprising any of the electronic calculation circuits disclosedabove or any of the electronic devices disclosed above.Figure 12 illustrates a wireless communications device 602 comprising any ofthe electronic calculation circuits disclosed above or any of the electronic devices disclosed above. The network node 601 and the wireless device 602 may be configured to performthe method actions of Figure 9 above.The embodiments herein may be implemented through a processor or one ormore processors, such as the processor 1504, 1604 of a processing circuitry in thenetwork node 601 and the wireless device 602 respectively and depicted in Figure 11 and12 together with computer program code for performing the functions and actions of theembodiments herein. The program code mentioned above may also be provided as a computer program product, for instance in the form of a data carrier carrying computerprogram code for performing the embodiments herein when being loaded into the networknode 601 and the wireless device 602 respectively. One such carrier may be in the formof a CD ROM disc. It is however feasible with other data carriers such as a memory stick. The computer program code may furthermore be provided as pure program code on aserver and downloaded to the network node 601 and the wireless device 602 respectively.The network node 601 and the wireless device 602 respectively may furthercomprise a memory 1502, 1602 comprising one or more memory units. The memorycomprises instructions executable by the processor in the network node 601 and thewireless device 602 respectively.The respective memory 1502, 1602 is arranged to be used to store e.g. information,data, configurations, and applications to perform the methods herein when beingexecuted in the network node 601 and the wireless device 602 respectively.In some embodiments, a computer program 1503, 1603 comprises instructions,which when executed by the at least one processor, cause the at least one processor ofthe network node 601 and the wireless device 602 respectively to perform the actionsabove. In some embodiments, a carrier 1505, 1605 comprises the computer program,wherein the carrier is one of an electronic signal, an optical signal, an electromagnetic signal, a magnetic signal, an electric signal, a radio signal, a microwave signal, or a computer-readable storage medium. The network node 601 and the wireless device 602 respectively may furthercomprise an input and output interface, I / O, 1506, 1606 configured to communicate withother devices. The input and output interface 1506, 1606 may comprise a receiver, suchas a wireless receiver, (not shown) and a transmitter, such as a wireless transmitter, (not shown). Those skilled in the art will also appreciate that the units described above may refer to a combination of analog and digital circuits, and / or one or more processors configuredwith software and / or firmware, e.g., stored in the network node 601 and the wirelessdevice 602 respectively, that when executed by the respective one or more processorssuch as the processors described above. One or more of these processors, as well as the other digital hardware, may be included in a single Application-Specific Integrated Circuitry (ASIC), or several processors and various digital hardware may be distributed among several separate components, whether individually packaged or assembled into a system-on-a-chip (SoC). Figure 13 illustrates a wireless communications network 170 in whichembodiments herein may be implemented. The wireless communications network 170 may use a number of different technologies, such as Wi-Fi, Long Term Evolution (LTE), LTE-Advanced, 5G, New Radio (NR), Wideband Code Division Multiple Access (WCDMA), Global System for Mobile communications / enhanced Data rate for GSM Evolution (GSM / EDGE), Worldwide Interoperability for Microwave Access (WiMax), or Ultra Mobile Broadband (UMB), just to mention a few possible implementations. Embodiments herein relate to recent technology trends that are of particular interest in a 5G context. However, embodiments are also applicable in further development of other existing wireless communication systems suchas e.g., WCDMA and LTE and in future wireless communication systems, such as 6Gsystems. Network nodes operate in the wireless communications network 170 such as thenetwork node 601. The network node 601 provides radio coverage over a geographicalarea, a service area referred to as a cell 15, which may also be referred to as a beam or a beam group of a first radio access technology (RAT), such as 5G, LTE, Wi-Fi or similar.There may be more than one cell. For example, there may be a second cell 16 as well.The network node 601 may be a NR-RAN node, transmission and reception point e.g. abase station, a radio access node such as a Wireless Local Area Network (WLAN) accesspoint or an Access Point Station (AP STA), an access controller, a base station, e.g. a radio base station such as a NodeB, an evolved Node B (eNB, eNode B), a gNB, a base transceiver station, a radio remote unit, an Access Point Base Station, a base station router, a transmission arrangement of a radio base station, a stand-alone access point or any other network unit capable of communicating with a wireless device within the servicearea depending e.g. on the radio access technology and terminology used. Therespective network node 601 may be referred to as a serving radio access node andcommunicates with a UE with Downlink (DL) transmissions to the UE and Uplink (UL) transmissions from the UE. A number of wireless communications devices operate in the wireless communication network 170, such as the wireless communications device 602. The wireless communications device 602 may be a mobile station, a non-accesspoint (non-AP) STA, a STA, a user equipment and / or a wireless terminal, thatcommunicate via one or more Access Networks (AN), e.g., RAN, e.g. via the networknode 601 to one or more core networks (CN) e.g. comprising a CN node 13, for examplecomprising an Access Management Function (AMF). It should be understood by the skilled in the art that “UE” is a non-limiting term which means any terminal, wireless communication terminal, user equipment, Machine Type Communication (MTC) device,Device to Device (D2D) terminal, or node e.g., smart phone, laptop, mobile phone,sensor, relay, mobile tablets or even a small base station communicating within a cell. When using the word "comprise" or “comprising” it shall be interpreted as non- limiting, i.e. meaning "consist at least of". The embodiments herein are not limited to the above-described preferred embodiments. Various alternatives, modifications and equivalents may be used.
Claims
CLAIMS1. An electronic device (200) for QR matrix decomposition of a target matrix ^comprising a plurality of columns, the electronic device (200) comprising: one or more first memristive calculation circuits (210a, 210b) eachcomprising: one or more first memristive crossbar arrays (211a, 211b), each first memristive crossbar array (211a, 211b) comprising multiple memristive devices in asingle column; and two second memristive crossbar arrays (212, 213); wherein a respective output of the one or more first memristive crossbar arrays (211a, 211b) is operatively connected to a respective input of the two second memristive crossbar arrays (212,213) and wherein an input of the one or more first memristive calculation circuits(210a, 210b) is operatively connected to an input of the one or more first memristive crossbar arrays (211a, 211b) and wherein an output of the two second memristive crossbar arrays (212, 213) is operatively connected to an output of the one or morefirst memristive calculation circuits (210a, 210b); a second memristive calculation circuit (220) comprising a third memristive crossbar array (221) comprising multiple memristive devices in multiple columns and multiple rows; a third memristive calculation circuit (230) comprising a fourth memristive crossbar array (232) comprising one row and multiple columns of memristors, and further comprising a fourth memristive calculation circuit (231), wherein an output of the fourth memristive calculation circuit (231) is operatively connected to an input ofthe fourth memristive crossbar array (232); and a fifth memristive calculation circuit (240) comprising a fifth memristive crossbar array (241) comprising multiple memristive devices in multiple columns and multiple rows; wherein an output of the one or more first memristive calculation circuits (210a, 210b) is operatively coupled to an input of the second memristive calculation circuit (220) and wherein an output of the second memristive calculation circuit (220) is operatively coupled to an input of the one or more first memristive calculation circuits (210a, 210b); wherein an output of the second memristive calculation circuit (220) is further operatively coupled to an input of the third memristive calculation circuit (230);wherein an output of the third memristive calculation circuit (230) is operatively coupled to an input of the fifth memristive calculation circuit (240).
2. The electronic device (200) according to claim 1, configured to decompose the matrixH into an upper triangular matrix ^ with real entries on a main diagonal and a unitaryand orthogonal matrix ^ satisfying ^^^ = ^, wherein ^ is an identity matrix,based on triangular orthogonalization.
3. The electronic device (200) according to claim 2, wherein:each one of the one or more first memristive calculation circuits (210a, 210b)is configured to calculate projection coefficients based on the target matrix ^;the second memristive calculation circuit (220) is configured to calculate oneor more orthogonal vectors based on previously calculated projection coefficients, the target matrix ^ to be decomposed and previously calculated lower indexed orthogonalvectors of the one or more orthogonal vectors;the third memristive calculation circuit (230) is configured to calculate unitaryvectors based on the orthogonal vectors; and the fifth memristive crossbar array (240) is configured to calculate a respective row of the upper triangular matrix ^ based on the target matrix H and the respectiveunitary vectors.
4. The electronic device (200) according to claim 3, wherein the target matrix ^comprises a plurality K of columns; wherein a respective memristive calculation circuit of the one or more first memristive calculation circuits (210a, 210b) is configured to calculate projection coefficients, defined as ^where ^^ is a j:th orthogonal vector of the one or more orthogonal vectors, ^ =1, … , ^ − 1 and ^^ is column k of the target matrix, ^ = 2, … , ^,wherein a first orthogonal vector ^^ equals a first column of the target matrix H;wherein each first memristive crossbar array (211a, 211b) of the one or morefirst memristive calculation circuits (210a, 210b) comprises K memristive devices andwherein each first memristive crossbar array (211a, 211b) of a j:th first memristive calculation circuit (210a, 210b) is configured to calculate either a first inner product of acorresponding j:th orthogonal vector ^^ with itself, 〈^^, ^^〉 or a second inner product ofthe corresponding j:th orthogonal vector ^^ with a k-th column of the target matrix ^^,wherein the two second memristive crossbar arrays (212, 213) are configuredto calculate a division of the second inner product by the first inner product;wherein the third memristive calculation circuit (230) is configured to normalizethe orthogonal vectors into unitary vectors and wherein the fourth memristive crossbar array (232) is configured to be programmed with the values of a respective j:th orthogonal vector ^^and wherein the fourth memristive calculation circuit (231) is configured to calculate a square root of a reciprocal of a square of the norm of therespective j:th orthogonal vector ^^.
5. The electronic device (200) according to claim 4, wherein a first of the two secondmemristive crossbar arrays (212, 213) is configured to calculate a reciprocal value of the first inner product of the respective j:th orthogonal vector ^^ with itself, 〈^^, ^^〉.
6. The electronic device (200) according to any of claims 4-5, wherein the secondmemristive calculation circuit (220) is configured to calculate a k:th orthogonal vector^^of the one or more orthogonal vectors by being configured to receive the previously calculated projection coefficients, the target matrix ^ to be decomposed andpreviously calculated lower indexed orthogonal vectors and apply the previouslycalculated projection coefficients, the target matrix ^ to be decomposed andpreviously calculated lower indexed orthogonal vectors to the third memristivecrossbar array (221) which is configured to calculate the one or more orthogonalvectors ^^as ^^^7. The electronic device (200) according to any of claims 3-6, wherein the thirdmemristive crossbar array (221) comprises (2^ − 2) rows and ^ columns.
8. The electronic device (200) according to claim 7, wherein a first ^ rows of the thirdmemristive crossbar array (221) are configured to be programmed by entries of the target matrix ^, such that the k-th column of ^ is programmed into the memristivedevices of the k-th row of the third memristive crossbar array (221), and wherein a last(^ − 2) rows of the third memristive crossbar array (221) are configured to beprogrammed with the orthogonal vectors ^^, where ^ = 2, … , ^ − 1.
9. A network node (601) of a wireless communications network (170), the network node(601) comprising the electronic device (200) of any of claims 1-8.
10. A wireless communications device (602) comprising the electronic device (200) of anyof claims 1-8.
11. A method, performed by an electronic device (200), for QR decomposition of a targetmatrix ^, the electronic device (200) comprising:one or more first memristive calculation circuits (210a, 210b) eachcomprising: one or more first memristive crossbar arrays (211a, 211b), each firstmemristive crossbar array (211a, 211b) comprising multiple memristive devices in asingle column; and two second memristive crossbar arrays (212, 213); wherein a respective output of the one or more first memristive crossbar arrays (211a, 211b) is operatively connected to a respective input of the two second memristive crossbar arrays (212,213) and wherein an input of the one or more first memristive calculation circuits(210a, 210b) is operatively connected to an input of the one or more first memristivecrossbar arrays (211a, 211b) and wherein an output of the one or more firstmemristive calculation circuits (210a, 210b) is operatively connected to an output of the two second memristive crossbar arrays (212, 213); a second memristive calculation circuit (220) comprising a third memristive crossbar array (221) comprising multiple memristive devices in multiple columns and multiple rows; a third memristive calculation circuit (230) comprising a fourth memristive crossbar array (232) comprising one row and multiple columns of memristors, and further comprising a fourth memristive calculation circuit (231), wherein an output of the fourth memristive calculation circuit (231) is operatively connected to an input of the fourth memristive crossbar array (232); and a fifth memristive calculation circuit (240) comprising a fifth memristive crossbar array (241) comprising multiple memristive devices in multiple columns and multiple rows;wherein an output of the one or more first memristive calculation circuits (210a,210b) is operatively coupled to an input of the second memristive calculation circuit (220) and wherein an output of the second memristive calculation circuit (220) is operatively coupled to an input of the one or more first memristive calculation circuits(210a, 210b); wherein an output of the second memristive calculation circuit (220) is further operatively coupled to an input of the third memristive calculation circuit (230); wherein an output of the third memristive calculation circuit (230) is operatively coupled to an input of the fifth memristive calculation circuit (240); the method comprising: calculating (901) projection coefficients defined as^where ^^ is a j:th orthogonal vector of one or more orthogonal vectors, ^ = 1, … , ^ − 1and ^^ is column k of the target matrix ^, ^ = 2, … ,with the one or more first memristive calculation circuits (210a, 210b) basedon the target matrix ^, wherein a first orthogonal vector ^^ equals a first column of thetarget matrix H; calculating (902), with the second memristive calculation circuit (220), the oneor more orthogonal vectors based on previously calculated projection coefficients, thetarget matrix ^ to be decomposed and previously calculated lower indexed vectors;and calculating (903), with the third memristive calculation circuit (230), unitaryvectors by normalizing the orthogonal vectors; and calculating (904), with the fifth memristive crossbar array (240), a respectiverow of the upper triangular matrix ^ based on the target matrix H and the respectiveunitary vectors.
12. The method according to claim 11, wherein decomposing the target matrix ^comprises decomposing the matrix H into an upper triangular matrix ^ with realentries on a main diagonal and a unitary and orthogonal matrix ^ satisfying ^^^ = ^,wherein ^ is an identity matrix, based on triangular orthogonalization.
13. The method according to claim 11 or 12, wherein calculating (902) a k:th orthogonalvector ^^ of the one or more orthogonal vectors comprises projecting a k-th column ofthe target matrix H to all of the calculated orthogonal vectors, ^^, ^^, ..., ^^^^, andsubtracting the projections from the original column, ^^, of the target matrix H toproduce the k-th orthogonal vector as ^^^14. A computer program (1403), comprising computer readable code units which whenexecuted on a computer causes the computer to perform the method according to any one of claims 11-13.
15. A carrier (1405) comprising the computer program (1403) according to the precedingclaim, wherein the carrier (1405) is one of an electronic signal, an optical signal, a radio signal and a computer readable medium.
Citation Information
Patent Citations
Eigenvalue decomposition with stochastic optimization
US11366876B2
Programmable Spatial Array for Matrix Decomposition
US20230297538A1