An electronic device and a method for permuting an electronic input signal of multiple samples
Patent Information
- Application Number
- EP2023716230
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-29
- Publication Date
- 2026-02-11
AI Technical Summary
Existing technologies face challenges in achieving high-throughput and real-time processing for interleaving, deinterleaving, and bit-reversal operations in wireless communication systems, particularly due to memory conflicts and limited parallelism in hardware implementation, which restricts the throughput of decoders and FFT processors.
An electronic device employing an electronic crossbar array with programmable memristive switches, configured according to a permutation matrix, enables fully parallel and flexible permutation of input signals, eliminating memory conflicts and supporting various permutation algorithms and patterns, thereby enhancing decoder and FFT processor performance.
The solution significantly improves throughput by eliminating memory conflicts and supporting complex permutation algorithms, leading to ultra-reliable communication and high-throughput processing in wireless communication systems.
Smart Images

Figure EP2023058094_03102024_PF_FP_ABST
Abstract
Description
[0001] AN ELECTRONIC DEVICE AND A METHOD FOR PERMUTING AN ELECTRONIC
[0002] INPUT SIGNAL OF MULTIPLE SAMPLES
[0003] TECHNICAL FIELD
[0004] The embodiments herein relate to an electronic device and a method for permuting an electronic input signal of multiple samples. A corresponding computer program and a computer program carrier are also disclosed.
[0005] BACKGROUND
[0006] In a typical wireless communication network, wireless devices, also known as wireless communication devices, mobile stations, stations (STA) and / or User Equipments (UE), communicate via a Local Area Network such as a Wi-Fi network or a Radio Access Network (RAN) to one or more core networks (CN). The RAN covers a geographical area which is divided into service areas or cell areas. Each service area or cell area may provide radio coverage via a beam or a beam group. Each service area or cell area is typically served by a radio access node such as a radio access node e.g., a Wi-Fi access point or a radio base station (RBS), which in some networks may also be denoted, for example, a NodeB, eNodeB (eNB), or gNB as denoted in 5G. A service area or cell area is a geographical area where radio coverage is provided by the radio access node. The radio access node communicates over an air interface operating on radio frequencies with the wireless device within range of the radio access node.
[0007] Specifications for the Evolved Packet System (EPS), also called a Fourth Generation (4G) network, have been completed within the 3rd Generation Partnership Project (3GPP) and this work continues in the coming 3GPP releases, for example to specify a Fifth Generation (5G) network also referred to as 5G New Radio (NR). The EPS comprises the Evolved Universal Terrestrial Radio Access Network (E-UTRAN), also known as the Long Term Evolution (LTE) radio access network, and the Evolved Packet Core (EPC), also known as System Architecture Evolution (SAE) core network. E- UTRAN / LTE is a variant of a 3GPP radio access network wherein the radio access nodes are directly connected to the EPC core network rather than to RNCs used in 3G networks. In general, in E-UTRAN / LTE the functions of a 3G RNC are distributed between the radio access nodes, e.g. eNodeBs in LTE, and the core network. As such, the RAN of an EPS has an essentially “flat” architecture comprising radio access nodes connected directly to one or more core networks, i.e. they are not connected to RNCs. To compensate for that, the E-UTRAN specification defines a direct interface between the radio access nodes, this interface being denoted the X2 interface.
[0008] Wireless communication systems in 3GPP
[0009] Figure 1 illustrates a simplified wireless communication system. Consider the simplified wireless communication system in Figure 1 , with a UE 12, which communicates with one or multiple access nodes 103-104, which in turn is connected to a network node 106. The access nodes 103-104 are part of the radio access network 10.
[0010] For wireless communication systems pursuant to 3GPP Evolved Packet System, (EPS), also referred to as Long Term Evolution, LTE, or 4G, standard specifications, such as specified in 3GPP TS 36.300 and related specifications, the access nodes 103-104 corresponds typically to Evolved NodeBs (eNBs) and the network node 106 corresponds typically to either a Mobility Management Entity (MME) and / or a Serving Gateway (SGW). The eNB is part of the radio access network 10, which in this case is the E-UTRAN (Evolved Universal Terrestrial Radio Access Network), while the MME and SGW are both part of the EPC (Evolved Packet Core network). The eNBs are inter-connected via the X2 interface, and connected to EPC via the S1 interface, more specifically via S1-C to the MME and S1-U to the SGW.
[0011] For wireless communication systems pursuant to 3GPP 5G System, 5GS (also referred to as New Radio, NR, or 5G) standard specifications, such as specified in 3GPP TS 38.300 and related specifications, on the other hand, the access nodes 103-104 corresponds typically to an 5G NodeB (gNB) and the network node 106 corresponds typically to either an Access and Mobility Management Function (AMF) and / or a User Plane Function (UPF). The gNB is part of the radio access network 10, which in this case is the NG-RAN (Next Generation Radio Access Network), while the AMF and UPF are both part of the 5G Core Network (5GC). The gNBs are inter-connected via the Xn interface, and connected to 5GC via the NG interface, more specifically via NG-C to the AMF and NG-U to the UPF.
[0012] Permutation, or reordering, of data samples of an electronic signal is important for several electronic applications. For example, both interleaving and bit-reversal in connection with Fast Fourier Transform (FFT) and Inverse FFT (IFFT) make use of permutation of data samples. Both applications are important in present and future communications system. Furthermore, high-throughput and real-time processing has become essential demands in such applications. Demands that can be challenging to fulfill with existing technology.
[0013] To describe these challenges in more detail these two applications and existing technology will be described.
[0014] Figure 2 illustrates a data processing chain 200 in a wireless communication system. On the transmitting side the data processing chain 200 may comprise an encoder 201 , such as a forward error-correcting encoder. The transmitting side of the data processing chain 200 may further comprise an interleaver 202 and a symbol mapper 203. The interleaver 202 may be part of the encoder 201.
[0015] On the receiving side the data processing chain 200 may comprise a symbol demapper 204, a deinterleaver 205 and a decoder 206. The deinterleaver 205 may be part of the decoder 206. Further, the decoder 206 may comprise further interleavers 207 as well.
[0016] The interleaver 202 is an important part of the data processing chain 200 in wireless communication systems. Interleavers are widely used in conjunction with some error correcting codes, i.e., channel coding, to improve error correction capabilities of coding schemes. The interleaver 202 when part of the encoder 201 is used to achieve time diversity in digital data communication. In general, the interleaver 202 disperses a sequence of bits or a sequence of symbols to minimize effects of burst errors during transmission. Burst errors are errors which are introduced by channels and they are localized in a short time-interval. Thus, such errors occur in multiple successive bits rather than occurring in bits independently of each other. The Interleaver 202 breaks up fading dips to mitigate the effects of fading channels (including the burst errors). Put differently, the interleaver 202 permutates bit streams to prevent all erroneous bits in a fading dip from ending up in the same codeword. This plays an important role in reducing bit error rate and hence improving transmission efficiency. It is worth to mention that “interleaving” does not change the information content and the code rate.
[0017] Moreover, interleavers may be used as a separate interleaver block 208, 209 in a baseband processing chain to improve the communication performance and provide a reliable data transmission. In this case, the bits within a bitstream are interleaved at the transmitter and then they will be rearranged again at the receiver side using the deinterleaver 205, which is basically the inverse interleaver. There are many applications like live video streaming and extended / augmented reality (XR, AR), which need a very high throughput. In order to achieve ever-increasing data rates, higher parallelism of channel coding schemes is required. On the other hand, the decoding is usually performed using iterative algorithms, which limits the decoder throughput due to the extra latency added by iterative processing.
[0018] An approach to increase the throughput and reduce the introduced latency by the iterative decoding is to use multiple sub-decoders in parallel to process the bitstreams concurrently. Although the use of parallel sub-decoders may increase the throughput significantly, there are critical problems in their hardware implementation. These problems are mainly related to the design of parallel interleavers, which are described in the next paragraphs.
[0019] Figure 3a illustrates a simplified architecture of a parallel decoder 300, which comprises multiple sub-decoders 301-1, 301-2, 301-3 ... 301 -S connected to several memories 302-1, 302-2, 302-3 ... 302-S using an interleaver and de-interleaver 303. All sub-decoders 301-1 , 301-2, 301-3 ... 301 -S operate concurrently, each of which receives a segment of a bitstream (codeword) from one or multiple memories and performs decoding. Then, the output stream of each sub-decoder will be written to one or multiple memories which may be same or different memories than the one or multiple memories from which the sub-decoders 301-1 , 301-2, 301-3 ... 301 -S receive the segment of the bitstream . Note that the input and output streams to the sub-decoders are permuted using the interleaver and de-interleaver 303. Consequently, in order to enable parallel decoding, a parallel interleaver and de-interleaver 303 is needed, which is an indispensable component for a parallel encoder or decoder.
[0020] Basically, parallelism increases the amount of data written to or read from the memories 302-1 , 302-2, 302-3 ... 302-S per clock cycle. The concurrent read and write operations result in memory conflicts or memory collisions in other words, which is mainly due to a permutation order of the interleavers 303. This means that in some time instances multiple sub-decoders 301-1, 301-2, 301-3 ... 301 -S want to write or to read from the same memory simultaneously, for example from memory 302-3 as indicated by the dashed arrows in Figure 3a. As a result, the degree of parallelism will be limited and hence an ultra-high throughput decoder is not feasible. Several methods have been proposed in the literature to address this issue and increase the level of parallelism, which may be categorized as follows. However, these methods have some limitations and drawbacks as described below.
[0021] • Contention-free algorithms like almost regular permutation (ARP) and quadratic permutation polynomial (QPP) may be used. See for example, R. Asghar, D. Wu, J. Eilert and D. Liu, "Memory Conflict Analysis and Interleaver Design for Parallel Turbo Decoding Supporting HSPA Evolution," 2009 12th Euromicro Conference on Digital System Design, Architectures, Methods and Tools, 2009, pp. 699-706, doi: 10.1109 / DSD.2009.178. However, in these algorithms, the level of parallelism is limited to specific interleaving patterns and certain design parameters, e.g., size of the codewords and the length of the interleaver. Moreover, these algorithms suffer from poor performance.
[0022] • Memory mapping. In this approach a contention-free memory access may be designed for any interleaving algorithm by finding the corresponding address mapping, which needs complex off-line computations. Then, all the address mapping scenarios should be stored in extra memories to be used for the desired interleaving pattern. This increases the hardware cost, due to larger area and power consumption, significantly. Thus, to reduce the amount of memory resources for storing the memory mappings, a limited number of interleaving patterns are supported.
[0023] • On-the-fly interleaver address generation is another method to tackle the memory conflicts. In this approach, a module called read / write conflict solver is designed to dynamically reschedule data access orders. This results in hardware overhead, due to extra buffers and control logic, as well as degraded throughput.
[0024] • Fully hard-wired interleaver is a solution in the literature to design a parallel decoder. In this method, the interleaver function is realized by hard-wire connections between sub-decoders and memories. Although, there is no memory collision, this approach is not flexible in terms of supported interleaver patterns and interleaver size.
[0025] Bit reversal is an algorithm that sorts a set of indexed data according to a reversing of the bits of the data index. This algorithm has been widely used to sort out the output samples of the FFT and IFFT. Figure 3b illustrates a processing chain of FFT / IFFT along with a corresponding sorting circuit. As shown in Figure 3b, an FFT / IFFT block 320 receives a sequence of K samples in natural order, where K is the length of FFT / IFFT. The output sequence of the FFT / IFFT block 320 includes K samples, which are in bit- reversed order or in any other non-natural order, depending on the selected FFT / IFFT algorithm. The FFT / IFFT block 320 is usually followed by other processing blocks which require natural-order input data. Thus, a sorting circuit 330 is usually needed to sort the output sequence of the FFT / IFFT block 320 as shown in Figure 3b.
[0026] To clarify the concept of bit-reversal algorithm, an example of 8-point FFT ( = 8) and the corresponding sorting circuit is mentioned in Table 1. An input sequence of FFT is entered to the FFT block 320 in natural order, -,x7, while the output samples of the FFT block 320 are generated in bit-reversed order, i.e., X0,X4, ...,X7. The corresponding decimal and binary values of the indices of input and output samples are represented in the left and middle parts of Table 1 , respectively. Finally, the sorted sequence is shown in the right side of T able 1 , where the output samples of the FFT block 320 are sorted in natural order.
[0027] Basically, the sorting circuit 330 sorts a sample with binary index of / n-1 / n-2 / n-3■■■I o t° a sample with the binary index of / 0 / i / 2...In-2In-1, where n = log2K. Note that the input / output samples of FFT, i.e., xtand Xt, are binary words.
[0028] Table 1. An example of sorting process for an 8-point FFT.
[0029] FFT / IFFT is used in various signal processing applications, such as wireless communication systems, and image and video signal processing. Moreover, high- throughput and real-time processing become essential demands in such applications. Thus, design of a high-throughput and parallel bit-reversal algorithm is necessary in many applications.
[0030] Figure 3c illustrates the processing flow of FFT and the corresponding sorting block, in which every input / output symbol includes K samples. The first output sample of FFT is generated after a certain number of clock cycles, which is called “Latency of FFT". Depending on the architecture of the FFT, more than one sample may be generated at each clock cycle. The output samples of FFT enter the sorting circuit and after the “Sorting Latency", the first samples of sorted sequence will be generated. In Figure 3c, the required time to sort all the K output samples of FFT in natural order, is called “Sorting Time”. As a result, the throughput of the whole processing chain, i.e., FFT and sorting, will be limited by the sorting time, which highly depends on the architecture of the sorting circuit.
[0031] Several designs have been proposed in the literature to reduce the sorting time for specific FFT architectures. For example, efficient memory addressing schemes have been presented to facilitate sorting process for memory-based FFT architectures. Also, for single-path pipelined FFT architectures the sorting may be performed using either doublebuffering strategy or using a single memory together with an address generator, which generates the memory addresses in natural and bit-reversed order, alternatively for even and odd sequences.
[0032] For parallel pipelined FFT architectures, design of the sorting circuit will be even more challenging since it should sort multiple concurrent FFT outputs simultaneously. The presented parallel sorting circuits in the literature suffer from a high hardware complexity and hardware cost, which is due to the large amount of required memory banks, extra buffers and multiplexers, and complex routing. Moreover, some of these designs need dual-port memories, such as disclosed in M. Garrido, J. Grajal and O. Gustafsson, "Optimum Circuits for Bit Reversal," in IEEE Transactions on Circuits and Systems II: Express Briefs, vol. 58, no. 10, pp. 657-661 , Oct. 2011 , doi: 10.1109 / TCSII .2011.2164141. This increases the hardware cost.
[0033] Another problem regarding the existing sorting solutions in the literatures is that some of them are only applicable to either specific order of FFT outputs or specific FFT algorithms / architectures. Also, the supported data parallelism is limited to certain values, which depend on the length of FFT, i.e., K. Moreover, the throughput of the sorting circuit will be constrained by the memory bandwidth, meaning that the number of samples per cycle which are received by the sorting circuit is limited to the memory bandwidth. Finally, the presented sorting schemes in the literature are not flexible in terms of radix of FFT computations. This means that a dedicated sorting circuit should be designed for a specific radix.
[0034] SUMMARY
[0035] There is thus a need for a more efficient approach for permuting an electronic input signal of multiple samples, such as for interleaving, deinterleaving or bit-reversal. The interleaving, deinterleaving, bit-reversal or re-ordering may for example be applied in digital communications.
[0036] An object of embodiments herein may be to obviate some of the problems related to interleaving, deinterleaving and bit-reversal mentioned above.
[0037] Embodiments herein disclose an electronic device for permuting an electronic input signal of multiple samples. The electronic device may be a fully parallel and flexible interleaver or de-interleaver or both, which is essentially needed to design a high- throughput decoder for a baseband processor.
[0038] The electronic device may also be a re-ordering circuit such as a bit reversal circuit, such as a fully parallel and flexible bit reversal circuit.
[0039] According to a first aspect, the object is achieved by a method, performed by an electronic device, for permuting an electronic input signal of multiple samples. The electronic device comprises an electronic crossbar array.
[0040] The method comprises applying the input signal of multiple samples to input conductors of the electronic crossbar array further comprising output conductors electrically connected to the input conductors via an array of electronic switches such that each output conductor is connected to a respective input conductor via a respective electronic switch configured according to a respective element of a permutation matrix.
[0041] The method further comprises reading an electronic output signal of multiple output samples, corresponding to the input signal of multiple samples, from the output conductors of the electronic crossbar array.
[0042] According to a second aspect, the object is achieved by an electronic device. The electronic device is configured to perform the method according to the first aspect above. According to a further aspect, the object is achieved by a computer program comprising instructions, which when executed by a processor, causes the processor to perform actions according to any of the aspects above.
[0043] According to a further aspect, the object is achieved by a carrier comprising the computer program of the aspect above, wherein the carrier is one of an electronic signal, an optical signal, an electromagnetic signal, a magnetic signal, an electric signal, a radio signal, a microwave signal, or a computer-readable storage medium.
[0044] Since the respective electronic switch of the electronic crossbar array is configured according to the respective element of the permutation matrix the electronic output signal will be permuted with respect to the electronic input signal.
[0045] Embodiments herein solve the problem of memory conflicts, which is a known problem in the context of for example interleaver design. As a result, the throughput of interleaver / de-interleaver is not limited since there will be no conflict in reading from or writing to memories anymore. This improves the throughput of decoders and consequently the throughput of baseband processors significantly since the latency of interleaving or de-interleaving methods disclosed by embodiments herein is only limited by the read cycle of the crossbar array, and it is not limited by memory collisions anymore.
[0046] In embodiments herein a permutation function, such as an interleaving function, may be converted to a matrix-vector multiplication, which may be done in a parallel manner. Thus, embodiments herein support fully parallel permutation of the input samples, such as fully parallel interleaving, de-interleaving and bit-reversal of the input samples. Embodiments herein further support partial-parallel and serial permutation of the input samples with the same hardware and with the same configuration of the crossbar array.
[0047] Embodiments herein support any permutation algorithm, such as any interleaving algorithm, and may realize any permutation pattern with the same hardware. As a result, a complex and sophisticated interleaving algorithm may be implemented to improve the performance of decoders significantly, which is important for ultra-reliable communication. To this end, a functionality of the targeted interleaving algorithm may be described via a permutation matrix and then it will be programmed to the crossbar array.
[0048] BRIEF DESCRIPTION OF THE DRAWINGS In the figures, features that appear in some embodiments are indicated by dashed lines.
[0049] The various aspects of embodiments disclosed herein, including particular features and advantages thereof, will be readily understood from the following detailed description and the accompanying drawings, in which:
[0050] Figure 1 illustrates a simplified wireless communication system,
[0051] Figure 2 is a block diagram schematically illustrating a data processing chain in a wireless communication system,
[0052] Figure 3a is a block diagram schematically illustrating a simplified architecture of a parallel decoder,
[0053] Figure 3b is a block diagram schematically illustrating a processing chain of FFT / IFFT along with a corresponding sorting circuit,
[0054] Figure 3c is a block diagram schematically illustrating a processing flow of FFT / IFFT and the corresponding sorting circuit,
[0055] Figure 4 is a block diagram schematically illustrating an electronic device according to some embodiments herein,
[0056] Figure 5 is a further block diagram schematically illustrating an electronic device according to some embodiments herein,
[0057] Figure 6 is a further block diagram schematically illustrating an electronic device according to some embodiments herein,
[0058] Figure 7 is a flowchart illustrating embodiments of a method for interleaving performed by an electronic device,
[0059] Figure 8 is a flowchart illustrating embodiments of a method of permutation performed by an electronic device comprising multiple crossbar arrays,
[0060] Figure 9a is a schematic block diagram illustrating embodiments of a permutation matrix and an electronic device,
[0061] Figure 9b is a schematic block diagram illustrating further embodiments of a permutation matrix and a further electronic device,
[0062] Figure 9c is a schematic block diagram illustrating further embodiments of a permutation matrix and an electronic device,
[0063] Figure 9d is a schematic block diagram illustrating further embodiments of a permutation matrix and a further electronic device,
[0064] Figure 9e is a schematic block diagram illustrating further embodiments of a permutation matrix and a further electronic device, Figure 10 is a flowchart illustrating embodiments of a method of permutation performed by an electronic device,
[0065] Figure 11 is a schematic block diagram illustrating further embodiments of a permutation matrix and a further electronic device,
[0066] Figure 12 is a schematic block diagram illustrating further embodiments of an electronic device,
[0067] Figure 13 is a schematic block diagram illustrating further embodiments of an electronic device,
[0068] Figure 14 is a flowchart illustrating embodiments of a method of interleaving or deinterleaving performed by an electronic device,
[0069] Figure 15 is a flowchart illustrating embodiments of a method of bit-reversal performed by an electronic device,
[0070] Figure 16 is a flowchart illustrating embodiments of a further method of bit-reversal performed by an electronic device,
[0071] Figure 17a is a schematic block diagram illustrating further embodiments of an electronic device,
[0072] Figure 17b is a schematic block diagram illustrating further embodiments of a bitreversal matrix,
[0073] Figure 18 is a flowchart illustrating embodiments of a further method of bit-reversal performed by an electronic device,
[0074] Figure 19 is a flowchart illustrating embodiments of a method of permutation performed by an electronic device,
[0075] Figure 20 is a schematic block diagram illustrating further embodiments of an electronic device.
[0076] Figure 21 is a block diagram schematically illustrating a wireless communication system.
[0077] DETAILED DESCRIPTION
[0078] Embodiments herein relate to electronic signal processing with electronic devices. More specifically, embodiments herein relate to methods and electronic devices for permuting electronic signals comprising multiple samples.
[0079] Figure 4 illustrates an electronic device 400 for permuting electronic signals.
[0080] For example, to realize a functionality of an interleaver, de-interleaver or both with a desired permutation pattern the electronic device 400 comprises an electronic crossbar array 410. A bit-reversal device may also be realized by the electronic device 400. The electronic crossbar array 410 comprises input conductors 401. The electronic crossbar array 410 further comprises output conductors 402 electrically connected to the input conductors 401 via an array of electronic switches such that each output conductor 402 is connected to a respective input conductor 401 via a respective electronic switch configured according to a respective element of a permutation matrix.
[0081] The crossbar array 410 enables performing massive multiply-accumulate (MAC) operations in parallel. More specifically, the crossbar array 410 computes matrix-vector multiplication (MVM) by calculating a dot-product of the input vector applied to crossbar rows (i.e. , word lines) and every column of the crossbar (i.e. , bit lines). The crossbar array 410 is a two-dimensional array that comprises an M x N array of electronic switches 411, 412, 421, 422, each of which may be programmed to represent an m-bit binary value. The electronic switches 411, 412, 421, 422 may be analog or digital switches. An analog switch may for example be implemented by a memristive device, also referred to as a memristor. In other words, the electronic switches 411, 412, 421, 422 of the crossbar array 410 may be memristors.
[0082] A memristor is a tunable resistor with memory. The memristor may essentially be a resistive switch. The memristor may comprise a dielectric layer sandwiched by two electrodes. A unique feature of memristors is that the conductance depends on historical electrical signals, making them capable of working as nonvolatile memory. In addition, memristors may store multibit information with continuously tunable conductance, in contrast to binary states “0” and “1” in traditional digital storage systems, equipping them with higher bit density. Thus, the m-bit binary value of the memristor may be set or programmed by applying a current to the memristor. The binary value may depend on the amplitude of the current.
[0083] Analog memristive devices have emerged as a new technology for storing and processing information in analog domain. These devices make it possible to perform computations in a place where data is stored. This concept is called in-memory computing, which eliminates the need for moving data from a memory to a processing unit.
[0084] There are different types of memristive devices, which are differentiated with respect to the used materials, switching principles, device endurance, retention, etc. The main types of memristive devices include phase change memory (PCM), resistive randomaccess memory (ReRAM), spin-transfer torque magnetic RAM (STT-MRAM), ferroelectric memristive devices (FeRAM). Memristive devices may support a limited bit precision, attributed to the limited number of conductance levels that may be reliably programmed in the device. For example, a PCM device may support around 50 conductance levels, meaning that it may represent around 6 bits.
[0085] A large number of such memristor devices may be organized to form an analog crossbar array. Thus, in some embodiments herein the crossbar array 410 is an analog crossbar array. Figure 5 illustrates an electronic device 500 comprising an analog crossbar array 510 that comprises M x N memristive devices 511, 512, 521, 522, each of which may be programmed to represent an m-bit binary value. The programming of the memristive devices 511 , 512, 521 , 522 may be performed by applying a current.
[0086] The analog crossbar array 510 further comprise parallel conductors, such as metal lines, termed word lines and bit lines, respectively, as electrodes of the memristors. The word lines and bit lines may be perpendicular to each other. The memristors are formed at the intersections of word and bit lines. In embodiments herein the input conductors 401 of the electronic device 400 corresponds to the word lines and the output conductors 402 of the electronic device 400 corresponds to the bit lines.
[0087] The analog crossbar array 510 computes MVM by calculating the dot-product of the input vector applied to crossbar rows (i.e., word lines) and every column of the crossbar (i.e., bit lines), all performed in analog domain using Ohm’s law for multiplication and Kirchhoff’s law for accumulation.
[0088] Thus, an M x N matrix of binary words, e.g., G, may be represented by the electronic crossbar array 410 and more specifically by the analog crossbar array 510. The input to the electronic crossbar array 410 is an electronic input signal of multiple samples, such as a vector of M binary values, e.g., V. When the analog crossbar array 510 is used and the electronic input signal is digital then the electronic device 500 further comprises one or more Digital to Analog Converters (DACs) 504 adapted to convert the input signal of multiple samples to corresponding analog voltages Vi, V2, ... VM. There may be one DAC 504 per input sample. In some other embodiments there may be less than one DAC 504 per input sample as one DAC 504 may be shared among several input samples by multiplexing. For example, two input samples may share the same DAC 504 as will be described in more detail below.
[0089] An output vector, e.g., / , is equal to the result of matrix-vector multiplication, i.e., I = G- V. This concept will be described in more detail below. In embodiments herein the matrix G is a permutation matrix which may perform for example interleaving, de-interleaving or bit-reversal of the elements of the input vector V. The above-mentioned property of the electronic crossbar array 410 may be utilized to realize a fully parallel and flexible interleaver or deinterleaver or both or a bit-reversal device.
[0090] The permutation matrix G is defined based on a targeted permutation pattern, such as an interleaving pattern. The entries of this permutation matrix are either 0 or 1, which may be programmed into the electronic switches 411 , 412, 421 , 422 of the electronic crossbar array 410. For example, the memristive devices 511, 512, 521, 522 of the analog crossbar array 510 may be programmed by the entries of the permutation matrix.
[0091] A respective electronic switch 411 of the electronic switches 411, 412, 421 , 422 is configured to conduct current when the corresponding matrix element of the permutation matrix is 1.
[0092] In some embodiments the input stream that consists of M binary words is first converted to the corresponding analog voltages and then they are applied to word lines (rows in Figure 4 and 5) of the analog crossbar array 510, e.g., at once. Output signals will be extracted from the bit lines (columns in Figure 4 and 5) of the crossbar array 410, 510. If the crossbar array 410, 510 is analog then the output values may be converted to digital values, which are equivalent to the interleaved version of the input sequence. In this way the interleaving function is realized in a fully parallel manner. Thus, the electronic device 500 may further comprise one or more Analog to Digital Converters (ADCs) 505 adapted to convert the output samples, comprising analog output current, to corresponding digital output values.
[0093] In embodiments herein in which the electronic crossbar array 410 is digital, e.g. when it comprises digital switches, there is no need to use DACs and ADCs. This will significantly reduce the hardware cost, especially in case of large crossbar arrays, such as for large interleavers. Also, there is no need to use DACs if the input signals are already analog and there is no need to use ADCs if the next block after the interleaver needs the analog signals as its input.
[0094] A corresponding de-interleaving may be easily realized using the same crossbar array 410, 510 without reprogramming. To this end, the input vector is applied to the crossbar columns instead of the crossbar rows. Consequently, the de-interleaved sequence will be extracted from the crossbar rows. This concept and also different implementation strategies for interleavers, de-interleavers and bit-reversal devices will be described in detail below.
[0095] In Figure 5 the entries of matrix G (an M x M matrix) are programmed to the memristive devices of an M x M crossbar array while the vector V (an M x 1 vector) is applied to the crossbar rows. Note that, the vector V corresponds to the actual input vector (Input 1, Input M), which are converted to analog voltages using the one or more DACs 504. As a result, the following MVM may be realized using the illustrated crossbar array 510: where the vector I is the output current of crossbar columns, which will be converted to the corresponding binary words using the one or more ADCs 505 as shown in Figure 5. This conversion may be done either separately for each crossbar column, i.e. , one ADC for each binary word, or in a time-multiplexed fashion and hence reduce ADC overhead i.e., multiple bit lines may share one ADC.
[0096] Equation Mapping
[0097] A functionality of the permutation, such as interleaving / de-interleaving, may be described using a permutation matrix, P, as follows. Let’s consider an interleaver of size K, where the size of a corresponding permutation matrix is K*K. A property of this matrix is that all the entries are either 0 or 1 , and each row of matrix P includes only one entry equal to 1 while the remaining entries are equal to zero. If the (i, j)-t entry of matrix P is 1 , the y-th entry of input vector will be permuted to the / -th position of the interleaved sequence. Thus, the functionality of the interleaver / de-interleaver may be mathematically expressed using the following equation mapping: where Vis the input vector of size K, and I is the interleaved sequence with the same size, i.e.,
[0098] I = P. V . (3)
[0099] This concept is clarified by the following example to interleave a vector with K = 6 entries:
[0100] 0 1 0 0 0 0 0 0 0 0 0 0
[0101] 1 0 0 0 0 0 0 1 -0 0 1 0
[0102] Figure 6 is a block diagram of the electronic device 400, 500 and of methods for permuting disclosed herein. A permutation algorithm is described by a permutation matrix, P.
[0103] Figure 7 illustrates a flowchart of a method of interleaving performed by the electronic device 400. In particular the method will be described as being performed by the electronic device 400, 500 comprising the analog crossbar array 510. However, although some of the actions of the method are only applicable for the analog crossbar array 510 the method may also be performed by the electronic device 400, 500 comprising a digital crossbar array.
[0104] Action 701
[0105] A permutation matrix (P) of size K*K is generated, or in other words created, according to the interleaving pattern.
[0106] Action 702
[0107] Given the size of the crossbar array, MxN, a number of sub-matrices per row and per column are calculated.
[0108] Action 703
[0109] Then a total number of required crossbar arrays to be programmed by P is equal to a = mxn = Ceil(K / M) x Ceil(K / N). Each of the a crossbar arrays may be considered as a tile within an mxn plane Action 704
[0110] If a = 1, the whole permutation matrix (P) will be programmed to one crossbar array, otherwise, it will be programmed into a crossbar arrays.
[0111] Action 705
[0112] The respective electronic switch 411, 412, 421, 422 may be configured such that electronic switches 411 , 412 of a k-th output conductor of the electronic crossbar array 410 are configured according to the elements of a k-th row of the permutation matrix of KxK elements, k=1 , ... , K.
[0113] Action 706
[0114] When a new input sequence is received, it may be converted to the corresponding analog voltages using the one or more DACs 504, i.e., each binary word in the input sequence may be converted to an analog voltage.
[0115] Action 707
[0116] These analog signals will be applied to the analog crossbar rows and after a read cycle the analog current signals will be extracted from the crossbar columns, which may be converted to the corresponding binary values using the one or more ADCs 505. The output words of the one or more ADCs 505 represent the interleaved version of the original sequence. This process may be repeated for further incoming sequences.
[0117] As long as the pattern and size of the interleaver are fixed, there is no need to reprogram the memristive devices of the crossbars. However, when a new interleaver size or pattern is requested, the corresponding permutation matrix may be generated and programmed to the crossbar arrays accordingly.
[0118] Sometimes the size of the permutation and consequently the size of the permutation matrix is large, such that it cannot be programmed to a single crossbar array. In these cases, the permutation matrix may be divided into several sub-matrices, which will be programmed to multiple crossbar arrays as detailed in a flowchart of Figure 8 and explained below.
[0119] Action 801 If the size of the permutation matrix ( x ) is larger than the size of the crossbar array the permutation matrix may be programmed to a = m*n crossbars, where m = Ceil(K / / W) and n = Ceil(K / / V) as mentioned above.
[0120] Action 802
[0121] Thus, the permutation matrix should be divided into a sub-matrices, i.e. , Gt , each of which has the size of N*M. To this end, the first sub-matrix, G1:1, is picked up from the upper left corner of the permutation matrix such that its upper left entry is P1 ±and its lower right entry is PNM. The next one, G12, is picked up from PliM+1to PNi2Mand so on.
[0122] Action 803a and 803b
[0123] As mentioned in Flowchart 3, if n #= K / N or m #= K / M the remaining entries of G^ will be considered as zero. The remaining entries are the entries in the last (nN-K) rows of the sub-matrices Gn,j_and the last (mM-K) columns of the sub-matrices Gj,m. As a result, a sub-matrices, G^, of size N*M are generated.
[0124] Action 804
[0125] The sub-matrices will be programmed into a crossbar arrays of the same size To this end, k-th row of G^ is mapped to k-th column of ( / , / )-th crossbar array.
[0126] To clarify this concept, an example of fully parallel interleaver of size K=20 is illustrated in Figure 9a, where it is assumed that the size of each crossbar array is 20x4. Thus, in Figure 9a a permutation matrix 900 is divided into a = 5 sub-matrices, i.e., G1:1, G2,I,G3,I .G4,I .G5,I . which are specified by different fill patterns in Figure 9. These submatrices will be programmed to the corresponding crossbar arrays, which are shown by corresponding fill patterns in Figure 9a.
[0127] Another example of mapping the permutation matrix of size 20x20 into multiple crossbars is illustrated in Figure 9b In this example, a size of each crossbar array is 10x10. Thus, the permutation matrix 900 is divided into a = 4 sub-matrices, i.e., G1;1, G12, G2;1,G2,2. which are specified by different patterns in Figure 9b. These sub-matrices will be programmed to corresponding crossbar arrays, which are shown by corresponding fill patterns in Figure 9b. For example, Figure 9b illustrates a first crossbar array 901, and a second crossbar array 902. Having considered the overall architecture in Figure 9b, the -th analog currentsignal is obtained by adding current-signals of multiple crossbar columns, which are programmed by the -th row of the permutation matrix. The accumulated current signal corresponds to -th word of the interleaved sequence. For example, in Figure 9b, the current signal for producing “Output 1” is obtained by adding I±and 1'^ which are generated by the first column of two different crossbars. The columns of the two crossbars that are used to represent the same row of the permutation matrix.
[0128] One way to add multiple current signals is to use an integrator 920, as shown in Figure 9b. The output signal of the integrator 920 will be sent to the one or more ADCs 505 to generate the binary word of the interleaved sequence.
[0129] Figure 9c illustrates an embodiment wherein the permutation matrix of size 20x20 may further be programmed into multiple crossbars which have different sizes. In Figure 9b the permutation matrix of size 20x20 is programmed into three crossbar arrays of size 20x4 and into one crossbar array of size 20x8.
[0130] Figure 9d illustrates a further embodiment wherein the permutation matrix may further be programmed into multiple crossbars which have different sizes. In Figure 9d the permutation matrix of size 20x20 is programmed into one crossbar array of size 10x20 and into two crossbar arrays of size 10x10.
[0131] It can be seen that each crossbar array in Figure 9a and Figure 9b has its own ADC modules, such that all the entries of interleaved sequence will be generated simultaneously. However, in order to reduce hardware cost, it is possible to share the ADC modules between all the crossbars as shown in Figure 9e. Thus, in case of M*N crossbars, N ADC modules may be used by multiple crossbars in a time-multiplexed manner using a multiplexer 930. In the example of Figure 9e four ADCs are shared between all crossbars.
[0132] Partial-parallel implementation
[0133] In some applications, the input sequence to the electronic device 400, 500 is not generated in one read cycle, i.e. , the full sequence to be permuted is generated through multiple read cycles. A read cycle corresponds to a time to produce the output signal of the electronic crossbar array 410 once the input signal has arrived at the electronic crossbar array 410.
[0134] A time interval between two successive input streams, is specified by a previous block which generates these inputs and therefore it is not dependent on the read cycle of the electronic crossbar array 410. For example, the time interval between two successive input streams may be more than one read cycle.
[0135] Moreover, to reduce the hardware cost of ADCs and DACs in case of an analog crossbar array, it is possible to lower the level of parallelism and share the DACs between multiple inputs and the ADCs between multiple outputs. In both of the above-mentioned cases, a partial parallel implementation of the electronic device 400, 500 is possible. A flowchart of such a partial parallel implementation of the electronic device 400, 500 comprising analog crossbars is illustrated in Figure 10 and described below.
[0136] Action 1001
[0137] As shown in Figure 10, a permutation matrix (P) of size K*K is generated according to the interleaving pattern as described above.
[0138] Action 1002
[0139] Given the size of the crossbar array, M*N, the number of required crossbar arrays to be programmed by P may be calculated as a = Ceil(K / / W) x Ceil(K / / V).
[0140] Action 1003a
[0141] If a = 1, the whole permutation matrix (P) will be programmed to one crossbar array.
[0142] Action 1003b
[0143] Otherwise, it will be programmed into a crossbar arrays.
[0144] Action 1004
[0145] Let’s consider that in every read cycle, Q binary words of the input sequence arrive. Thus, the whole input sequence is received after q = K / Q cycles. In each read cycle, the incoming Q binary words will be converted to the corresponding analog voltages using DAC modules.
[0146] Action 1005 Then, the outputs of ADC modules will be extracted after the “read cycle” of the crossbar array.
[0147] Action 1006
[0148] If the generated binary-word of -th ADC is not zero, it will be considered as the -th word of the interleaved sequence.
[0149] This process will be repeated for the remaining parts of the input sequence, q =
[0150] I , ... , K / Q as detailed in the flowchart of Figure 10.
[0151] Action 1007
[0152] The method may comprise checking whether or not a new interleaver size or a new interleaver pattern is to be used.
[0153] As long as the pattern and size of the interleaver are fixed, the above procedure will be performed for the next incoming sequences. However, when a new interleaver size or pattern is requested, the corresponding permutation matrix may be generated and programmed to the electronic switches 411 , 412, 421 , 422 of the crossbar arrays accordingly.
[0154] An example of partial parallel permuting device of size K=20 is illustrated in Figure
[0155] I I , where it is assumed that the input sequence to be interleaved is received in two read cycles. Thus, a required size of the crossbar array is 10x20, and the permutation matrix is divided into a = 2 sub-matrices, i.e., G1:1and G12, which are specified by different patterns in Figure 11. These sub-matrices will be programmed to corresponding further first and second crossbar arrays 1101, 1102, which are shown by corresponding patterns in Figure 11. In the following disclosure, the further first and second crossbar arrays 1101,
[0156] 1102 of Figure 11 will be referred to as just the first and second crossbar arrays 1101, 1102 as it will be clear from the context whether the first and second crossbar arrays refer to Figure 9b or Figure 11. Thus, G1:1will be programmed to the first crossbar array 1101, while G12will be programmed to the second crossbar array 1102.
[0157] The following description of Figure 11 assumes that the input signal to the electronic device 400, 500 is digital.
[0158] Each DAC 504 is shared between the two crossbars 1101 , 1102 since in every read cycle only one crossbar receives inputs from the DACs 504. This may be realized using multiplexers and de-multiplexers as shown in Figure 11. That is, in Figure 11 an input (e.g., Input 1) to the first crossbar array 1101 is multiplexed with a corresponding input (e.g., Input 11) to the second crossbar array 1102 by one or more input multiplexers 1111. A selector of the one or more input multiplexers 1111 controls which input is passed to the DAC 504. After the DAC 504 the digital input is routed to the correct crossbar array, e.g., by letting the signal pass through a demultiplexer 1112 comprising a selector that controls to which crossbar array the output is sent.
[0159] Moreover, one or more output multiplexers 1113 is used at the inputs of the one or more ADCs 505 to switch between the generated signals by different crossbars.
[0160] Thus, in some embodiments wherein the input signal is digital then the first crossbar array 1101 of alpha multiple crossbar arrays 1101, 1102 is adapted to receive a first set of K multiple samples of the digital input signal in a first read cycle of the alpha multiple crossbar arrays 1101 , 1102 and a second crossbar array 1102 of the alpha multiple crossbar arrays 1101 , 1102 is adapted to receive a second set of the K multiple samples of the digital input signal in a second read cycle of the alpha multiple crossbar arrays 1101, 1102 after the first set of the K multiple samples of the digital input signal has been received in the first read cycle.
[0161] According to Figure 11 the electronic device 500 may further comprise the one or more input multiplexers 1111 adapted to select the first set of the K multiple samples of the digital input signal or the second set of the K multiple samples of the digital input signal as input to the one or more DACs 504 adapted to convert the digital input signal of the corresponding analog voltages.
[0162] According to Figure 11 the electronic device 500 may further comprise the one or more de-multiplexers 1112 adapted to select a first set of multiple analog voltages from the one or more DACs 504, corresponding to the first set of the K multiple samples of the digital input signal, for distribution to the first crossbar array 1101 or a second set of multiple analog voltages from the one or more DACs 504, corresponding to the second set of the K multiple samples of the digital input signal for distribution to the second crossbar array 1102.
[0163] According to Figure 11 the electronic device 500 may further comprise the one or more output multiplexers 1113 adapted to select a first set of multiple analog currents from the first crossbar array 1101 , or a second set of multiple analog currents from the from the second crossbar array 1102 for distribution to the one or more ADCs 505.
[0164] The electronic device 500 selects a non-zero output from the one or more ADCs 505 unless both outputs from the one or more ADCs 505 are zero in which case the electronic device 500 may select any of the outputs. When both outputs of the same column in two crossbars are zero, it means that the corresponding input sample was “zero”. In such cases the output is also zero.
[0165] Bit-Serial Inputs
[0166] The supported bit resolution of ADCs determines the quantization error; the higher supported bit resolution the lower the quantization error and hence the better performance. Therefore, in case that either the ADCs do not support the required resolution or in order to improve the performance, each binary word of the input vector may be sent to the DACs in a bit-serial manner.
[0167] Figure 12 illustrates the electronic device 500 comprising the analog crossbar array 510 which performs the same operation as the one in Figure 5 while a bit-serial adapted method is employed for the input vector. At each time instance, which corresponds to the read cycle of the analog crossbar array 510, one bit of all input binary-words is applied to the corresponding DAC. In Figure 12, the / -th bit of y-th input and output words are shown by Vj and I- , respectively. Thus, after W read cycles, the MVM is completed where W is the number of bits per input binary-word.
[0168] For the bit-serial adapted method, “Shift and Add’’ circuits may be used to calculate the final results of each crossbar column. Thus, the electronic device 500 may further comprise one or more digital shift and add circuits 1240 connected to a respective output of the one or more ADCs 505.
[0169] In some embodiments the final results of each crossbar column may be obtained via time multiplexing and sharing the one or more digital shift and add circuits 1240 between multiple columns. However, if all the outputs are to be generated simultaneously in a parallel manner, corresponding to the received inputs, then the same amount of shift and add circuits as the amount of outputs is required.
[0170] De-interleaver design
[0171] Many applications, which include an interleaver, need a de-interleaver as well. Embodiments herein of the fully parallel and flexible interleaver may be further be employed to design a fully parallel and flexible de-interleaver as well. To this end, a corresponding permutation matrix, which describes a de-interleaver pattern may be created. Having considered the functionality of the interleaver and the de-interleaver, the de-interleaving matrix is the inverse of the interleaving matrix. Since every row and column of the interleaving matrix only has one entry equal to 1 , there is no need to do an explicit matrix inversion to obtain the de-interleaving matrix. More specifically, if the (i, y)-th entry of interleaving matrix is 1 , the y-th entry of the input vector will be permuted to the / - th position of the interleaved sequence. Thus, if (i, j)-t entry of interleaving matrix is 1 , the (j, / )-th entry of corresponding de-interleaving matrix will be equal to 1. As a result, the inverse of such a permutation matrix is equal to the transpose of that matrix, i.e., p i=pT
[0172] After creating the permutation matrix of de-interleaver, it will be programmed to one or multiple crossbars, e.g., following Figure 8, which depends on the size of the deinterleaver and the available crossbar arrays.
[0173] The embodiments disclosed herein relating to the interleaver is equally applicable to the de-interleaver. For example, similar to the interleaver design, it is possible to implement the de-interleaver in a fully parallel or partial parallel manner, following Figure 7 and Figure 10, respectively.
[0174] Implementation of the interleaver and the de-interleaver using the same hardware
[0175] In traditional designs, the interleaver and the de-interleaver are implemented separately in hardware. However, embodiments herein realize both interleaver and deinterleaver using the same hardware. Thus, some embodiments herein disclose an electronic interleaver or de-interleaver comprising the electronic device 400, 500 disclosed herein. Thus, depending on the configuration of the electronic device 400, 500 it may act as either an interleaver or a de-interleaver. In some embodiments the crossbar array 410, 510 may be reprogrammed to change from interleaver to de-interleaver and vice versa. However, some embodiments herein employ the above-described fact that the inverse of the permutation matrix is equal to the transpose of that matrix. By doing so there is no need for reprogramming of the crossbar array. Instead, the input and output may be redirected depending on a requested functionality of interleaver or de-interleaver.
[0176] Therefore, embodiments herein may determine the direction of the input and output paths, through which the input samples are sent to the crossbar and the output samples are received from the crossbar array 410, respectively.
[0177] More specifically, the crossbar array 410 is programmed with the permutation matrix of the interleaver. Then, if the interleaver function is requested, the input sequence is applied to the crossbar rows while the interleaved sequence is obtained from the crossbar columns. If the de-interleaver function is needed, the input sequence is applied to the crossbar columns while the de-interleaved sequence is received from the crossbar rows. The proposed scheme is depicted in Figure 13, wherein one or more demultiplexers 1301 is used to distribute the input sequence to the crossbar array rows or columns and one or more multiplexers 1302 is used to manage the inputs to the ADCs (i.e., the output from the crossbar array 410). As shown in Figure 13, by configuring selectors of the one or more multiplexers 1301 and the de-multiplexers to 0 and 1 , the functionality of interleaver and de-interleaver may be realized, respectively.
[0178] Thus, in some embodiments an input conductor for de-interleaving is an output conductor for interleaving and an output conductor for de-interleaving is an input conductor for interleaving. Then the electronic device 500 further comprises: a one or more de-multiplexers 1301 adapted to distribute a respective sample of the input signal to the input conductors 401 of the crossbar array 410 based on a requested functionality of interleaving or de-interleaving and b one or more multiplexers 1302 adapted to read a respective sample of the output signal from the output conductors 402 of the crossbar array 410 based on the requested functionality of interleaving or de-interleaving.
[0179] Figure 14 illustrates a flowchart, which describes how to realize a fully parallel and flexible interleaver and de-interleaver with the same hardware.
[0180] Action 1401
[0181] A permutation matrix (P) of size K*K is generated according to the interleaving pattern.
[0182] Action 1402
[0183] This permutation matrix will be programmed to a = Ceil(K / / W) x Ceil(K / / V) crossbar arrays, where the size of the crossbar array is M*N.
[0184] Action 1403a and 1403b
[0185] Depending on the implementation strategy, a fully parallel or partial parallel implementation may be realized following the flowchart of Figure 7 or the flowchart of Figure 10, respectively.
[0186] Action 1404a
[0187] If the interleaver function is requested, the input sequence will be applied to the crossbar rows. Action 1405a
[0188] Then, the output signals of the crossbar columns will be sent to the ADCs to generate the interleaved sequence after a crossbar read cycle.
[0189] Action 1404b
[0190] If the de-interleaver function is needed, the input sequence will be applied to the crossbar columns.
[0191] Action 1405b
[0192] The output signals of the crossbar rows will be sent to the ADCs to generate the deinterleaved sequence.
[0193] As long as the pattern and size of the interleaver are fixed, the above procedure may be performed for next incoming sequences. If a new interleaver size or pattern is requested, the corresponding permutation matrix may be generated and programmed to the switching devices of the crossbar arrays accordingly.
[0194] Bit-reversal device
[0195] Embodiments herein also disclose a bit-reversal device comprising the electronic device 400, 500.
[0196] The bit-reversal device is adapted to configure the electronic switches 411, 412 of a k-th column of the crossbar array 410 by elements of a k-th row of a bit-reversal matrix.
[0197] The bit reversal function may be mathematically represented using a KxK matrix, where K is the size of the FFT. The entries of the bit-reversal matrix are either 0 or 1. The bit-reversal matrix may be programmed into the electronic switches 411 , 412, 421 , 422 of the crossbar array 410. Then, FFT output samples (i.e. , K binary words) may be converted to the corresponding analog voltages using DACs, which will be applied to the crossbar rows. Finally, the generated signals from crossbar columns will be converted to the digital values using ADCs. In the following, it is demonstrated that the outputs of ADCs are equivalent to a sorted sequence of FFT outputs.
[0198] Embodiments herein enable to implement a fully parallel and flexible bit-reversal algorithm. This is essentially needed to achieve a high throughput in many applications, which include FFT processors. Moreover, the proposed scheme is applicable for any FFT algorithm, radix, FFT size, and FFT architecture. Also, it supports variable FFT sizes using the same crossbar array without reprogramming.
[0199] Embodiments herein convert the bit-reversal function to a bit-reversal matrix, R. This process is detailed in Figure 15 and described below for any -point FFT algorithm with radix 2n,n = 1, ..., log2K. Similarly, a corresponding procedure may be determined for the other FFT radices with different order of output samples.
[0200] Let’s consider an FFT algorithm of size K points, which produces the output samples in a bit-reversed order. The size of corresponding bit-reversal matrix ( / ?) is KxK, in which each column / row includes only one entry equal to 1 while the remaining entries are 0. Thus, in order to create such a matrix, the position of 1 entries should be identified.
[0201] Action 1501 (initialization)
[0202] Alpha = 1. n = 1. Alpha and n are parameters which are used to describe a bitreversal algorithm according to embodiments herein.
[0203] Action 1502
[0204] The indices of 1 entries will be stored in vector J, which is a 1x vector. All the elements of this vector are initialized by zero.
[0205] Action 1503
[0206] All entries of the bit-reversal matrix are initialized by zero.
[0207] Action 1504
[0208] The procedure is started form the first row of the bit-reversal matrix. The index of the 1 entry is calculated as detailed in Figure 15, n, 1=1 , and it will be stored as the first element of vector J.
[0209] Action 1505
[0210] This process will be repeated for the remaining rows of the bit-reversal matrix, i.e., n = 1, ..., log2K, as described in the flowchart of Figure 15 to find an index of next 1 entry, save it in vector J, and update the reordering matrix accordingly. As a result, a KxK bitreversal matrix, which includes K entries equal to 1 will be generated. This matrix will be used to perform bit-reversal operation for the FFT output samples. Parallel Bit-Reversal Scheme
[0211] Embodiments herein employ the electronic device 400, 500 comprising the crossbar array 410, 510 to realize a parallel bit-reversal algorithm for the FFT outputs. To this end, the functionality of bit-reversal algorithm is considered as the multiplication of the bitreversal matrix, R, and the vector of FFT output samples, V. This may be mathematically expressed as:
[0212] I = R. V, where V includes K output samples of FFT, which are in bit-reverse order, and I includes the same sample but in natural order. Detailed embodiments are disclosed in relation to a flowchart of Figure 16 and described below.
[0213] Action 1601
[0214] In order to realize the multiplication of the bit-reversal matrix, R, and the vector of FFT output samples, V, which is equivalent to bit-reversal operation, a KxK bit-reversal matrix R is generated according to the flowchart of Figure 15, where K is the FFT size ( output samples).
[0215] Action 1602
[0216] Then, the bit-reversal matrix will be programmed to the electronic switches 411 , 412, 421 , 422 of the crossbar array 410. To this end, the switches 411, 412 of -th column of the crossbar array are programmed by the elements of -th row of the bit-reversal matrix.
[0217] As shown in Figure 16, another input to the flowchart is a parallelism degree (P). This parameter represents the number of generated FFT output samples at each time instant, which is determined based on the FFT architecture.
[0218] Action 1603a
[0219] If all of the FFT outputs are generated at once (i.e., K = P), the fully-parallel realization of bit-reversal algorithm is needed, which is described in the left part of the flowchart of Figure 15. In this case, the FFT output samples will be sent to the DACs. The FFT output samples correspond to the input signal of the electronic device 400, 500 when it operates as a bit-reversal device.
[0220] Action 1604a
[0221] The DACs convert the FFT output samples to the corresponding analog voltages, i e., V1,V2. VK.
[0222] These analog voltages may be applied to the crossbar rows. Having considered the structure of the crossbar, the basic multiply-accumulate (MAC) operations in equation (4) are performed by using Ohm’s law and Krichhoff’s current law.
[0223] Action 1605a
[0224] Next, the ADCs convert the output signals of the crossbar columns to the corresponding binary words. These binary words represent the FFT output samples, which are now in natural order. The ADC outputs are usually valid after a read cycle of the crossbar array.
[0225] The above-mentioned process may be repeated for the next sequence of FFT output samples, i.e. , next K samples. As a result, the output samples of a -point FFT may be sorted in a fully parallel manner, which only takes one read cycle.
[0226] Supporting various Parallelism Degrees (P-parallel Architecture)
[0227] Embodiments herein may be used for any degree of parallelism (P). This means that the bit-reversal device may accept P FFT outputs per cycle, where P = 1, In the previous section the fully-parallel bit-reversal scheme was presented, where P = K.
[0228] Now, let’s consider that the FFT architecture generates P samples per cycle, where P < K, i.e., the full sequence is generated through K / P cycles. For such scenarios the proposed scheme works as detailed in the right part of the flowchart of Figure 16 and described below.
[0229] Moreover, sometimes the partial-parallel approach may be employed to reduce the hardware cost by reducing the number of DACs. This is supported by embodiments herein by lowering the level of parallelism (i.e., P).
[0230] Description of right part of flowchart of Figure 16
[0231] Action 1603b As shown in the right part of the flowchart of Figure 16, P FFT outputs are received in each cycle, which are converted into the corresponding analog voltages using the DACs.
[0232] Action 1604b
[0233] The first P voltages are applied to the first P crossbar rows.
[0234] Action 1605b
[0235] Then, the non-zero outputs of the ADCs will be picked up, which correspond to a part of the sorted sequence. For example, if the output of -th ADC is non-zero, this binary word may be considered as the -th sample of the sorted sequence, i.e., the sequence with natural order. In the next cycle, the second bunch of P voltages may be applied to the (P + l)-th to (2P)-th crossbar rows and the non-zero outputs of ADCs may be picked up.
[0236] This process will be repeated K / P times to receive all the FFT outputs and generate the natural-order sequence. If the number of non-zero samples is less than K, it means that the remaining FFT outputs are zero.
[0237] Figure 17a illustrates an example of 4-parallel (P = 4) bit-reversal device 1700, which is used to sort a sequence of K=20 samples. A corresponding 20x20 bit-reversal matrix 1701 is illustrated in Figure 17b. Each part of the bit-reversal matrix in Figure 17b is programmed into certain memristive devices of the crossbar array of the bit-reversal device 1700, which are specified by a corresponding pattern fill in Figure 17a. This embodiment demonstrates how the number of DACs and consequently the hardware cost may be reduced.
[0238] Note that each pattern in Figure 17a represents either a separate crossbar or a part of a large crossbar. Thus, depending on the size of the bit-reversal matrix, it may be programmed into one or multiple crossbars, following the pattern fill in Figure 17a.
[0239] Variable-length bit-reversal
[0240] An advantage of the bit-reversal device disclosed herein is that it is flexible in terms of the length of input sequence, i.e., the FFT size ( ). More specifically, if the crossbar array is programmed to perform sorting for a sequence of K samples, it may also be used to sort sequences with a smaller size (K), where K’< K. In cases where the sequence is smaller K’< K) there is no need to create a new bitreversal matrix, which means that reprogramming is not needed and the original bitreversal matrix of size KxK may be used. This concept is detailed in a flowchart of Figure 18. As shown in Figure 18, the inputs to the flowchart are the number of samples to be sorted, K, and the parallelism degree, P. The first step is to create the bit-reversal matrix of size K x K, following the flowchart of Figure 15. The bit-reversal matrix, R, is programmed to the switching devices of the crossbar array such that Zc-th row of R is programmed to the Zc-th column of crossbar.
[0241] When a new sequence of size K' arrives at the input, one of the following cases may happen:
[0242] 1. If the size of incoming sequence corresponds to the size of programmed bitreversal matrix, i.e. , K' = K, the current configuration of the crossbar may be kept. In this case, Zc-th sample of the input sequence is sent to the Zc-th crossbar row, and the sorting procedure will be performed following the flowchart of Figure 16.
[0243] 2. If the size of incoming sequence is less than K, i.e., K' < K, the current configuration of the crossbar may be kept. However, in this case ’-th sample of the input sequence is sent to the (— i + l)-th crossbar row, where k' = 1, ...,K' and i = 0, ...,K' - 1. Then, the sorting procedure will be performed following the flowchart of Figure 16.
[0244] 3. If the size of incoming sequence is larger than K, i.e., K' > K, the current configuration of the crossbar should be changed. Thus, a new bit-reversal matrix of size K' x K' will be created following the flowchart of Figure 15. Then, the new bit-reversal matrix may be programmed to the crossbar and the sorting may be performed as described in case 1.
[0245] Usually the FFT sizes, which are used in any application / system are known in advance. Thus, to avoid reprogramming in a certain application / system, the crossbar may be programmed with the largest required size, i.e., largest K to be supported in the target application / system. As a result, the above-mentioned case 3 will not happen and therefore the same hardware may be used to sort the incoming sequences with different sizes in the target application / system.
[0246] Large Bit-Reversal Matrix
[0247] Sometimes the size of the input sequence of bit-reversal circuit, K, is large such that the corresponding bit-reversal matrix cannot be programmed to a single crossbar array. In these cases, the bit-reversal matrix may be divided into multiple smaller matrices, which may be programmed to multiple crossbar arrays. An example of this concept is shown in Figure 17a, in which the bit-reversal matrix is divided into four parts and programmed into four smaller crossbar arrays.
[0248] It is even possible to map the bit-reversal matrix into multiple crossbars of different sizes to improve the hardware utilization. This concept is presented in Figure 9c and Figure 9d.
[0249] Bit-Serial Inputs
[0250] As explained above, in case that either the ADCs do not support the required resolution or in order to improve the performance, each input sample, which is a binary word, may be sent to the DACs in a bit-serial manner. Thus, at each time instance, which corresponds to the read cycle of the crossbar array, one bit of all input binary-words is applied to the corresponding DACs. Therefore, after W read cycles, the matrix-vector multiplication is completed where W is the number of bits per input binary-word. In order to calculate the final results of each crossbar column in the bit-serial scheme, multiple “Shift and Add” blocks may be used after the ADC modules.
[0251] Exemplifying methods according to embodiments herein will now be described with reference to a flow chart in Figure 19.
[0252] The flow chart of Figure 19 illustrates a method, performed by the electronic device 400, 500. As mentioned above, the electronic device 400, 500 comprises the electronic crossbar array 410, 510.
[0253] The method is for permuting an electronic input signal of multiple samples. The permutation may be fully parallel.
[0254] The method may be for interleaving or de-interleaving binary words.
[0255] In some other embodiments the method is for bit-reversal of the input signal of multiple samples. Then the permutation matrix is a bit-reversal matrix. The method for bit-reversal supports variable length of the input signal.
[0256] The method actions may be performed in any suitable order. Some method actions may be optional.
[0257] In some embodiments herein there are K multiple samples of the input signal. Then the number of input conductors is K and the number of output conductors is K. K is a natural number of two or more. The respective electronic switch 411, 412, 421, 422 is configured according to the respective element of the permutation matrix of KxK elements.
[0258] The respective electronic switch 411, 412, 421, 422 may be configured such that electronic switches 411 , 412 of a k-th output conductor of the electronic crossbar array 410 are configured according to the elements of a k-th row of the permutation matrix of KxK elements, k=1 , ... , K.
[0259] Action 1901
[0260] In some embodiments herein there are K multiple samples of the input signal and the number of input conductors is M and the number of output conductors is N. K is a natural number of two or more and K is larger than M or N. Then the method may further comprise calculating a number alpha of required crossbar arrays of size MxN for permuting the K multiple samples of the input signal according to alpha = Ceil(K / M)xCeil(K / N).
[0261] Action 1902
[0262] When K is larger than M or N then the method may further comprise generating the number alpha of sub-matrices which together represent the permutation matrix.
[0263] Action 1903
[0264] The method may further comprise configuring the respective electronic switch 411 , 412, 421 , 422 according to the respective element of the permutation matrix.
[0265] When K is larger than M or N then the method may further comprise configuring the required crossbar arrays with the alpha sub-matrices.
[0266] Action 1904
[0267] When the electronic switches 411, 412, 421, 422 of the electronic crossbar array 410 are memristors, then the method further comprises converting the multiple samples of the input signal, applied to the input conductors 401, to corresponding analog input voltages and converting the output samples, comprising analog output current, to corresponding digital output values.
[0268] Action 1905 The method comprises applying the input signal of multiple samples to the input conductors 401 of the electronic crossbar array 410 further comprising the output conductors 402 electrically connected to the input conductors 401 via the array of electronic switches 411, 412, 421, 422 such that each output conductor 402 is connected to a respective input conductor 401 via a respective electronic switch 411 , 412, 421 , 422 configured according to a respective element of the permutation matrix.
[0269] When a number of the multiple samples of the input signal K’ is less than the number of input conductors K then applying the input signal of multiple samples comprises applying a k’-th sample of the input signal to the — i + 1-th row of the crossbar array 410,
[0270] Action 1906
[0271] The method may further comprise reading an electronic output signal of multiple output samples, corresponding to the input signal of multiple samples, from the output conductors 402 of the electronic crossbar array 410.
[0272] As explained in detail above, the output signal is the permuted signal which is formed by matrix-vector multiplication of the input signal and the values of the crossbar array 410.
[0273] In case the electronic device 500 performs partial-parallel permutation then the electronic device 500 may select a non-zero output from the one or more ADCs 505 unless both outputs from the one or more ADCs 505 are zero in which case the electronic device 500 is adapted to select any of the outputs.
[0274] Figure 20 illustrates further optional details of the electronic device 400, 500. The electronic device 400, 500 is configured to perform the method actions of Figure 19 above.
[0275] Thus, the electronic device 400, 500 is configured for permuting an electronic input signal of multiple samples.
[0276] The electronic device 400, 500 is further configured to apply the input signal of multiple samples to input conductors 401 of the electronic crossbar array 410 further comprising output conductors 402 electrically connected to the input conductors 401 via an array of electronic switches 411 , 412, 421 , 422 such that each output conductor 402 is connected to a respective input conductor 401 via a respective electronic switch 411, 412, 421 , 422 configured according to a respective element of a permutation matrix. The electronic device 400, 500 is further configured to read an electronic output signal of multiple output samples, corresponding to the input signal of multiple samples, from the output conductors 402 of the electronic crossbar array 410.
[0277] In some embodiments herein the electronic crossbar array 410 comprises a number alpha multiple crossbar arrays 1201 , 1202. Then the permutation matrix is represented by alpha sub-matrices and the alpha multiple crossbar arrays 1201, 1202 are configured with the alpha sub-matrices.
[0278] In some embodiments herein the alpha multiple crossbar arrays 1201, 1202 each are of size MxN and there are K multiple samples of the input signal and the number of input conductors 401 is M and the number of output conductors 402 is N and K is a natural number of two or more and K is larger than M or N. Then alpha = Ceil(K / M)xCeil(K / N). That is the number of multiple crossbar arrays may be calculated as Ceil(K / M)xCeil(K / N).
[0279] In some other embodiments herein at least two of the alpha multiple crossbar arrays 1201, 1202 are of different sizes.
[0280] In some embodiments herein the first crossbar array 901 of the alpha multiple crossbar arrays 901 , 902 is adapted to receive a first set of the K multiple samples of the input signal and the second crossbar array 902 of the alpha multiple crossbar arrays 901 , 902 is adapted to receive a second set of the K multiple samples of the input signal. Then the electronic device 400, 500 may further comprise one or more integrators 920 to add the output samples, comprising the analog output current, from corresponding output conductors of the respective first and second crossbar array 901 , 902.
[0281] Thus, in some embodiments wherein the input signal is digital then the first crossbar array 1101 of alpha multiple crossbar arrays 1101, 1102 is adapted to receive a first set of K multiple samples of the digital input signal in a first read cycle of the alpha multiple crossbar arrays 1101, 1102 and a second crossbar array 1102 of the alpha multiple crossbar arrays 1101 , 1102 is adapted to receive a second set of the K multiple samples of the digital input signal in a second read cycle of the alpha multiple crossbar arrays 1101 , 1102 after the first set of the K multiple samples of the digital input signal has been received in the first read cycle.
[0282] As mentioned above, according to Figure 11 the electronic device 500 may further comprise the one or more input multiplexers 1111 adapted to select the first set of the K multiple samples of the digital input signal or the second set of the K multiple samples of the digital input signal as input to the one or more DACs 504 adapted to convert the digital input signal of the corresponding analog voltages.
[0283] According to Figure 11 the electronic device 500 may further comprise the one or more de-multiplexers 1112 adapted to select a first set of multiple analog voltages from the one or more DACs 504, corresponding to the first set of the K multiple samples of the digital input signal, for distribution to the first crossbar array 1101 or a second set of multiple analog voltages from the one or more DACs 504, corresponding to the second set of the K multiple samples of the digital input signal for distribution to the second crossbar array 1102.
[0284] According to Figure 11 the electronic device 500 may further comprise the one or more output multiplexers 1113 adapted to select a first set of multiple analog currents from the first crossbar array 1101 , or a second set of multiple analog currents from the from the second crossbar array 1102 for distribution to the one or more ADCs 505.
[0285] The electronic device 500 may be further adapted to select a non-zero output from the one or more ADCs 505 unless both outputs from the one or more ADCs 505 are zero in which case the electronic device 500 is adapted to select any of the outputs.
[0286] The embodiments herein may be implemented through a processor or one or more processors, such as the processor 1004, of a processing circuitry in the electronic device 400, 500, and depicted in Figure 20 together with computer program code for performing the functions and actions of the embodiments herein. The program code mentioned above may also be provided as a computer program product, for instance in the form of a data carrier carrying computer program code for performing the embodiments herein when being loaded into the electronic device 400, 500. One such carrier may be in the form of a CD ROM disc. It is however feasible with other data carriers such as a memory stick. The computer program code may furthermore be provided as pure program code on a server and downloaded to the electronic device 400, 500.
[0287] The electronic device 400, 500 may further comprise a memory 1002 comprising one or more memory units. The memory comprises instructions executable by the processor in the electronic device 400, 500. The respective memory 1002 is arranged to be used to store e.g. information, data, configurations, and applications to perform the methods herein when being executed in the electronic device 400, 500.
[0288] In some embodiments, a computer program 1003 comprises instructions, which when executed by the at least one processor, cause the at least one processor of the electronic device 400, 500 to perform the actions above.
[0289] In some embodiments, a carrier 1005 comprises the computer program, wherein the carrier is one of an electronic signal, an optical signal, an electromagnetic signal, a magnetic signal, an electric signal, a radio signal, a microwave signal, or a computer- readable storage medium.
[0290] The electronic device 400, 500 may be comprised in a node of a wireless communications network. For example, the electronic device 400, 500 may be comprised in an access node such as a radio access node. In some other examples the electronic device 400, 500 may be comprised in a wireless communications device. As mentioned above, there are other applications as well. For example, the electronic device 400, 500 may be comprised in an electronic device for image and video signal processing.
[0291] Figure 21 illustrates a wireless communications network 100 in which embodiments herein may be implemented.
[0292] The wireless communications network 100 may use a number of different technologies, such as Wi-Fi, Long Term Evolution (LTE), LTE-Advanced, 5G, New Radio (NR), Wideband Code Division Multiple Access (WCDMA), Global System for Mobile communications / enhanced Data rate for GSM Evolution (GSM / EDGE), Worldwide Interoperability for Microwave Access (WiMax), or Ultra Mobile Broadband (UMB), just to mention a few possible implementations. Embodiments herein relate to recent technology trends that are of particular interest in a 5G context. However, embodiments are also applicable in further development of other existing wireless communication systems such as e.g. WCDMA and LTE and in future wireless communication systems, such as 6G systems.
[0293] Access nodes operate in the wireless communications network 100 such as a radio access node 111. The radio access node 111 provides radio coverage over a geographical area, a service area referred to as a cell 115, which may also be referred to as a beam or a beam group of a first radio access technology (RAT), such as 5G, LTE, Wi-Fi or similar. There may be more than one cell. For example, there may be a second cell 116 as well. The radio access node 111 may be a NR-RAN node, transmission and reception point e.g. a base station, a radio access node such as a Wireless Local Area Network (WLAN) access point or an Access Point Station (AP STA), an access controller, a base station, e.g. a radio base station such as a NodeB, an evolved Node B (eNB, eNode B), a gNB, a base transceiver station, a radio remote unit, an Access Point Base Station, a base station router, a transmission arrangement of a radio base station, a stand-alone access point or any other network unit capable of communicating with a wireless device within the service area depending e.g. on the radio access technology and terminology used. The respective radio access node 111 may be referred to as a serving radio access node and communicates with a UE with Downlink (DL) transmissions to the UE and Uplink (UL) transmissions from the UE.
[0294] A number of wireless communications devices operate in the wireless communication network 100, such as a wireless communications device 121.
[0295] The wireless communications device 121 may be a mobile station, a non-access point (non-AP) STA, a STA, a user equipment and / or a wireless terminal, that communicate via one or more Access Networks (AN), e.g. RAN, e.g. via the radio access node 111 to one or more core networks (CN) e.g. comprising a CN node 130, for example comprising an Access Management Function (AMF). It should be understood by the skilled in the art that “UE” is a non-limiting term which means any terminal, wireless communication terminal, user equipment, Machine Type Communication (MTC) device, Device to Device (D2D) terminal, or node e.g. smart phone, laptop, mobile phone, sensor, relay, mobile tablets or even a small base station communicating within a cell.
[0296] Those skilled in the art will also appreciate that the units described above may refer to a combination of analog and digital circuits, and / or one or more processors configured with software and / or firmware, e.g. stored in the electronic device 400, 500, that when executed by the respective one or more processors such as the processors described above. One or more of these processors, as well as the other digital hardware, may be included in a single Application-Specific Integrated Circuitry (ASIC), or several processors and various digital hardware may be distributed among several separate components, whether individually packaged or assembled into a system-on-a-chip (SoC). When using the word "comprise" or “comprising” it shall be interpreted as nonlimiting, i.e. meaning "consist at least of".
[0297] The embodiments herein are not limited to the above-described preferred embodiments. Various alternatives, modifications and equivalents may be used.
Claims
CLAIMS1. A method, performed by an electronic device (400, 500), for permuting an electronic input signal of multiple samples, wherein the electronic device (400, 500) comprises an electronic crossbar array (410, 510) and wherein the method comprises: applying (1905) the input signal of multiple samples to input conductors (401) of the electronic crossbar array (410, 510) further comprising output conductors (402) electrically connected to the input conductors (401) via an array of electronic switches (411 , 412, 421 , 422) such that each output conductor (402) is connected to a respective input conductor (401) via a respective electronic switch (411, 412, 421, 422) configured according to a respective element of a permutation matrix; and reading (1906) an electronic output signal of multiple output samples, corresponding to the input signal of multiple samples, from the output conductors (402) of the electronic crossbar array (410, 510).
2. The method according to claim 1, wherein the respective electronic switch (411, 412, 421 , 422) of the electronic switches (411, 412, 421 , 422) is configured to conduct current when the corresponding matrix element of the permutation matrix is 1.
3. The method according to claim 1 or 2, wherein the electronic switches (411, 412, 421, 422) of the electronic crossbar array (410, 510) are memristors, the method further comprising: converting (1904) the multiple samples of the input signal, applied to the input conductors (401), to corresponding analog input voltages; and converting the output samples, comprising analog output current, to corresponding digital output values.
4. The method according to any of the claims 1-3, wherein there are K multiple samples of the input signal and wherein the number of input conductors is K and the number of output conductors is K, wherein K is a natural number of two or more, wherein the respective electronic switch (411 , 412, 421 , 422) is configured according to the respective element of the permutation matrix of KxK elements.
5. The method according to claim 4, wherein the respective electronic switch (411, 412, 421 , 422) is configured such that electronic switches (411 , 412) of a k-th output conductor of the electronic crossbar array (410, 510) are configured according to the elements of a k-th row of the permutation matrix of KxK elements, k=1 , ... , K.
6. The method according to any of the claims 1-3, wherein there are K multiple samples of the input signal and wherein the number of input conductors is M and the number of output conductors is N and wherein K is a natural number of two or more and K is larger than M or N, the method further comprising: calculating (1901) a number alpha of required crossbar arrays of size MxN for permuting the K multiple samples of the input signal according to alpha = Ceil(K / M)xCeil(K / N); generating (1902) a number alpha of sub-matrices which together represent the permutation matrix; and configuring (1903) the required crossbar arrays with the alpha sub-matrices.
7. The method according to any of the claims 1-6, further comprising: configuring (1903) the respective electronic switch (411 , 412, 421 , 422) according to the respective element of the permutation matrix.
8. The method according to any of the claims 1-7, for interleaving or de-interleaving binary words.
9. The method according to any of the claims 1-7, for bit-reversal of the input signal of multiple samples, wherein the permutation matrix is a bit-reversal matrix.
10. The method according to any of the claims 1-8, wherein a number of the multiple samples of the input signal K’ is less than the number of input conductors K and wherein applying (1905) the input signal of multiple samples comprises applying a k’- th sample of the input signal to the (— Kf i + l)-th row of the crossbar array (410), where11. An electronic device (400, 500) for permuting an electronic input signal of multiple samples, wherein the electronic device (400, 500) comprises an electronic crossbar array (410, 510) and wherein the electronic device (400, 500) is adapted to: apply the input signal of multiple samples to input conductors (401) of the electronic crossbar array (410, 510) further comprising output conductors (402) electrically connected to the input conductors (401) via an array of electronic switches (411 , 412, 421 , 422) such that each output conductor (402) is connected to arespective input conductor (401) via a respective electronic switch (411, 412, 421 , 422) configured according to a respective element of a permutation matrix; and read an electronic output signal of multiple output samples, corresponding to the input signal of multiple samples, from the output conductors (402) of the electronic crossbar array (410, 510).
12. The electronic device (400, 500) according to claim 11, wherein the electronic switches (411 , 412, 421 , 422) of the crossbar array (410) are memristors.
13. The electronic device (400, 500) according to any of the claims 11-12, the electronic crossbar array (410, 510) comprising a number alpha multiple crossbar arrays (1201, 1202), wherein the permutation matrix is represented by alpha sub-matrices; and wherein the alpha multiple crossbar arrays (1201, 1202) are configured with the alpha sub-matrices.
14. The electronic device (400, 500) according to claim 13, the alpha multiple crossbar arrays (1201 , 1202) each are of size MxN, wherein there are K multiple samples of the input signal and wherein the number of input conductors (401) is M and the number of output conductors (402) is N and wherein K is a natural number of two or more and K is larger than M or N, wherein alpha = Ceil(K / M)xCeil(K / N).
15. The electronic device (400, 500) according to claim 13, wherein at least two of the alpha multiple crossbar arrays (1201 , 1202) are of different sizes.
16. The electronic device (400, 500) according to any of the claims 11-15, further comprising one or more DACs (604) adapted to convert the input signal of multiple samples to corresponding analog voltages.
17. The electronic device (400, 500) according to any of the claims 11-16, further comprising one or more ADCs (505) adapted to convert the output samples, comprising analog output current, to corresponding digital output values.
18. The electronic device (1400) according to any of the claims 13-15 or claims 16-17 when dependent on any of the claims 13-15, wherein a first crossbar array (901) ofthe alpha multiple crossbar arrays (901 , 902) is adapted to receive a first set of the K multiple samples of the input signal and wherein a second crossbar array (902) of the alpha multiple crossbar arrays (901, 902) is adapted to receive a second set of the K multiple samples of the input signal, wherein the electronic device (400, 500) further comprises one or more integrators (920) to add the output samples, comprising the analog output current, from corresponding output conductors of the respective first and second crossbar array (901 , 902).
19. The electronic device (400, 500) according to any of the claims 13-15, wherein the input signal is digital and wherein a first crossbar array (1101 ) of the alpha multiple crossbar arrays (1101, 1102) is adapted to receive a first set of the K multiple samples of the digital input signal in a first read cycle of the alpha multiple crossbar arrays (1101 , 1102) and wherein a second crossbar array (1102) of the alpha multiple crossbar arrays (1101, 1102) is adapted to receive a second set of the K multiple samples of the digital input signal in a second read cycle of the alpha multiple crossbar arrays (1101, 1102) after the first set of the K multiple samples of the digital input signal has been received in the first read cycle, wherein the electronic device (400, 500) further comprises: one or more input multiplexers (1111) adapted to select the first set of the K multiple samples of the digital input signal or the second set of the K multiple samples of the digital input signal as input to the one or more DACs adapted to convert the digital input signal of the corresponding analog voltages; one or more de-multiplexers (1112) adapted to select a first set of multiple analog voltages from the one or more DACs, corresponding to the first set of the K multiple samples of the digital input signal, for distribution to the first crossbar array (1101) or a second set of multiple analog voltages from the one or more DACs, corresponding to the second set of the K multiple samples of the digital input signal for distribution to the second crossbar array (1102); and one or more output multiplexers (1113) adapted to select a first set of multiple analog currents from the first crossbar array (1101), or a second set of multiple analog currents from the from the second crossbar array (1102) for distribution to the one or more ADCs.
20. The electronic device (400, 500) according to claim 19, further adapted to select a non-zero output from the one or more ADCs (505) unless both outputs from the one ormore ADCs (505) are zero in which case the electronic device (400, 500) is adapted to select any of the outputs.
21. The electronic device (400, 500) according to any of the claims 17-20, further comprising one or more digital shift and add circuits (1240) connected to a respective output of the one or more ADCs (505).
22. An electronic interleaver or de-interleaver comprising the electronic device (400, 500) according to any of the claims 11-21.
23. The electronic interleaver and de-interleaver according to claim 22, wherein an input conductor for de-interleaving is an output conductor for interleaving and wherein an output conductor for de-interleaving is an input conductor for interleaving, the electronic device (400, 500) further comprising: a) one or more de-multiplexers (1301) adapted to distribute a respective sample of the input signal to the input conductors (401) of the crossbar array (410) based on a requested functionality of interleaving or de-interleaving and b) one or more multiplexers (1302) adapted to read a respective sample of the output signal from the output conductors (402) of the crossbar array (410) based on the requested functionality of interleaving or de-interleaving.
24. An electronic bit-reversal device comprising the electronic device (400, 500) according to any of the claims 11-21, wherein the bit-reversal device is adapted to: configure the electronic switches (411 , 412) of a k-th column of the crossbar array (410) by elements of a k-th row of a bit-reversal matrix.
25. A computer program (1003), comprising computer readable code units which when executed on an electronic device (400, 500) causes the electronic device (400, 500) to perform the method according to any one of claims 1-10.
26. A carrier (1005) comprising the computer program according to the preceding claim, wherein the carrier (1005) is one of an electronic signal, an optical signal, a radio signal and a computer readable medium.