Computing device in resistive memory
By adopting direct-drive successive approximation analog-to-digital converter technology, the problem of high amplifier overhead in the traditional ACIM core is solved, achieving more efficient computing power and more flexible operating modes, and improving the performance of calculating the resistor matrix in analog memory.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NOKIA NETWORKS OY
- Filing Date
- 2025-10-27
- Publication Date
- 2026-04-28
AI Technical Summary
Traditional von Neumann architectures are inefficient in compute-intensive applications, and analog-in-memory computing accelerators (ACIMs) suffer from problems such as high amplifier overhead power, high noise, limited output voltage swing, and limited bandwidth, which restrict their operating modes and scalability.
By employing direct-drive successive approximation analog-to-digital converter (SAR ADC) technology, and utilizing a self-biased circuit system and a direct-drive ADC, performance-limited amplifiers can be replaced, enabling flexible expansion and reconfigurability of the calculated resistor matrix in analog memory.
The performance of the ACIM core has been improved, enabling more efficient computing power and greater operational flexibility, while reducing power consumption and noise, and expanding the applicability of the resistor matrix.
Smart Images

Figure CN121938428A_ABST
Abstract
Description
Technical Field
[0001] The following example embodiments relate to information technology and computing technology. Background Technology
[0002] Information technology is having an increasingly significant impact on our society. Modern IT applications, from wireless networks to computer vision and large language models, require ever-increasing computing power. For many computationally intensive applications, the traditional von Neumann architecture can be replaced by more efficient application-specific computing solutions, such as ACIM (Analog In-Memory Computation) accelerators. ACIM accelerators can be deployed for, for example, computing matrix operations such as vector-matrix multiplication and VMM (Video Memory Model), or for implementing machine learning accelerators independently or in conjunction with other accelerators, such as digital in-memory computing accelerators. Summary of the Invention
[0003] The scope of protection sought by the various example embodiments is defined by the claims. Example embodiments and features (if any) described in this specification that are not within the scope of the claims should be interpreted as examples helpful in understanding the various embodiments.
[0004] The independent claims are provided with respect to several aspects. Several other aspects are defined in the dependent claims.
[0005] According to one aspect, an apparatus is provided, the apparatus comprising at least one processor and at least one memory storing instructions, the instructions, when executed by the at least one processor, causing the apparatus to at least: provide an analog signal to at least one input to an analog memory included in the apparatus for calculating a resistance matrix; and receive a digital signal from at least one output of the at least one analog memory for calculating the resistance matrix.
[0006] According to one embodiment, the provided analog signal is converted from a digital signal by at least one directly driven digital-to-analog converter from the input of the at least one analog memory that calculates the resistance matrix.
[0007] According to one embodiment, the calculation of the resistance matrix in at least one analog memory is implemented using at least one resistance memory, which includes at least one resistive element within the at least one resistance memory.
[0008] According to one embodiment, the calculation of the resistance matrix in at least one analog memory is implemented using at least one memory element and at least one resistive element.
[0009] According to one embodiment, at least one direct-drive analog-to-digital converter is a direct-drive successive approximation analog-to-digital converter.
[0010] According to one embodiment, at least one direct-drive successive approximation analog-to-digital converter (DADC) has a self-biasing circuit system comprising: a comparator in the input of the DADC, the comparator being arranged to compare a reference node voltage with an input node voltage from an output of an intrinsic digital-to-analog converter (ADC), the intrinsic ADC receiving an output signal of the DADC having reverse polarity as an input signal.
[0011] According to one embodiment, at least one direct-drive successive approximation analog-to-digital converter is arranged to provide the comparator's output signal to at least one other digital-to-analog converter via synchronous successive approximation logic of the at least one direct-drive successive approximation analog-to-digital converter.
[0012] According to one embodiment, at least one direct-drive successive approximation analog-to-digital converter is arranged to provide logic block output signals to at least one other digital-to-analog converter.
[0013] According to one embodiment, at least one direct-drive successive approximation analog-to-digital converter is arranged to provide the output signal of the intrinsic digital-to-analog converter to at least one load.
[0014] According to one embodiment, the apparatus includes an activation function between the output of the synchronous successive approximation logic and the input of the intrinsic digital-to-analog converter, or between the output of the synchronous successive approximation logic and the input of the at least one other digital-to-analog converter.
[0015] According to one embodiment, in this device, at least one digital output of a direct-drive analog-to-digital converter from the output of a calculation resistance matrix in a first analog memory is connected to at least one digital input of a direct-drive digital-to-analog converter from the input of a calculation resistance matrix in a second analog memory.
[0016] According to one embodiment, in this device, the output columns of the resistance matrix calculated in the first analog memory and the output columns of the resistance matrix calculated in the second analog memory are combined and connected to at least one direct-drive analog-to-digital converter in the output of the resistance matrix calculated in the second analog memory.
[0017] According to one embodiment, in this device, the output of the at least one intrinsic digital-to-analog converter that directly drives the successive approximation analog-to-digital converter is connected back to the input row of the at least one analog memory for calculating the resistance matrix, for matrix inverse calculation.
[0018] According to one embodiment, in this apparatus, the output of the at least one intrinsic digital-to-analog converter that directly drives the successive approximation analog-to-digital converter is connected back to the input row of the at least one analog memory for calculating the resistance matrix via at least one comparator, for use in Ising machine calculations.
[0019] According to one embodiment, the device is an electronic device, a user device, a node element, a base station, an access network component, a served device, a downlink device, a mobile device, a terminal device, a communication device, a user equipment, a subscriber station, a portable subscriber station, a mobile station, or an access terminal.
[0020] According to one aspect, a non-transitory computer-readable medium is provided, the medium comprising program instructions that, when executed by a device, cause the device to perform at least the following: providing an analog signal to at least one input to an analog memory included in the device for calculating a resistance matrix; and directly driving an analog-to-digital converter to receive a digital signal from at least one output of the at least one analog memory for calculating the resistance matrix.
[0021] According to one aspect, a user device / user equipment is provided, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the device to at least: provide an analog signal to at least one input to an analog memory included in the device for calculating a resistance matrix; and receive a digital signal from at least one output of the at least one analog memory for calculating the resistance matrix.
[0022] According to one aspect, a node element is provided, the node element comprising: at least one processor; and at least one memory storing instructions, which, when executed by the at least one processor, cause the device to at least: provide an analog signal to at least one input to an analog memory included in the device for calculating a resistance matrix; and receive a digital signal from at least one output of the at least one analog memory for calculating the resistance matrix.
[0023] According to one aspect, a base station is provided, the base station comprising: at least one processor; and at least one memory storing instructions, which, when executed by the at least one processor, cause the means to at least: provide an analog signal to at least one input to an analog memory included in the means for calculating a resistance matrix; and receive a digital signal from at least one output of the at least one analog memory for calculating the resistance matrix. Attached Figure Description
[0024] In the following description, various exemplary embodiments will be described in more detail with reference to the accompanying drawings, in which:
[0025] Figure 1 The illustration shows an example of an ACIM core based on a programmable resistor matrix;
[0026] Figure 2An example of an extended ACIM core based on a programmable resistor matrix is illustrated.
[0027] Figure 3 Another example of an extended ACIM core based on a programmable resistor matrix is illustrated.
[0028] Figure 4 The illustration shows an example of a conventional successive approximation ADC with varying input voltage levels;
[0029] Figure 5 An example of a wireless communication network is illustrated;
[0030] Figure 6 The illustration shows an example embodiment of a direct-drive successive approximation analog-to-digital converter;
[0031] Figure 7 The illustration shows an example embodiment of the ACIM core based on a programmable resistor matrix and a direct-drive analog-to-digital converter;
[0032] Figure 8 The illustration shows another example embodiment of the ACIM core based on a programmable resistor matrix and a direct-drive analog-to-digital converter;
[0033] Figure 9 An example embodiment illustrating the further propagation of the conversion result of a directly driven successive approximation analog-to-digital converter is shown;
[0034] Figure 10 Another example embodiment illustrating the further propagation of the conversion result of a direct-drive successive approximation analog-to-digital converter is shown;
[0035] Figure 11 A third example embodiment illustrating the further propagation of the conversion result of a directly driven successive approximation analog-to-digital converter is shown.
[0036] Figure 12A The diagram illustrates the digital domain activation function based on a digital rectified linear unit;
[0037] Figure 12B The diagram illustrates the digital domain activation function based on a digital comparator;
[0038] Figure 13 A fourth example embodiment illustrating the further propagation of the conversion result of a directly driven successive approximation analog-to-digital converter is shown.
[0039] Figure 14 The illustration shows an example embodiment of a digital-to-analog converter architecture for resistive analog memory computing cores based on R-2R topology and its voltage source equivalent;
[0040] Figure 15The illustration shows an example embodiment of a 4-bit implementation of a programmable resistor in a computation matrix in analog memory controlled by the tuning word w;
[0041] Figure 16A The illustration shows an example embodiment of signal routing analysis in the first column of the computation matrix core in simulated memory;
[0042] Figure 16B The illustration shows an equivalent circuit model of an example embodiment of signal routing analysis in the first column of the computation matrix core in simulated memory;
[0043] Figure 17 The illustration shows a third example embodiment of an extended ACIM core based on a programmable resistor matrix;
[0044] Figure 18 The illustration shows a fourth example embodiment of the extended ACIM core based on a programmable resistor matrix;
[0045] Figure 19 The illustration shows an example embodiment of the matrix inversion ACIM core based on a programmable resistor matrix and a direct-drive analog-to-digital converter;
[0046] Figure 20 The illustration shows an example embodiment of the ACIM core of the Ising machine based on a programmable resistor matrix and a direct-drive analog-to-digital converter;
[0047] Figure 21 The illustration shows an example embodiment of calculating the differential derivative of a 4-bit implementation of a single-ended weighted cell of a programmable resistor in a matrix in simulated memory;
[0048] Figure 22 The illustration shows an example embodiment of a hardware implementation that supports a direct-drive ACIM core with negative weights;
[0049] Figure 23 The illustration shows another example embodiment of a hardware implementation that supports negative weights for directly driving the ACIM core;
[0050] Figure 24 The illustration shows a third example embodiment of a hardware implementation that supports negative weights and directly drives the ACIM core.
[0051] Figure 25 The illustration shows an example of an apparatus including components for performing one or more of the example embodiments described above;
[0052] Figure 26 The illustration shows an example of an apparatus including components for performing one or more of the example embodiments described above;
[0053] Figure 27Examples of apparatuses combined with user devices / user equipment according to some embodiments of the present invention are illustrated;
[0054] Figure 28 Examples of devices combined with node elements according to some embodiments of the present invention are illustrated; and
[0055] Figure 29 Examples of apparatuses combined with base stations according to some embodiments of the present invention are illustrated. Detailed Implementation
[0056] The following embodiments are exemplary. Although the specification may refer to "a," "an," or "some" embodiments in several places in the text, this does not necessarily mean that each reference to the same embodiment(s) or a particular feature applies only to a single embodiment. Individual features of different embodiments may also be combined to provide other embodiments within the scope of the claims. Furthermore, the words "comprising" and "including" should be understood as not limiting the described embodiments to consisting only of those features already mentioned, and these embodiments may also include features not specifically mentioned. Reference numerals in the specification and / or claims are used to illustrate embodiments with reference to the accompanying drawings, and not to limit the embodiments to these examples.
[0057] Some of the example embodiments described herein can be implemented in wireless communication networks, including radio access networks based on one or more of the following radio access technologies (RATs): Global System for Mobile Communications (GSM) or any other second-generation radio access technology, Universal Mobile Telecommunications System (UMTS, 3G) based on Basic Wideband Code Division Multiple Access (W-CDMA), High-Speed Packet Access (HSPA), Long Term Evolution (LTE), LTE Advanced, Fourth Generation (4G), Fifth Generation (5G), 5G New Radio (NR), Advanced 5G (i.e., 3GPP NR Rel-18 and later), or Sixth Generation (6G). Some examples of radio access networks include: Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN), Evolved Universal Terrestrial Radio Access (E-UTRA), or Next Generation Radio Access Network (NG-RAN). The wireless communication network may also include a core network, and some example embodiments may also be applied to the network functions of the core network.
[0058] It should be noted that the embodiments are not limited to the wireless communication network given as an example, and those skilled in the art can apply the solution to other wireless communication networks or systems with the necessary characteristics. For example, some example embodiments can also be applied to communication systems based on the IEEE 802.11 standard or communication systems based on the IEEE 802.15 standard. IEEE is an abbreviation for the Institute of Electrical and Electronics Engineers.
[0059] Figure 1 The illustration shows an example of an ACIM core based on a programmable resistor matrix. An Analog In-Memory Computation Accelerator (ACIM Accelerator) can be implemented based on one or more Analog In-Memory Computation Cores (ACIM Cores). A typical approach to implementing an ACIM core is based on a programmable resistor matrix. Figure 1 The resistor matrix shown in the example implements a VMM (Voice-Matrix Multiplication) operation in the analog domain: y = Ax, where x is the vector of the input signal, A is a matrix, and y is the vector of the output signal of the VMM operation, given by Equation 1. (1) (2)
[0060] Figure 1 The matrix A described in the example and Equation 2 is 4×4 in size, while the real ACIM matrix is typically significantly larger, such as 128×128. Each element a in matrix A... xx Typically composed of a programmable resistor R u or programmable resistor R u Array composition. In large ACIM cores, large resistors R are preferred. u A resistor is used to keep the current level in the matrix controllable. The input vector x can initially be analog or digital.
[0061] When the input signal is initially digital, and the vector of the input signal (input vector x) is initially digital, the input vector x is converted to the analog domain by a digital-to-analog converter (DAC) array, such as... Figure 1 As shown in the example. The converted analog domain input signal is then typically buffered by an amplifier array, such as... Figure 1 As shown in the example.
[0062] The output signal of the ACIM core is usually also buffered by an amplifier array, such as Figure 1 As shown in the example. In addition to buffering, the output amplifier also biases the columns of the matrix to the desired reference potential level v. ref ,like Figure 1 As shown in the example, the output signal vector (output vector y) of the ACIM core can be used in either the analog or digital domain.
[0063] When the output vector y of the ACIM core is used in the analog domain, it can be used directly without conversion. When used in the digital domain, the output vector y of the ACIM core must first be converted to the digital domain by an analog-to-digital converter (ADC) array, such as... Figure 1 As shown in the example.
[0064] Figure 2 An example of an extended ACIM core based on a programmable resistor matrix is illustrated. The applicability of the ACIM core can be expanded through different operating modes and extensions. Figure 2 In the example shown, the applicability of the ACIM cores is expanded by concatenating the output vector of one ACIM core to the input of another sequential ACIM core. In machine learning transformers, the output of the ACIM cores can be processed with additional operators before propagation; a common operator is ReLU (Rectified Linear Unit), such as... Figure 2 As shown in the example.
[0065] Figure 3 Another example of an extended ACIM core based on a programmable resistor matrix is illustrated. The applicability of the ACIM core can be expanded through different operating modes and extensions. Figure 3 In another example shown, the applicability of the ACIM core is expanded through another extension mode, where two ACIM cores are merged into a single effective VMM core, doubling the length of its input vector. For clarity, Figure 3 The different switches and mode control signals required for switching between different extended modes are omitted.
[0066] Currently, the ACIM core is typically buffered by an analog amplifier, such as Figure 1 As shown in the figure. Due to the limitations of the amplifier-based approach, the feasible implementation of extending the ACIM core with different operating modes and extensions remains limited.
[0067] The resistor matrix of an ACIM core is typically driven and biased by an array of operational amplifiers. However, the overhead power, overhead noise, overhead area, finite output voltage swing, and finite bandwidth of these amplifiers generally limit the performance of resistive ACIM cores. The amplifier's feedback network has a considerable impact on its closed-loop bandwidth and stability, which limits the operating modes, i.e., configurability, of the ACIM core. Therefore, no general-purpose ACIM core with multiple operating modes or extensions has been proposed in the literature. Finally, analog signal processing operators and memories are highly nonlinear, mismatched, and dependent on process, voltage, temperature, and aging PVTA effects (PVTA, process, voltage, temperature, aging).
[0068] Figure 4 The illustration shows examples of traditional successive approximation ADCs with different input voltage levels. A successive approximation ADC, or SAR ADC (SAR, successive approximation register), is an analog-to-digital converter that uses successive approximation steps to convert a continuous analog waveform into a discrete digital representation. Discrete-time SAR ADCs typically employ a binary search algorithm. Figure 4In the example of the successive approximation ADC shown, the SAR ADC does not drive its input node; therefore, it requires a preamplifier stage to properly bias the VMM core matrix columns. Figure 4 The input voltage level (i.e., the output of the bias amplifier) of the conventional SAR ADC described in the example will vary for each successive approximation ADC.
[0069] Computing technology and resistive memory computing devices can be used in communication and wireless technologies.
[0070] Figure 5 An example of a simplified wireless communication network is depicted, showing some physical and logical entities. Figure 5 The connections shown can be physical or logical connections. Those skilled in the art will understand that wireless communication networks may also include, in addition to... Figure 5 Other physical and logical entities besides the physical and logical entities shown.
[0071] However, the exemplary embodiments described herein are not limited to the wireless communication networks given as examples, but those skilled in the art can apply the exemplary embodiments described herein to other wireless communication networks with the necessary characteristics.
[0072] Figure 5 The example wireless communication network shown includes a radio access network (RAN) and a core network 110.
[0073] Figure 5 User equipment (UE) 100, 102 is shown that is configured to wirelessly connect to access node 104 of radio access network on one or more communication channels in radio cell.
[0074] Access node 104 may include a computing device configured to control the radio resources of access node 104 and be in a radio connection with one or more UEs 100 and UE 102. Access node 104 may also be referred to as a base station, base transceiver station (BTS), access point, cell site, network node, radio access network node, or RAN node. In this specification, the terms "access node" and "radio access network node" are used interchangeably.
[0075] Access node 104 may be, for example, an evolved NodeB (eNB or eNodeB) providing a radio cell, a next-generation evolved NodeB (ng-eNB), or a next-generation NodeB (gNB or gNodeB). Access node 104 may include or be coupled to a transceiver. From the transceiver of access node 104, a connection to an antenna element may be provided, which establishes a bidirectional radio link to one or more UEs 100 and UE 102. The antenna element may include an antenna or antenna element, or multiple antennas or antenna elements.
[0076] The radio connection (e.g., a radio link) from UE 100, UE 102 to access node 104 may be referred to as an uplink (UL) or reverse link, while the radio connection from access node 104 to UE 100, UE 102 may be referred to as a downlink (DL) or forward link. UE 100 may also communicate directly with another UE 102 via a radio connection commonly referred to as a sidelink (SL), and vice versa. It should be understood that access node 104 or its functionality may be implemented using any node, host, server, access point, or other entity suitable for providing such functionality.
[0077] A radio access network may include more than one access node 104, in which case the access nodes may also be configured to communicate with each other via wired or wireless links. These links between access nodes may be used to send and receive control plane signaling, or to route data from one access node to another.
[0078] Access node 104 can also be connected to core network (CN) 110. Core network 110 may include an evolved packet core (EPC) network and / or a fifth-generation core network (5GC). EPC may include network entities such as a serving gateway (S-GW for routing and forwarding data packets), a packet data network gateway (P-GW) for providing connectivity to external packet data networks for the UE, and / or a mobility management entity (MME). 5GC may include one or more network functions such as at least one of the following: user plane function (UPF), access and mobility management function (AMF), location management function (LMF), and / or session management function (SMF).
[0079] The core network 110 can also communicate with or utilize services provided by one or more external networks 113, such as the public switched telephone network or the Internet. For example, in a 5G wireless communication network, the UPF of the core network 110 can be configured to communicate with an external data network via the N6 interface. In an LTE wireless communication network, the P-GW of the core network 110 can be configured to communicate with an external data network.
[0080] It should also be understood that, compared to LTE or 5G, the functional allocation between core network operations and access node operations may differ or even not exist in future wireless communication networks.
[0081] The UE 100 and UE 102 shown are a type of device that can be allocated and assigned resources on an air interface. UE 100 and UE 102 may also be referred to as wireless communication devices, subscriber units, mobile stations, remote terminals, access terminals, user terminals, terminal equipment, mobile phones, smartphones, or user equipment, to name just a few. UE 100 and UE 102 can be computing devices operating with or without a Subscriber Identity Module (SIM), including but not limited to the following types of computing devices: mobile phones, smartphones, personal digital assistants (PDAs), handheld devices, computing devices including wireless modems (e.g., alarm or measuring devices), laptop computers, desktop computers, tablet computers, game consoles, laptops, multimedia devices, RedCap devices, wearable devices with radio components (e.g., watches, medical or health devices, headphones or glasses), sensors including wireless modems, or computing devices including wireless modems integrated in vehicles.
[0082] Any features of the UE described in this document can also be implemented using corresponding devices, such as relay nodes. An example of such a relay node could be a Layer 3 relay (self-backhaul relay) toward an access node. A self-backhaul relay node can also be referred to as an Integrated Access and Backhaul (IAB) node. An IAB node may include two logical parts: a Mobile Terminal (MT) part, which is responsible for (multiple) backhaul links (i.e., (multiple) links between the IAB node and the donor node (also called the parent node); and a Distributed Unit (DU) part, which is responsible for (multiple) access links, i.e., (multiple) sub-links (multiple hop scenarios) between the IAB node and (multiple) UEs and / or between the IAB node and other IAB nodes.
[0083] Another example of such a relay node could be a Layer 1 relay, referred to as a repeater. A repeater can amplify signals received from an access node and forward them to the UE, and / or amplify signals received from a UE and forward them to the access node.
[0084] It should be understood that UE 100 and UE 102 can also be almost exclusively uplink-only devices, examples of which could be cameras or camcorders that load image or video clips onto the network. UE 100 and UE 102 can also be devices capable of operating in Internet of Things (IoT) networks, a scenario that can provide objects with the ability to transmit data over a network without requiring human-to-human or human-to-computer interaction.
[0085] Wireless communication networks can also support the use of cloud services. For example, at least a portion of the core network operation can be performed as a cloud service (this is in...). Figure 5 (This is represented by "Cloud" 114). UE 100 and UE 102 can also utilize Cloud 114. In some applications, computations for a given UE can be performed in Cloud 114 or in another UE.
[0086] Wireless communication networks can also include a central control entity, such as a Network Management System (NMS). An NMS is a centralized suite of software and hardware used to monitor, control, and manage network infrastructure. The NMS is responsible for a wide range of tasks, such as fault management, configuration management, security management, performance management, and billing management. The NMS enables network operators to efficiently manage and optimize network resources to ensure the network delivers high performance, reliability, and security.
[0087] 5G enables the use of multiple-input multiple-output (MIMO) antennas in access node 104 and / or UE 100, UE 102, far more base stations or access nodes than LTE networks (the so-called small cell concept), including macro sites operating in cooperation with smaller sites, and the use of various radio technologies depending on service requirements, use cases, and / or available spectrum. 5G wireless communication networks can support a wide range of use cases and related applications, including video streaming, augmented reality, different data sharing methods, and various forms of machine-type applications such as (massive) machine-type communication (mMTC), including vehicle safety, various sensors, and real-time control.
[0088] In 5G wireless communication networks, access nodes and / or UEs can have multiple radio interfaces, such as sub-6 GHz, centimeter wave (cmWave), and millimeter wave (mmWave), and can also be integrated with traditional radio access technologies such as LTE. For example, integration with LTE can be implemented as a system where LTE provides macro coverage, and 5G radio interface access can be aggregated to LTE from small cells. In other words, 5G wireless communication networks can support both RAT interoperability (such as interoperability between LTE and 5G) and RI interoperability (interoperability between radio interfaces, such as sub-6 GHz, cmWave, and mmWave).
[0089] 5G wireless communication networks can also apply network slicing, in which multiple independent and dedicated virtual subnets (network instances) can be created within the same physical infrastructure to run services with different requirements for latency, reliability, throughput and mobility.
[0090] 5G enables analytics and knowledge generation at the data source. This approach can involve leveraging resources that may not have continuous network connectivity, such as laptops, smartphones, tablets, and sensors. Multi-access edge computing (MEC) can provide a distributed computing environment for application and service hosting. It can also store and process content near cellular subscribers for faster response times. Edge computing can encompass a wide range of technologies, such as wireless sensor networks, mobile data acquisition, mobile signature analytics, collaborative distributed peer-to-peer self-organizing networks and processing (which can also be categorized as local cloud / fog computing and grid / mesh computing), dew computing, mobile edge computing, cloudlets, distributed data storage and retrieval, autonomous self-healing networks, remote cloud services, augmented and virtual reality, data caching, the Internet of Things (massive connectivity and / or time-critical), and critical communications (autonomous vehicles, traffic safety, real-time analytics, time-critical control, healthcare applications).
[0091] In one embodiment, access node 104 may include: a radio unit (RU) 103 comprising radio transceivers (TRXs), i.e., transmitters (Tx) and receivers (Rx); one or more distributed units (DUs) 105, which may be used for so-called Layer 1 (L1) processing and real-time Layer 2 (L2) processing; and a central unit (CU) 108 (also called a centralized unit), which may be used for non-real-time L2 and Layer 3 (L3) processing. CU 108 may be connected to one or more DUs 105, for example, via an F1 interface. Such an embodiment of access node 104 enables the centralization of CUs relative to cell sites and DUs, while DUs may be more distributed and may even be retained at the cell site. CUs and DUs together may also be referred to as baseband or baseband unit (BBU). CUs and DUs may also be included in a radio access point (RAP).
[0092] CU 108 may be a logical node that hosts the Radio Resource Control (RRC), Service Data Adaptation Protocol (SDAP), and / or Packet Data Convergence Protocol (PDCP) of the NR protocol stack of access node 104. CU 108 may include a control plane (CU-CP), which may be a logical node that hosts the control plane portions of the RRC and PDCP protocols of the NR protocol stack of access node 104. CU 108 may also include a user plane (CU-UP), which may be a logical node that hosts the user plane portion of the PDCP protocol and the SDAP protocol of the CU of access node 104.
[0093] DU 105 can be a logical node for the Radio Link Control (RLC), Media Access Control (MAC), and / or Physical (PHY) layers of the NR protocol stack of the managed access node 104. The operation of DU 105 can be controlled at least partially by CU 108. It should also be understood that the functional distribution between DU 105 and CU 108 can vary depending on the implementation.
[0094] Cloud computing systems can also be used to provide CU 108 and / or DU 105. CUs provided by cloud computing systems can be referred to as virtualized CUs (vCUs). In addition to vCUs, virtualized DUs (vDUs) provided by cloud computing systems can also exist. Furthermore, there can be a combination where DUs can be implemented on so-called bare-metal solutions, such as application-specific integrated circuits (ASICs) or customer-specific standard product (CSSP) system-on-chips (SoCs).
[0095] By leveraging Network Functions Virtualization (NFV) and Software-Defined Networking (SDN), edge cloud can be introduced into the radio access network. Using edge cloud can represent access node operations that should be performed at least partially in a computing system operatively coupled to the Remote Radio Head (RRH) or Radio Unit (RU) 103 of access node 104. Alternatively, access node operations can be performed on a distributed computing system or cloud computing system located at access node 104. The application of cloud RAN architecture enables real-time RAN functions to be performed at the radio access network (e.g., in DU 105) and non-real-time functions to be performed in a centralized manner (e.g., in CU 108).
[0096] 5G (or New Radio) wireless communication networks can support multiple hierarchical structures, in which multiple access edge computing (MEC) servers can be placed between the core network 110 and the access node 104. It should be understood that MEC can also be applied to LTE wireless communication networks.
[0097] 5G wireless communication networks (5G networks) can also include non-terrestrial communication networks, such as satellite communication networks, to enhance or supplement the coverage of 5G radio access networks. For example, satellite communication can support data transmission between the 5G radio access network and the core network 110, thereby enabling wider network coverage. Possible use cases could include providing service continuity for machine-to-machine (M2M) or Internet of Things (IoT) devices or for passengers on board vehicles, or ensuring the availability of services for critical communications and future rail, maritime, or air communications. Satellite communication can utilize geostationary orbit (GEO) satellite systems or low Earth orbit (LEO) satellite systems, particularly mega-constellations (i.e., systems deploying hundreds of (nano) satellites). A given satellite 106 in a mega-constellation can cover the network entity of several enabling satellites that create a terrestrial cell. Terrestrial cells can be created by ground relay access nodes or by access nodes located on the ground or in satellites.
[0098] Those skilled in the art will understand that Figure 5 The access node 104 shown is merely an example of a portion of the radio access network. In practice, the radio access network may include multiple access nodes 104, UE 100 and UE 102 may access multiple radio cells, and the radio access network may also include other devices, such as physical layer relay access nodes or other entities. At least one of the access nodes may be a home eNodeB or a home gNodeB. A home gNodeB or a home eNodeB is an access node that can be used to provide indoor coverage in a home, office, or other indoor environment.
[0099] Furthermore, within the geographical area of the radio access network, multiple different types of radio cells and multiple radio cells can be provided. Radio cells can be macrocells (or umbrella cells), which can be large areas with diameters of up to tens of kilometers, or smaller cells, such as microcells, femtocells, or picocells. Figure 5 The (multiple) access nodes 104 can provide any type of these cells. A cellular radio network can be implemented as a multi-layered access network comprising several types of radio cells. In a multi-layered access network, one access node can provide one or more types of radio cells, and therefore providing such a multi-layered access network may require multiple access nodes.
[0100] To meet the need for improved performance in radio access networks, the concept of "plug-and-play" access nodes can be introduced. Besides the home eNodeB or home gNodeB, radio access networks capable of using "plug-and-play" access nodes can also include home node B gateways (HNB-GW). Figure 5(Not shown in the image). An HNB-GW, which can be installed within an operator's radio access network, can aggregate traffic from a large number of home eNodeBs or home gNodeBs back to the operator's core network 110.
[0101] 6G wireless communication networks are expected to employ flexible decentralized and / or distributed computing systems and architectures, along with ubiquitous computing, based on mobile edge computing, artificial intelligence, short packet communication, and blockchain technologies, to achieve local spectrum licensing, spectrum sharing, infrastructure sharing, and intelligent automated management. Key features of 6G may include intelligent interconnected management and control capabilities, programmability, integrated sensing and communication, reduced energy footprint, trusted infrastructure, scalability, and affordability. Furthermore, 6G targets new use cases, including integrating location and sensing capabilities into the system definition to unify the user experience in the physical and digital worlds.
[0102] However, the following uses the principles and terminology of 5G radio access technology to describe some example embodiments, without limiting the example embodiments to 5G radio access technology.
[0103] Figure 6 The illustration shows an example embodiment of a direct-drive successive approximation analog-to-digital converter (ADC). A successive approximation ADC, or SAR ADC (SAR, successive approximation register), is an analog-to-digital converter that uses successive approximation steps to convert a continuous analog waveform into a discrete digital representation. Discrete-time SAR ADCs typically employ a binary search algorithm, while continuous-time SAR ADCs can also use a linear search algorithm. Figure 6 In the example embodiment of the direct-drive successive approximation ADC shown, the SAR ADC drives its input node. Therefore, the direct-drive SAR ADC can be self-biased using its inherent digital-to-analog converter iDAC (iDAC, Intrinsic Digital-to-Analog Converter).
[0104] Assume the ACIM core has a large unit resistance R. u With resistors, the iDAC does not require a sequential buffer amplifier. The input voltage level of the directly driven SAR ADC remains constant, making comparator design much easier.
[0105] In a direct-drive configuration, the polarity of the iDAC is reversed because the direct-drive iDAC attempts to cancel out any changes in the input signal level of the SAR ADC. The output of the direct-drive iDAC is directly connected to the input of the SAR ADC, such as... Figure 6 As shown in the diagram. In another embodiment, the output of the directly driven iDAC is connected to the input of the SAR ADC via a feedback network.
[0106] exist Figure 6In the example embodiment shown, an apparatus for calculating a resistance matrix in analog memory including a direct-drive successive approximation analog-to-digital converter (DADC) is provided, at least one of the DADCs having a self-biasing circuit system comprising: a comparator in the input of the DADC, the comparator being arranged to compare a reference node voltage with an input node voltage from an output of an intrinsic digital-to-analog converter (ADC), the intrinsic ADC receiving an output signal of the DADC having reverse polarity as an input signal.
[0107] In an example embodiment of the device that includes a resistor matrix computed in analog memory, direct-drive analog-to-digital converter technology is used to enable flexible scalability and reconfigurability for the ACIM core.
[0108] In calculating the resistance matrix in analog memory using direct-drive analog-to-digital converter (ADC) technology, the ADC is not very sensitive to changes in its feedback network as long as the network's time constant is small compared to the ADC's clock rate. Therefore, the conversion result of the direct-drive ADC can be further propagated.
[0109] Figure 7 An example embodiment of the ACIM core based on a programmable resistor matrix and a direct-drive analog-to-digital converter is illustrated. Figure 7 The resistance matrix shown in the example performs the same VMM operation y=Ax in the analog domain, where x is a vector, A is a matrix, and y is a vector, as shown below. Figure 1 Examples and those described in Equations 1 and 2.
[0110] Figure 7 The example embodiment and the matrix A described in Equation 2 have a size of 4×4, while the real ACIM matrix is typically significantly larger, such as, but not limited to, 128×128. Each element a in matrix A xx Typically composed of a programmable resistor R u or programmable resistor R u Array composition. In large ACIM cores, large resistors R are preferred. u Resistors are used to keep the current level in the matrix controllable. Figure 7 In the example embodiment shown, the input vector x in the resistance matrix is simulated. Figure 7 The resistor matrix deployment shown in the example embodiment directly drives the ADC, replacing the performance-limited amplifier in the ACIM core.
[0111] exist Figure 7In the example embodiment shown, an apparatus for calculating a resistance matrix in analog memory is provided as an input to the analog memory included in the apparatus. When a vector-matrix multiplication operation is performed in the apparatus, a digital signal is received from the output of the analog memory-based resistance matrix calculation.
[0112] Figure 8 Another example embodiment of the ACIM core based on a programmable resistor matrix and a direct-drive analog-to-digital converter is illustrated. Figure 8 The resistor matrix shown in the example embodiment implements the same VMM operation y=Ax in the analog domain, where x is a vector, A is a matrix, and y is a vector, as shown below. Figure 1 Examples and those described in Equations 1 and 2. Figure 8 The example embodiment and the matrix A described in Equation 2 have a size of 4×4, while the real ACIM matrix is typically significantly larger, such as, but not limited to, 128×128. Each element a in matrix A xx Typically composed of a programmable resistor R u or programmable resistor R u Array composition. In large ACIM cores, large resistors R are preferred. u Resistors are used to keep the current level in the matrix controllable.
[0113] exist Figure 8 In the example embodiment shown, the input vector x in the resistance matrix is digital. Figure 8 The resistor matrix deployment shown in the example embodiment directly drives the DAC and ADC, replacing the performance-limited amplifier in the ACIM core.
[0114] exist Figure 8 In the example embodiment shown, the apparatus for calculating the resistance matrix in analog memory can be implemented using a unit resistor R. u A large resistor matrix can be used to deploy a direct-drive DAC without requiring a sequential buffer amplifier. This is because the unit resistor R... u With a relatively large resistance, directly driving the DAC would result in a high-impedance resistive load. Although a buffer is omitted, its output swing will not be significantly attenuated. Variable analog gain corresponds to logic shift operations in the digital domain.
[0115] exist Figure 8 In the example embodiment shown, an apparatus for calculating a resistance matrix in analog memory is provided with an analog signal as an input to the analog memory included in the apparatus. Figure 8In an example embodiment, the provided analog signal is converted from a digital signal by a direct-drive digital-to-analog converter (DAC) at the input of the resistance matrix calculation in the analog memory. When vector-matrix multiplication is performed in this device, the digital signal is received from the output of the direct-drive DAC at the output of the resistance matrix calculation in the analog memory.
[0116] Figure 9 An example embodiment illustrating further propagation of the conversion result from a directly driven successive approximation analog-to-digital converter is shown. Figure 9 In the example embodiment shown, an apparatus for calculating a resistance matrix in analog memory with directly driven successive approximation analog-to-digital converters is provided with the comparator's output signal via synchronous successive approximation logic of the at least one directly driven successive approximation analog-to-digital converter. Propagation from the comparator output requires only a single signal line, making routing easier in dense environments, but requiring parallel synchronous SAR logic blocks. Dashed lines mark potentially long signal lines or potentially long data buses.
[0117] Figure 10 Another example embodiment illustrating the further propagation of the conversion result of a directly driven successive approximation analog-to-digital converter is shown. Figure 10 In the example embodiment shown, an apparatus for calculating a resistance matrix in analog memory, including direct-drive successive approximation analog-to-digital converters, at least one of the direct-drive successive approximation analog-to-digital converters is arranged to provide a logic block output signal to at least one other digital-to-analog converter. Propagation from the SAR logic block output requires no additional hardware resources, but requires routing of a fully digital bus. Dashed lines mark potentially long signal lines or potentially long data buses.
[0118] Figure 11 A third example embodiment illustrating the further propagation of the conversion result from a directly driven successive approximation analog-to-digital converter is shown. Figure 11 In the example embodiment shown, the apparatus for calculating a resistor matrix in analog memory, including at least one of the at least one directly driven successive approximation analog-to-digital converters, is arranged to provide the output signal of the inherent digital-to-analog converter to at least one load. Propagation from the output of the digital-to-analog converter requires no additional hardware resources, but some time-domain multiplexing is required. Dashed lines mark potentially long signal lines or potentially long data buses.
[0119] In one embodiment of an apparatus for calculating a resistance matrix in analog memory having a direct-drive successive approximation analog-to-digital converter, the apparatus includes an activation function between the output of the synchronous successive approximation logic and the input of the intrinsic digital-to-analog converter or between the output of the synchronous successive approximation logic and the input of the at least one other digital-to-analog converter.
[0120] Figure 12A The diagram illustrates the digital domain activation function based on a digital rectified linear unit. Figure 12A In the example embodiment, a simple multiplexer for N-bit digital words is used to implement the signal processing operations of the digital rectified linear unit in the digital domain. To implement the digital rectified linear unit, Figure 12A The multiplexer uses the sign bit in_d[N-1] of its input data to determine whether to propagate the input data as is or to propagate zeros. In a device that includes analog memory for calculating the resistance matrix with a direct-drive successive approximation analog-to-digital converter, Figure 12A The numerical domain activation function shown in the example embodiment can be deployed in a machine learning accelerator.
[0121] Figure 12B The diagram illustrates a digital domain activation function based on a digital comparator. Figure 12B In the example embodiment, a simple multiplexer for N-bit digital words is used to implement the signal processing operations of a digital comparator in the digital domain, such as level comparison. To implement a digital comparator, Figure 12B The multiplexer uses the sign bit of its input data to determine whether to propagate a positive or negative number, i.e., by comparing it with a zero input data value. In a device that includes analog memory for calculating the resistance matrix with a direct-drive successive approximation analog-to-digital converter, Figure 12B The digital domain activation function shown in the example embodiment can be deployed in the Ising machine.
[0122] Figure 13 A fourth example embodiment illustrating the further propagation of the conversion result from a directly driven successive approximation analog-to-digital converter is shown. Figure 13 In the example embodiment shown, the apparatus for calculating a resistance matrix in analog memory, including a direct-drive successive approximation analog-to-digital converter (DADC), is arranged to provide the output signal of the intrinsic digital-to-analog converter (DAC) to at least one load. Propagation from the output of the DAC requires no additional hardware resources, but some time-domain multiplexing is necessary. Figure 13 In an example embodiment, a rectified linear unit is added between the SAR logic block output and the input of an analog-to-digital converter with inverted polarity. Figure 11 In an example embodiment, dashed lines mark potential long signal lines or potential long data buses.
[0123] In addition to various signal processing operators, non-volatile memory is also easy to implement in the digital domain, while analog signal processing operators and memory suffer from leakage, charge injection, high nonlinearity, mismatch, and dependence on process, voltage, temperature, and aging PVTA (PVTA, process, voltage, temperature, aging).
[0124] Figure 14 The illustration shows an example embodiment of a digital-to-analog converter (DAC) architecture for a computational core in resistive analog memory based on an R-2R topology and its voltage source equivalent. The potential DAC architecture for a computational core in resistive analog memory is based on an R-2R topology, where a 4-bit R-2R DAC is as follows: Figure 14 As shown, the unit resistance is r u .
[0125] The control of the digital-to-analog converter is performed via switching branches controlled by the tuning word x. Bit x3 represents the most significant bit, and bit x0 represents the least significant bit. One bias rail of each branch is connected to the reference potential v. ref The other track is connected to one of the DAC's power supplies. The sign bit S of the control word indicates that the connected power supply is V. dd or v ss Potential. In many practical implementations, , However, for the sake of convenience, let's assume here... , This assumption will not affect the final results of the analysis. Assuming positive sign control is used, our N... x The corresponding driving voltage v of the equivalent voltage source model of the bit-to-analog converter x As given by Equation 3: (3)
[0126] Figure 15 The illustration shows an example embodiment of a 4-bit implementation of a programmable resistor in an analog-memory compute matrix controlled by a tuning word w. In the presented example embodiment of the 4-bit implementation of the programmable resistor in the analog-memory compute matrix, the architecture used in the digital-to-analog converter architecture for the resistive analog-memory compute core is based on an R-2R topology. In the presented example embodiment, a unit resistor R is used. u Furthermore, the programmable R-2R resistors of the R-2R topology are controlled by the tuning word w.
[0127] The input V of the programmable resistor r In the voltage domain, the output i of the programmable resistor cIn the current domain, the input voltage of the programmable resistor is set by the direct-drive digital-to-analog converter at the corresponding matrix row in memory, while the output voltage v of the programmable resistor... c The intrinsic digital-to-analog converter, driven by a direct successive approximation analog-to-digital converter, calculates the matrix column settings in the corresponding memory. By default, the output voltage of the programmable resistor is set to be the same as the reference voltage v. ref Same level. Assuming functional output voltage bias, N w The output current of the bit programmable resistor is i C As given by Equation 4: (4)
[0128] A key feature of programmable R-2R resistors is that their effective input resistance is always R. u And it has nothing to do with the word "tuning".
[0129] Figure 16A The illustration shows an example implementation of signal routing analysis in the first column of a computational matrix core in simulated memory. In the following analysis, we examine... Figure 16A The first column of the ACIM core, where the analyzed signal routes are highlighted in black. For clarity, the core size in the simulated memory shown is 4×4, but our analysis assumes a core size of N. R ×N C , where N R N is the number of rows in the computing core within simulated memory. C It is the number of columns in the computing core in simulated memory.
[0130] Figure 16B The diagram illustrates an equivalent circuit model of an example embodiment of signal routing analysis in the first column of a computation matrix core in simulated memory.
[0131] To simplify the following analysis, we assume that all input digital-to-analog converters are controlled identically. Furthermore, we assume that all programmable resistors at the analyzed columns are controlled identically so that they use maximum relative weights to push the maximum current to the analyzed columns. Additionally, we examine a rectangular ACIM core, i.e., N... R =N C The goal of this analysis is to find the conditions under which the input digital-to-analog converter and the direct-drive analog-to-digital converter utilize their maximum dynamic characteristics, i.e., x=y.
[0132] The input digital-to-analog converter is modeled by its equivalent voltage source, which sees N C The net parallel resistance of a programmable resistor, i.e., R u / N C Therefore, the voltage level v at the analyzed row is... rEquation 5 expresses this as: (5) Where r ux It is the unit resistance level of the input digital-to-analog converter.
[0133] The net current of the column being analyzed is N. R The total current of the programmable resistors and the reverse current of the inherent digital-to-analog converter (DAC) that directly drives the ADC. The integrated iDAC of the ADC is modeled using its equivalent current source, whose polarity is easily accounted for by the inverted polarity of the iDAC. The unit resistance level of the inherent DAC is r. uy .
[0134] The column being analyzed must be biased to 0V for the computation in the simulated memory to run correctly, and therefore, the net current of the column must be zero. In the following analysis, we can write:
[0135] By inserting the output current i in Equation 4 C We can write:
[0136] By inserting the row voltage level v in Equation 5 r We can write:
[0137] By taking into account that the matrix is rectangular, i.e., N R =N C We can write:
[0138] By considering the maximum weight applied to the programmable resistor, and therefore considering the expression Approaching 1, that is We can write:
[0139] By considering the relative unit resistor size applied between the input DAC and the matrix, and therefore considering the unit resistance... We can write: ⇒
[0140] Because of R u The resistivity level of the ACIM core is relatively high, thus allowing it to achieve the proposed relative size. This is achieved by inserting the output current i in Equation 4. C We can write: Where N y It is an integrated analog-to-digital converter that directly drives the number of bits in the digital-to-analog converter, where y k It is the tuning bit of the intrinsic digital-to-analog converter, and where v ddx It is the power supply for the input digital-to-analog converter.
[0141] By considering that the input digital-to-analog converter and the intrinsic digital-to-analog converter have the same number of bits, i.e., N x =N y And considering the target having the same digital swing level at the input digital-to-analog converter and the successive approximation analog-to-digital converter, i.e., x k =y k We obtain the following results, as shown in Equation 6: (6)
[0142] Equation 6 shows that the output data value y of the analyzed signal path is proportional to the input data value x. It is worth noting that when r... uy =2r ux At that time, x k =y k Equation 6 also shows that power scaling can also be used to set the output code of a successive approximation analog-to-digital converter to be equal to the input code of the input digital-to-analog converter.
[0143] Figure 17 The illustration shows a third example of an extended ACIM core based on a programmable resistor matrix. Figure 17 In this context, the ACIM core based on a programmable resistor matrix is extended as the conversion results of the direct-drive successive approximation analog-to-digital converter are further propagated. Figure 17 The propagation from the output of a computation matrix core in analog memory to the input of another ACIM core is described, with rectified linear unit activation functions applied between the ACIM cores.
[0144] Figure 18 The illustration shows a fourth example of an extended ACIM core based on a programmable resistor matrix. Figure 18 In this design, two ACIM cores based on programmable resistor matrices are combined to form a dual-input vector length ACIM core. Figure 18 The paper describes merging a computation matrix core in analog memory into another ACIM core to achieve an efficient ACIM core with double the length of the input vector.
[0145] Figure 19 The illustration shows an example embodiment of the ACIM (Anchored Matrix Inversion) core based on a programmable resistor matrix and a direct-drive analog-to-digital converter. Figure 19In some embodiments, direct-drive techniques are applied to compute matrix reconfiguration in simulated memory to support matrix inversion calculations. For example, using... Figure 19 The configuration allows us to calculate the inverse of equation 1, which gives the result for equation 7: (7)
[0146] Figure 20 An example embodiment of the Ising machine ACIM core based on a programmable resistor matrix and a direct-drive analog-to-digital converter is illustrated. Figure 20 In some embodiments, direct-drive techniques are applied to the reconfiguration of the simulated in-memory computation matrix to support Ising machine computation. For example, Figure 20 The embodiment implements the Ising machine, which makes the Ising Hamiltonian H in Equation 8... p Minimize: (8)
[0147] Figure 21 The illustration shows an example embodiment of calculating the differential derivative of a 4-bit implementation of a single-ended weighted cell of a programmable resistor in a matrix within simulated memory. Some applications may require negative weights for vector-matrix multiplication, i.e., some elements in matrix A in Equation 1 should be negative.
[0148] exist Figure 21 The diagram illustrates a 4-bit implementation of a programmable resistor in an ACIM, controlled by a tuning word w and including differential weighted unit branches. Each differential weighted unit branch can be connected to one of two outputs. One output is connected to the column representing a positive signal (+), and the other is connected to the column representing a negative signal (-). The corresponding net current i... c+ and i c- This is expressed in equations 9 and 10 as follows: (9) (10)
[0149] The net current i in Equations 9 and 10 c+ and i c- They can be distinguished by two parallel direct-drive analog-to-digital converters, such as... Figure 22 and Figure 23 As shown. The differential current is given by Equation 11: (11)
[0150] Equation 11 shows that, depending on the applied weights, the differential current can acquire positive or negative values; that is, the net weights can have positive or negative values as expected.
[0151] Figure 22The illustration shows an example embodiment of a hardware implementation that supports a directly driven ACIM core with negative weights. For clarity, Figure 22 Only one matrix column is shown. Figure 22 In an example embodiment, a digital adder is used to extract the differential current i c+ -i c- The net signal is represented. For example... Figure 22 As shown, two parallel direct-drive analog-to-digital converters can distinguish the net current i c+ and i c- Dashed lines mark potential long signal lines or potential long data buses.
[0152] Figure 23 Another example embodiment of a hardware implementation supporting a directly driven ACIM core with negative weights is illustrated. For clarity, Figure 23 Only one matrix column is shown in the diagram. Figure 23 In an example embodiment, the analog summation at the output of the sequential digital-to-analog converter is used to extract the differential current i c+ -i c- The net signal is represented. For example... Figure 23 As shown, two parallel direct-drive analog-to-digital converters can distinguish the net current i c+ and i c- .exist Figure 23 In the propagation method of the example embodiment, at least one of the direct-drive successive approximation analog-to-digital converters is arranged to provide logic block output signals to at least one other digital-to-analog converter. Dashed lines mark potentially long signal lines or potentially long data buses.
[0153] Figure 24 The illustration shows a third example embodiment of a hardware implementation that supports a directly driven ACIM core with negative weights. For clarity, Figure 24 Only one matrix column is shown in the diagram. Figure 24 In the example implementation, a single-ended weighting unit and a serial switch are used to select whether to push the output current to the positive or negative signal line. Dashed lines mark potential long signal lines or potential long data buses.
[0154] Figure 25 The illustration shows an example of an apparatus including components for performing one or more of the example embodiments described above. For example, apparatus 9700 may be an apparatus such as user equipment (UE) 100, 102, etc., or an apparatus including the UE, or an apparatus included in the UE. User equipment may also be referred to as wireless communication device, subscriber unit, mobile station, remote terminal, access terminal, user terminal, terminal equipment, or user equipment.
[0155] Apparatus 9700 may include a circuit system, module, or chipset suitable for implementing one or more of the example embodiments described above. For example, apparatus 9700 may include at least one processor 9710. At least one processor 9710 interprets instructions (e.g., computer program instructions) and processes data. At least one processor 9710 may include one or more programmable processors. At least one processor 9710 may include programmable hardware with embedded firmware, and alternatively or additionally may include one or more application-specific integrated circuits (ASICs).
[0156] At least one processor 9710 is coupled to at least one memory 9720. The at least one processor is configured to write data to and read data from the at least one memory 9720. The at least one memory 9720 may include one or more memory cells. Memory cells may be volatile or non-volatile. It should be noted that one or more non-volatile memory cells and one or more volatile memory cells may be present, or alternatively, one or more non-volatile memory cells may be present, or alternatively, one or more volatile memory cells may be present. Volatile memory may be, for example, random access memory (RAM), dynamic random access memory, or synchronous dynamic random access memory (SDRAM). Non-volatile memory may be, for example, read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), flash memory, optical storage devices, or magnetic storage devices. Generally, memory may be referred to as a non-transitory computer-readable medium. The term "non-transitory" as used herein refers to a limitation on the medium itself (i.e., tangible, not tactile), rather than a limitation on the persistence of data storage (e.g., RAM and ROM). At least one memory 9720 stores computer-readable instructions that are executed by at least one processor 9710 to perform one or more of the example embodiments described above. For example, non-volatile memory stores computer-readable instructions, and at least one processor 9710 uses volatile memory for temporary storage of data and / or instructions to execute instructions. Computer-readable instructions may refer to computer program code.
[0157] Computer-readable instructions may have been pre-stored in at least one memory 9720, or alternatively or additionally, they may be received by the device via an electromagnetic carrier signal, and / or copied from a physical entity such as a computer program product. Execution of the computer-readable instructions by at least one processor 9710 causes the device 9700 to perform one or more of the above-described example embodiments. That is, at least one processor storing the instructions and at least one memory may provide components for providing or causing execution of any of the above methods and / or blocks.
[0158] In the context of this document, "memory" or "a computer-readable medium" or "a plurality of computer-readable media" can be any one or more nontransitory media or components that can contain, store, communicate, propagate, or transmit instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer. The term "nontransitory" as used herein is a limitation on the medium itself (i.e., tangible, not tactile), not a limitation on the persistence of data storage (e.g., RAM and ROM).
[0159] The device 9700 may also include or be connected to the input unit 9730. The input unit 9730 may include one or more interfaces for receiving input. These interfaces may include, for example, one or more temperature, motion, and / or orientation sensors, one or more cameras, one or more accelerometers, one or more microphones, one or more buttons, and / or one or more touch detection units. Furthermore, the input unit 9730 may include interfaces to which external devices can be connected.
[0160] The device 9700 may also include an output unit 9740. The output unit may include or be connected to one or more displays capable of rendering visual content, such as a light-emitting diode (LED) display, a liquid crystal display (LCD), and / or a liquid crystal on silicon (LCoS) display. The output unit 9740 may also include one or more audio outputs. The one or more audio outputs may be, for example, speakers.
[0161] Device 9700 also includes a connection unit 9750. Connection unit 9750 enables wireless connectivity to one or more external devices. Connection unit 9750 includes at least one transmitter and at least one receiver, which may be integrated into device 9700 or connected to the transmitter and receiver. The at least one transmitter includes at least one transmitting antenna, and the at least one receiver includes at least one receiving antenna. Connection unit 9750 may include an integrated circuit or set of integrated circuits providing wireless communication capabilities to device 9700. Alternatively, the wireless connection may be a hardwired application-specific integrated circuit (ASIC). Connection unit 9750 may also provide components for performing at least some of the blocks or functions of one or more of the example embodiments described above. Connection unit 9750 may include one or more components controlled by a corresponding control unit, such as a power amplifier, digital front-end (DFE), analog-to-digital converter (ADC), digital-to-analog converter (DAC), frequency converter, modulator (demodulator), and / or encoder / decoder circuitry.
[0162] It should be noted that device 9700 may also include Figure 25 Various components are not shown. These various components can be hardware components and / or software components.
[0163] Figure 26 An example of an apparatus 9800 including components for performing one or more of the example embodiments described above is illustrated. For example, apparatus 9800 may be an apparatus such as access node 104, or an apparatus including access node 104, or an apparatus included in access node 104.
[0164] Apparatus 9800 may include, for example, a circuit system, module, or chipset suitable for implementing one or more of the example embodiments described above. Apparatus 9800 may be an electronic device including one or more electronic circuit systems. Apparatus 9800 may include a communication control circuit system 9810 (such as at least one processor) and at least one memory 9820 storing instructions 9822, which, when executed by at least one processor, cause apparatus 9800 to perform one or more of the example embodiments described above. For example, such instructions 9822 may include computer program code (software). At least one processor and at least one memory storing instructions may provide components for providing or causing execution of any of the methods and / or blocks described above.
[0165] A processor is coupled to memory 9820. The processor is configured to read data from and write data to memory 9820. Memory 9820 may include one or more memory cells. Memory cells may be volatile or non-volatile. It should be noted that one or more non-volatile memory cells and one or more volatile memory cells may be present, or alternatively, one or more non-volatile memory cells may be present, or alternatively, one or more volatile memory cells may be present. Volatile memory may be, for example, random access memory (RAM), dynamic random access memory (DRAM), or synchronous dynamic random access memory (SDRAM). Non-volatile memory may be, for example, read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), flash memory, optical storage device, or magnetic storage device. Generally, memory may be referred to as a non-transitory computer-readable medium. The term "non-transitory" as used herein is a limitation on the medium itself (i.e., tangible, not tactile), rather than a limitation on the persistence of data storage (e.g., RAM vs. ROM). Memory 9820 stores computer-readable instructions that are executed by the processor. For example, non-volatile memory stores computer-readable instructions, while the processor uses volatile memory for temporary storage of data and / or instructions to execute instructions.
[0166] The computer-readable instructions may have been pre-stored in memory 9820, or alternatively or additionally, they may be received by the device via an electromagnetic carrier signal, and / or copied from a physical entity such as a computer program product. Execution of the computer-readable instructions causes device 9800 to perform one or more of the functions described above.
[0167] The memory 9820 can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and / or removable memory. The memory may include a configuration database for storing configuration data, such as a current list of neighboring cells, and, in some example embodiments, the structure of frames used in detected neighboring cells.
[0168] The device 9800 may also include or be connected to a communication interface 9830, such as a radio unit, which includes hardware and / or software for establishing a communication connection with one or more wireless communication devices according to one or more communication protocols. The communication interface 9830 includes at least one transmitter (Tx) and at least one receiver (Rx), which may be integrated into the device 9800 or connected to it. The communication interface 9830 may provide components for performing some blocks and / or functions (e.g., transmitting and receiving) of one or more of the above-described example embodiments. The communication interface 9830 may include one or more components, such as a power amplifier, a digital front-end (DFE), an analog-to-digital converter (ADC), a digital-to-analog converter (DAC), a frequency converter, a modulator (demodulator), and / or an encoder / decoder circuit system, which are controlled by corresponding control units.
[0169] Communication interface 9830 provides the device with radio communication capabilities for communication in a wireless communication network. For example, the communication interface may provide a radio interface to one or more UEs 100 and UE 102. Device 9800 may also include or be connected to another interface toward core network 110, such as a network coordinator device or AMF, and / or include or be connected to access node 104 of the wireless communication network.
[0170] The apparatus 9800 may also include a scheduler 9840 configured to allocate radio resources. The scheduler 9840 may be configured together with the communication control circuitry system 9810 or may be configured separately.
[0171] It should be noted that device 9800 may also include Figure 26 Various components are not shown. These various components can be hardware components and / or software components.
[0172] Figure 27Examples of apparatuses combined with user devices / user equipment according to some embodiments of the present invention are illustrated. Figure 27 The illustration shows apparatus configured to perform the functions described above in conjunction with a user device / user equipment. Each apparatus 500 may include one or more communication control circuitry systems, such as at least one processor 502; and at least one memory 504, which includes one or more algorithms 503, such as computer program code (software), wherein the at least one memory and the computer program code (software) are configured, together with the at least one processor, to cause the apparatus to perform any of the exemplary functions of the user device. The apparatus may also include different communication interfaces 501 and one or more user interfaces 501'.
[0173] refer to Figure 27 At least one communication control circuitry in device 500 is configured to perform application-specific computational solutions, perform computational matrix operations such as vector-matrix multiplication, implement a machine learning accelerator independently or in conjunction with other accelerators (such as in-memory computation), or perform the aforementioned operations via one or more circuitry systems. Figures 1 to 24 The described function.
[0174] Figure 28 Examples of devices combined with node elements according to some embodiments of the present invention are illustrated. Figure 28 The illustration shows a device configured to perform the functions described above in conjunction with node elements. Each device 600 may include one or more communication control circuitry systems, such as at least one processor 602; and at least one memory 604, which includes one or more algorithms 603, such as computer program code (software), wherein the at least one memory and the computer program code (software) are configured, together with the at least one processor, to cause the device to perform any of the exemplary functions of the node element. The device may also include different communication interfaces 601 and one or more user interfaces 601'.
[0175] refer to Figure 28 At least one communication control circuitry in device 600 is configured to perform application-specific computational solutions, perform computational matrix operations such as vector-matrix multiplication, implement a machine learning accelerator independently or in conjunction with other accelerators (such as in-memory computation), or perform the aforementioned operations via one or more circuitry systems. Figures 1 to 24 The described function.
[0176] Figure 29 Examples of apparatuses combined with base stations according to some embodiments of the present invention are illustrated. Figure 29The illustration shows an apparatus configured to perform the functions described above in conjunction with a base station. Each apparatus 700 may include one or more communication control circuitry systems, such as at least one processor 702; and at least one memory 704, which includes one or more algorithms 703, such as computer program code (software), wherein the at least one memory and the computer program code (software) are configured, together with the at least one processor, to cause the apparatus to perform any of the exemplary functions of the base station. The memory 704 may include a database for storing various information, such as contact information and maintenance requirements for the apparatus. The apparatus may also include various communication interfaces 701.
[0177] refer to Figure 29 At least one communication control circuitry in device 700 is configured to perform application-specific computational solutions, perform computational matrix operations such as vector-matrix multiplication, implement a machine learning accelerator independently or in conjunction with other accelerators (such as in-memory computation), or perform the aforementioned operations via one or more circuitry systems. Figures 1 to 24 The described function.
[0178] The device can be any type of electronic device, such as, but not limited to, user equipment and other electronic devices that may require such a memory device. The device can be any electrical device that can be connected to an access network and configured to wirelessly connect to the access network components (e.g., access network devices) providing the cell on one or more communication channels, including one or more control channels. The physical link from the device to the access network and then to the core network is referred to as an uplink or reverse link, while the physical link to the device is referred to as a downlink or forward link. By way of example and not limitation, the device can be referred to as a served device, a downlink device, a mobile device, a terminal device, a communication device, a user equipment (UE), a subscriber station (SS), a portable subscriber station, a mobile station (MS), or an access terminal (AT). A non-limiting list of examples of the device, or the objects that the device may include, or the objects that the device may include, includes mobile phones, cellular phones, smartphones, Voice over Internet Protocol (VoIP) phones, wireless local loop phones, devices using wireless modems, portable computers, desktop computers, laptop embedded devices (LEE), laptop mounted devices (LME), smart devices, multimedia devices, image capture terminal devices (such as digital cameras), gaming terminal devices, music storage and playback devices, drones, vehicles, automated guided vehicles, autonomous connected vehicles, in-vehicle wireless terminal devices, wireless endpoints, Internet of Things (IoT) devices, Industrial Internet of Things (IIoT) devices, devices operating in industrial and / or automated processing chain environments, consumer electronics devices, consumer IoT devices, mobile robots, mobile robotic arms, sensors, surveillance cameras, electronic health-related devices, medical monitoring devices, such as medical devices for remote surgery, wearable devices (such as smartwatches, smart rings, head-mounted displays (HMDs)), personal devices, etc. The device may also be part of a group of devices that are considered a device (i.e., a mobile device) by a wireless network.
[0179] The device can be any access network component of the access network. The access network domain can be based on any type of access network, such as cellular access networks (e.g., 5G, 5G Advanced, 6G, etc.), non-terrestrial networks, traditional cellular radio access networks (e.g., 4G or older generation networks), or non-cellular access networks (e.g., wireless LANs), or any combination thereof. To provide radio access, the access network includes access network components, such as access network devices or access equipment. An access equipment component can provide one or more cells, each cell may have different cell accessibility, but one cell is provided by one access equipment. However, overlapping cells can exist, such as macrocells provided by an access equipment that works in conjunction with access nodes providing smaller cells (e.g., microcells, femtocells, or picocells) that at least partially overlap within the macrocell. There are a wide variety of access network components. A non-limiting list of examples of access network components includes different types of base stations, such as eNBs, gNBs, split gNBs, transmit / receive points, network control repeaters, nodes operatively coupled to one or more remote radio heads, satellites, donor nodes in integrated access and backhaul (IAB), fixed IAB nodes, mobile IAB nodes mounted on vehicles, etc. At least some of the devices in the access network can provide an abstraction platform to decouple the abstraction of network functions from the processing hardware.
[0180] Furthermore, it should be noted that some components can be multi-domain components. For example, a device component can also provide services to other device components, i.e., it can also operate as an access network component, such as a relay node, a mobile IAB node, or a mobile terminal portion of an IAB node. Therefore, the term "mobile device" in this document is used for a device component or device component function in a multi-domain component, and the term "access network device" is used for an access network component or access network component function in a multi-domain component.
[0181] Replacing the traditional buffer amplifier-based resistive ACIM method with a direct-drive approach offers several advantages. First, there are potential advantages in area, power, and speed. Second, the direct-drive approach eliminates various drawbacks of analog amplifiers, including limited output voltage range, increased noise, and limited linearity. Third, the direct-drive ADC is far less sensitive to its feedback network than a direct-drive amplifier, making the ACIM core more easily supportive of various extensions and operating modes. Some signal processing operators, such as ReLU, comparators, and memory, are easier to implement in the digital domain (supported by the direct-drive approach) than in the analog domain (supported by traditional methods). Finally, the proposed direct-drive approach also supports negative weights.
[0182] ACIM accelerators can be used for machine learning and advanced computing in mobile network base stations. However, machine learning accelerators also have wide applications outside of mobile networks. ACIM accelerators can be implemented on integrated circuits.
[0183] As used in this application, the term "resistive memory" may refer to one or more of the following: a resistive memory may be the same (resistor / memory) element, that is, the memory element itself also acts as a resistor, and a resistive memory may include one or more of the following, for example: memristor, FeFET, RRAM, MRAM, OxRAM, CNTRAM.
[0184] The term "memory element" as used in this application may refer to one or more of the following: a memory element may include, for example, SRAM, DRAM, non-volatile flash memory, or memory based on, for example, RRAM.
[0185] The term "resistive element" as used in this application may refer to one or more of the following: a resistive element may include, for example, a resistor, a memristor, a switch-controlled resistor, a switch-controlled resistor in an R2R network, or a varactor diode.
[0186] The term "activation function" as used in this application may refer to one or more of the following: an activation function may include, for example, a linear activation function, a rectified linear unit, a ReLU, a binary step, a Heaviside activation function, a logic activation function (several related activation functions are listed herein), or any other type of known activation function.
[0187] As used in this application, the term "circuit system" may refer to one or more or all of the following: a) a hardware circuit implementation only (such as an implementation only in analog and / or digital circuit systems); and b) a combination of hardware circuits and software, such as (if applicable): i) a combination of (multiple) analog and / or digital hardware circuits with software / firmware, and (multiple) hardware processors having software (including (multiple) digital signal processors, software, and (multiple) portions of memory, which work together to cause a device (such as a mobile phone) to perform various functions); and c) (multiple) hardware circuits and / or (multiple) processors, such as (multiple) microprocessors or portions of (multiple) microprocessors, which require software (e.g., firmware) to operate, but may be absent when the software is not required to operate.
[0188] This definition of "circuit system" applies to all uses of the term in this application, including in any claim. As another example, as used in this application, the term "circuit system" also covers only the implementation of hardware circuitry or a processor (or processors) or portions thereof, and their accompanying software and / or firmware. For example, if applicable to a particular claim element, the term "circuit system" also covers baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices.
[0189] The techniques and methods described herein can be implemented in various ways. For example, these techniques can be implemented in hardware (one or more devices), firmware (one or more devices), software (one or more modules), or a combination thereof. For hardware implementation, the apparatus(s) of the example embodiments can be implemented within one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to perform the functions described herein, or a combination thereof. For firmware or software, implementation can be achieved by modules (e.g., processes, functions, etc.) of at least one chipset that perform the functions described herein. Software code can be stored in memory cells and executed by a processor. Memory cells can be implemented within the processor or external to the processor. In the latter case, as is known in the art, memory cells can be communicatively coupled to the processor in various ways. Furthermore, those skilled in the art will understand that the components of the systems described herein can be rearranged and / or supplemented by additional components to facilitate the implementation of the various aspects described, etc., and they are not limited to the precise configurations illustrated in the given figures.
[0190] It will be apparent to those skilled in the art that, with advancements in technology, the inventive concept can be implemented in various ways within the scope of the claims. Embodiments are not limited to the exemplary embodiments described above, but can vary within the scope of the claims. Therefore, all words and expressions should be interpreted broadly, and they are intended to illustrate rather than limit the embodiments.
Claims
1. An apparatus comprising: At least one processor; as well as At least one memory storing instructions, which, when executed by the at least one processor, cause the means to at least: Provide an analog signal to at least one input for calculating the resistance matrix in the analog memory included in the device; At least one direct-drive analog-to-digital converter receives digital signals from the output of the resistance matrix calculated from the at least one analog memory. The at least one direct-drive analog-to-digital converter is a direct-drive successive approximation analog-to-digital converter, and The at least one of the direct-drive successive approximation analog-to-digital converters has a self-biasing circuit system comprising: a comparator in the input of the direct-drive successive approximation analog-to-digital converter, the comparator being arranged to compare a reference node voltage with an input node voltage from an output of an intrinsic digital-to-analog converter, the intrinsic digital-to-analog converter receiving an output signal of the direct-drive successive approximation analog-to-digital converter having reverse polarity as an input signal.
2. The apparatus of claim 1, wherein the provided analog signal is converted from a digital signal by at least one of the inputs of the at least one analog memory that calculates the resistance matrix.
3. The apparatus according to claim 1 or 2, wherein the calculation of the resistance matrix in the at least one analog memory is implemented using at least one resistive memory, the at least one resistive memory comprising: At least one resistive element within the at least one resistive memory.
4. The apparatus according to claim 1 or 2, wherein the calculation of the resistance matrix in the at least one analog memory is implemented using at least one memory element and at least one resistive element.
5. The apparatus of claim 1, wherein the at least one direct-drive successive approximation analog-to-digital converter is arranged to provide the output signal of the comparator to at least one other digital-to-analog converter via synchronous successive approximation logic of the at least one direct-drive successive approximation analog-to-digital converter.
6. The apparatus of claim 1, wherein the at least one direct-drive successive approximation analog-to-digital converter is arranged to provide a logic block output signal to at least one other digital-to-analog converter.
7. The apparatus of claim 1, wherein the at least one direct-drive successive approximation analog-to-digital converter is arranged to provide the output signal of the intrinsic digital-to-analog converter to at least one load.
8. The apparatus according to any one of claims 5 to 7, comprising: An activation function between the output of the synchronous successive approximation logic and the input of the intrinsic digital-to-analog converter, or between the output of the synchronous successive approximation logic and the input of the at least one other digital-to-analog converter.
9. The apparatus according to any one of claims 1 to 4, wherein at least one digital output of a direct-drive analog-to-digital converter from the output of the calculation resistor matrix in the first analog memory is connected to at least one digital input of a direct-drive digital-to-analog converter from the input of the calculation resistor matrix in the second analog memory.
10. The apparatus according to any one of claims 1 to 4, wherein the output column of the resistance matrix calculated in the first analog memory and the output column of the resistance matrix calculated in the second analog memory are combined and connected to the at least one direct-drive analog-to-digital converter in the output of the resistance matrix calculated in the second analog memory.
11. The apparatus of claim 1, wherein the output of the at least one intrinsic digital-to-analog converter that directly drives the successive approximation analog-to-digital converter is connected back to the input row of the at least one analog memory for calculating the resistance matrix, for matrix inverse calculation.
12. The apparatus of claim 1, wherein the output of the at least one intrinsic digital-to-analog converter that directly drives the successive approximation analog-to-digital converter is connected back to the input row of the at least one analog memory via at least one comparator for Ising machine calculation.
13. The apparatus according to any one of the preceding claims, wherein the apparatus is an electronic device, a user device, a node element, a base station, an access network component, a served device, a downlink device, a mobile device, a terminal device, a communication device, a user equipment, a subscriber station, a portable subscriber station, a mobile station, or an access terminal.
14. A non-transitory computer-readable medium comprising program instructions that, when executed by a device, cause the device to perform at least the following: Provide an analog signal to at least one input for calculating the resistance matrix in the analog memory included in the device; At least one direct-drive analog-to-digital converter receives digital signals from the output of the resistance matrix calculated from the at least one analog memory. The at least one direct-drive analog-to-digital converter is a direct-drive successive approximation analog-to-digital converter, and The at least one direct-drive successive approximation analog-to-digital converter described herein has a self-biasing circuit system, the self-biasing circuit system comprising: The comparator in the input of the direct-drive successive approximation analog-to-digital converter is arranged to compare a reference node voltage with an input node voltage from the output of an intrinsic digital-to-analog converter, which receives the output signal of the direct-drive successive approximation analog-to-digital converter with reverse polarity as an input signal.