Numerical comparator, B2S converter, random calculation integrated circuit and design method thereof
By integrating a binary digital comparator within DRAM, and using capacitor charge and discharge and sense amplifiers to achieve B2S conversion, the high area overhead and data movement problems of SNG in DRAM-PIM architecture are solved, and the computing performance and efficiency are improved.
Patent Information
- Application Number
- CN202510553010.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-29
AI Technical Summary
In the existing DRAM-PIM architecture, the high area overhead and frequent data movement problems of random number generators (SNGs) limit the performance and efficiency of in-memory computing, especially the insufficient design of binary digital comparators leads to poor computing performance.
The binary digital comparator is directly integrated into the DRAM, and the capacitance charge and discharge and sense amplifier (SA) functions of DRAM can realize time-weighted B2S conversion, eliminate the need for external SNG modules, and reduce hardware costs and data transmission.
Significantly reduce hardware costs, reduce data transmission, improve the accuracy and efficiency of in-memory computing, reduce latency, and enhance system reliability and compatibility.
Smart Images

Figure CN120469663A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of integrated circuits, and in particular relates to a DRAM-based storage and computing integrated numerical comparator, a B2S converter, a random computing integrated circuit, and a design method thereof. Background Art
[0002] General Matrix-Vector Multiplication (GEMV) is a core operation in many computing workloads, including scientific computing, machine learning, and large language models (LLMs). In high-performance computing, GEMV is a critical, compute-intensive operation and is often used to optimize the performance of matrix operations. For example, efficient matrix-vector multiplication can be implemented on GPUs or multi-core CPUs to accelerate scientific computing and engineering applications.
[0003] However, GEMV is still inherently limited by memory and is often constrained by data transfer bottlenecks and memory access latency. Processing in memory (PIM) aims to alleviate these problems by integrating computation directly into the memory array and reducing energy-intensive data transfer. Dynamic random access memory (DRAM) has become a hot topic for developing DRAM-based processing in memory (DRAM-PIM) architectures due to its maturity, scalability, and high storage density. Although known prototypes such as Samsung's FIMDRAM[7], Hynix's AiM[8], and UPMEM[6] have demonstrated progress in this field, the DRAM-PIM architecture itself still has limitations, especially in terms of computational performance. For example, in DRISA[5], multiplication operations take up to 1600 nanoseconds, which limits its application in computationally intensive tasks.
[0004] To address this issue, Samsung has introduced stochastic computing (SC) into DRAM-based in-memory processing (DRAM-PIM), improving peak performance by up to 4 times.
[0005] However, despite the enormous potential of integrating SC with DRAM, existing research has primarily focused on optimizing random operations (such as random multiplication) while neglecting a key component: the random number generator (SNG), which typically accounts for 90% of the area in the entire SC system. Current implementations typically place the SNG (B2S module) near the memory rather than within it, resulting in increased area overhead and additional data movement. This placement undermines the advantages of PIM, as the external location of the SNG negates the benefits of reduced data transfer and energy consumption. A core challenge lies in the lack of efficient in-memory SNG designs, which limits the full potential of SC in DRAM-PIM-based systems.
[0006] To address this limitation, this disclosure proposes a novel approach to integrate the SNG directly into DRAM, eliminating the need for an external SNG module. Specifically, we focus on the binary digital comparator, a key component in the SNG. Unlike existing approaches that typically require hardware modifications, the design of this disclosure leverages the inherent characteristics of DRAM, such as capacitor charging and discharging, and the sense amplifier (SA), to efficiently implement these functions without modifying the DRAM circuitry itself or adding any additional circuitry. This approach not only reduces area overhead but also minimizes data handling, thereby improving the overall accuracy and efficiency of the in-memory SC. Summary of the Invention
[0007] One embodiment of the present disclosure is a numerical comparator, the comparator comprising a DRAM memory cell array, wherein the memory cells of the memory cell array are connected to form an array through word lines and bit lines.
[0008] The first number involved in the comparison is stored in storage units of different word lines on the same bit line according to the bit sequence, and the different word lines on the same bit line constitute a first word line group.
[0009] The second number involved in the comparison is stored in a storage unit of a word line different from any word line in the first word line group on the same complementary bit line (BLB) corresponding to the bit line (BL), and the word lines different from any word line in the first word line group constitute a second word line group.
[0010] The bit line and the complementary bit line are connected to the input end of the sense amplifier corresponding to the bit line.
[0011] One of the embodiments of the present disclosure is a B2S converter (SNG), which includes the numerical comparator described above, wherein the first number is an input binary number and the second number is a random number after conversion.
[0012] The disclosed embodiments address the high area overhead and frequent data movement issues of the B2S module in the existing DRAM-PIM architecture. By completely implementing the B2S function within the DRAM core array, external comparator and random number generator modules are eliminated, thereby significantly reducing hardware costs, reducing data transmission, and improving the energy efficiency and performance of in-memory computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily apparent by reading the following detailed description with reference to the accompanying drawings, in which several embodiments of the present invention are shown by way of example and not limitation, in which:
[0014] Figure 1 Schematic diagram of the three basic modules of existing stochastic computing (SC).
[0015] Figure 2 A schematic diagram comparing the present disclosure with the existing random computing and DRAM-PIM architecture according to one embodiment of the present disclosure.
[0016] Figure 3 A schematic diagram comparing chip area ratios of the present disclosure and existing circuit designs according to one embodiment of the present disclosure.
[0017] Figure 4 Schematic diagram comparing delays of different design architectures under different GEMV dimensions according to one embodiment of the present disclosure.
[0018] Figure 5 Schematic diagram of a B2S conversion circuit according to one embodiment of the present disclosure.
[0019] Figure 6 A schematic diagram comparing SPICE simulation results of the sense amplifier (SA) performance under different bit line configurations according to one embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] In the existing in-memory computing (PIM) architecture based on dynamic random access memory (DRAM), the binary to random number (B2S) conversion module is the core component for implementing stochastic computing (SC).
[0021] Figure 1 This is an overview of stochastic computing (SC): (a) binary to stochastic (B2S) conversion, (b) the computation process including AND gate (AND)-based multiplication operation (MUL) and multiplexer (MUX)-based addition operation (ADD), and (c) stochastic to binary (S2B) conversion. Figure 1 This figure illustrates the binary-to-random (B2S) conversion module in stochastic computing (SC). A traditional B2S module typically consists of a random number generator (RNS) and a comparator. The comparator requires dedicated logic circuits, such as CMOS comparators or lookup tables, to compare binary values with random numbers and generate a random bit stream. The English terms used in the figure have the following meanings:
[0022] 1. Random number generator (Random number source RNS) - Generates a string of random numbers (shown in binary form in the figure) for comparison with a given binary number.
[0023] 2. Binary number - A binary number that is compared with the random number generated by the random number generator.
[0024] 3. Comparator (R vs. B) - Receives a random number (R) from the random number generator and a specific binary number (B). Its function is to compare the two numbers and generate an output signal (S) based on the comparison result. The binary data length is 3 bits, and the random bit stream length is 2 3 =8 bits: The comparator's truth table shows the output under different comparison conditions: If the random number R is less than the binary number B (for example, R = 001, B = 100), the output S is 1. If the random number R is greater than or equal to the binary number B (for example, R = 111, B = 100), the output S is 0. This comparison process is the key step in the random number generator to produce random bit stream data.
[0025] 4. Stochastic number (or stochastic bit stream)—This is the output of a comparator, also known as a "stochastic number." It represents data based on the probability of a "1" bit. For example, "1011 0010" represents "4 / 8," and "0000 1111" also represents "4 / 8," regardless of the position of the "1." Another example is "0011 1100 00001111" represents "8 / 16." This data reflects the result of comparing a random number with a given binary number.
[0026] 5. AND-based MUL - "MUL" is the abbreviation of "Multiplication", which means "multiplication operation".
[0027] 6. MUX-based ADD—ADD stands for “addition,” and “MUX” is the abbreviation for “Multiplexer.”
[0028] 7. Stochastic to Binary (S2B) Conversion - "S2B" is the abbreviation of "Stochastic-to-Binary", which means "stochastic to binary conversion", converting the random results generated by random calculations into a deterministic binary number.
[0029] Figure 1 AND gate-based multiplication (MUL) is a method for implementing multiplication in random calculations, using AND gates to perform random multiplication. Multiplexer (MUX)-based addition (ADD) uses a multiplexer (MUX) to implement random addition. Random-to-binary (S2B) conversion is the process of converting the random results of random calculations back into binary numbers.
[0030] Figure 1 The B2S module shown in Figure 1 is a key component in random computing, generating a random bit stream by comparing a randomly generated number with a predetermined binary number. The entire random computing process also includes random-to-binary conversion and complex computational operations based on these random bit streams.
[0031] Furthermore, the existing SCOPE architecture deploys the B2S module in the near memory area near the DRAM storage unit, e.g. Figure 2 As shown in (a). Figure 2 are stochastic computing architectures: (a) the SCOPE architecture with binary-to-stochastic (B2S) and random-to-binary (S2B) modules located near the memory bank [3]; and (b) the proposed design that embeds the binary-to-stochastic (B2S) module and random adder into the dynamic random access memory (DRAM) memory bank. Figure 2 The English terms in include:
[0032] DRAM – Dynamic Random Access Memory
[0033] Bank——Memory
[0034] SA - Sense Amplifier
[0035] Binary ADD——Binary addition
[0036] Stochastic MUL——Stochastic Multiplication
[0037] B2S——Binary to Random Number
[0038] S2B——Random Number to Binary
[0039] Near Bank Hardware
[0040] data_in——input data
[0041] data_out——output data
[0042] Stochastic ADD——random addition
[0043] Shifter
[0044] Figure 2 This includes the location and integration method of the binary-to-random number (B2S) module and the random-to-binary (S2B) module in the random computing architecture design. Figure 2(a) The SCOPE architecture deploys the B2S module in the near-memory area near the dynamic random access memory (DRAM) storage unit. This reduces the distance data must travel during computation, thereby reducing latency and energy consumption. The processing flow includes:
[0045] (1) The input binary data (Binary data_in) is first converted into a random bit stream by the B2S module.
[0046] (2) The converted random bit stream participates in the random multiplication operation (Stochastic MUL).
[0047] (3) The result of the random multiplication is then converted back into binary form by the S2B module.
[0048] (4) The final binary result (Binary data_out) is output from the memory bank.
[0049] Figure 2 (b) The B2S module and stochastic adder (Stochastic ADD) are further embedded directly into the DRAM memory bank. This integration method can more closely combine storage and computing, further improving computing efficiency and reducing energy consumption. The processing flow includes:
[0050] (1) The input binary data is also first converted into a random bit stream through the B2S module.
[0051] (2) The converted random bit stream directly participates in the random multiplication operation (Stochastic MUL) in the memory.
[0052] (3) The result of random multiplication is converted back to binary form through the S2B module and the shifter.
[0053] (4) The final binary result is also output from the storage body.
[0054] Figure 2 The invention discloses that the performance of random computing is optimized by embedding computing modules (especially binary-to-random number module and random adder) into memory.
[0055] Figure 3The paper reveals the following: (a) The binary-to-random (B2S) module in the SCOPE architecture accounts for 13.6% of the total area; (b) the proposed design reduces the area share of the binary-to-random (B2S) and random-to-binary (S2B) modules to less than 1%. "Area share" refers to the proportion of the total chip area occupied by a specific functional module in integrated circuit (IC) design. This is because chip area directly affects chip cost, power consumption, heat dissipation requirements, performance, and reliability. Figure 3 The English terms in include:
[0056] B2S——Binary to Random Number Module
[0057] S2B——Random Number to Binary Module
[0058] DRAM Bank——Dynamic Random Access Memory Bank
[0059] Proposed——The architecture proposed in this disclosure
[0060] Shifter
[0061] Figure 3 (a) In the SCOPE architecture, the binary-to-random-number (B2S) module accounts for 13.6% of the area, while the dynamic random access memory (DRAM) bank accounts for 86%. The S2B module accounts for less than 1%. This indicates that the B2S module consumes a relatively large amount of hardware resources in the SCOPE architecture design. Because the B2S module occupies a considerable area, it limits the space available for data storage in the memory, affecting overall storage density and cost-effectiveness. Figure 3 (b) is the design proposed in this disclosure. The area share of the binary-to-random (B2S) module and the random-to-binary (S2B) module is reduced to less than 1%, while the area share of the DRAM memory is increased to 99%. This design significantly reduces the area occupied by the B2S and S2B modules by integrating them into the memory. This integration method greatly improves the storage efficiency of the memory because it allows more space to be used for data storage rather than for conversion modules.
[0062] exist Figure 3In the figure, the area share comparison shows the proportion of the random computing modules (B2S and S2B) in the overall chip area in different design schemes. These modules occupy a relatively large area (13.6%) in the SCOPE architecture, but in the proposed design, these modules are integrated into the dynamic random access memory (DRAM) memory bank, occupying a significantly smaller area (less than 1%). This design improvement can bring higher storage density and possible performance advantages while reducing cost and power consumption.
[0063] Figure 4 It is revealed that data transmission constitutes the main overhead in random computing. Figure 4 The English terms in include:
[0064] Latency
[0065] ns——nanosecond (time unit)
[0066] Data Movement
[0067] Calculation
[0068] GEMV dimension
[0069] Proposed——(method or architecture)
[0070] Figure 4 The figure reveals the latency comparison between the existing method (SCOPE) and the method proposed in this disclosure under different GEMV (General Matrix-Vector) dimension configurations. The GEMV dimension in the figure is defined by three parameters: [m,n]x[n,1], which represent the number of rows, columns and dimensions of the matrix, respectively. The latency is decomposed into two parts: data movement (blue) and calculation (yellow). The vertical axis in the figure represents the latency of the operation in nanoseconds (ns). The lower the latency, the better the performance. The blue bar graph represents the time required for data to move during the calculation process. The yellow bar graph represents the time required to actually perform the calculation operation. The figure uses dotted lines to separate the different GEMV dimensions and shows the latency comparison between the SCOPE method and the method proposed in this disclosure under each dimension:
[0071] 1. When the GEMV dimension is [m=8, n=8], the method proposed in this disclosure is 6.0 times faster than the SCOPE method.
[0072] 2. When the GEMV dimension increases to [m=8, n=16], the method proposed in this disclosure has a 6.8-fold speed increase.
[0073] 3. When the GEMV dimension is further increased to [m=16, n=16], the method proposed in this disclosure has a 7.7-fold speed increase.
[0074] 4. When the GEMV dimension is [m=16, n=32], the method proposed in this disclosure has an 8.5-fold speed increase.
[0075] 5. When the maximum GEMV dimension is [m=32, n=32], the method proposed in this disclosure has a 9.1-fold speed improvement.
[0076] therefore, Figure 4 The results reveal that data transfer (mobility) is the primary overhead in random computations, and that the proposed method significantly improves performance over existing methods (SCOPE) under different GEMV dimension configurations. This demonstrates that the proposed method can more effectively reduce latency and improve computational efficiency when processing large-scale matrix-vector multiplications.
[0077] This paper analyzes the existing solutions and summarizes the following problems with the existing solutions:
[0078] 1. High area overhead.
[0079] The area overhead of the binary-to-random (B2S) module is large. The B2S module is crucial for converting binary values into random bit streams, and in a stochastic computing (SC) system, its hardware cost accounts for up to 80% [4]. For example, in the SCOPE architecture ( Figure 3 ), the B2S module occupies 13.6% of the circuit area, which is approximately equivalent to one-sixth of a dynamic random access memory (DRAM) bank[3].
[0080] 2. Frequent data movement.
[0081] In the SCOPE framework ( Figure 2 (a)), the computation process requires multiple data transfers: (1) transferring random data to the dynamic random access memory (DRAM) core for multiplication (MUL); (2) transferring intermediate results for random to binary (S2B) conversion; (3) returning data to the DRAM memory bank for binary addition (ADD); and (4) transferring the final result. Figure 4 As shown, these frequent data transmissions significantly increase latency and energy consumption.
[0082] 3. Limited scalability.
[0083] Traditional comparators are difficult to integrate directly into the DRAM core array, which limits the deep integration and performance improvement of SC in DRAM-PIM.
[0084] The existing DRAM structure mainly includes the following units:
[0085] (1) Memory cell array
[0086] Each DRAM memory cell consists of a transistor and a capacitor (1T1C structure). The transistor acts as an access switch, and the capacitor is used to store data.
[0087] Memory cells are arranged in rows and columns to form a memory array. The transistor gates of each row of memory cells are connected to a word line, and each column of memory cells is connected to a bit line. The word line controls access to memory cells in a specific row, and data is read or written via the bit line.
[0088] (2) Auxiliary circuit
[0089] Sense Amplifier - used to detect tiny voltage changes on the bit line to read the data in the storage cell. In a standard read operation, the bit line BL and the complementary bit line BLB are first precharged to VDD / 2, then the word line WL is turned on, and the storage capacitor and the BL capacitor begin to share charge. If the storage capacitor stores "1", the BL voltage will increase from VDD / 2. At this time, the BLB potential is still VDD / 2. The SA will compare the two potentials and continue to raise the BL voltage to "1", that is, read "1". If the storage capacitor stores "0", the BL voltage will drop from VDD / 2. At this time, the BLB potential is still VDD / 2. The SA will compare the two potentials and continue to pull the BL voltage down to "0", that is, read "0".
[0090] According to one or more embodiments, the method proposed in the present disclosure utilizes the charge sharing process between the storage capacitor and the bit line (BL) capacitor. Figure 5 As shown in (a), charge sharing is a gradual process, not a transient one. Therefore, the industry typically limits this charge sharing time to ensure sufficient charge sharing. For example, Micron's DDR5 has a t_RCD of 16 nanoseconds.
[0091] When two binary numbers need to be compared (two 8-bit INT8 numbers), they will be stored in 8 lines of WL on the same BL / BLB from high bit to low bit, such as Figure 5As shown in Figure 1. Two numbers are stored in BL and BLB, respectively. Then, the control transistors of eight rows are turned on for charge sharing. It's important to note that the bits stored in different WLs have different weights for the binary digits. Therefore, the turn-on times of the eight rows of WLs are weighted incrementally, with the higher-order bits having a greater impact, meaning they are allocated a longer time. Finally, the DRAM's sense amplifiers (SAs) compare and amplify the voltages after charge sharing, creating a time-weighted comparator for comparing two binary values without requiring additional hardware.
[0092] At this point, a comparison operation is completed, that is, B2S ( Figure 1 One bit of the bitstream generated in (a) is generated. The hardware required for this operation is a security element (SA) and eight rows of write-loops (WLs) corresponding to a BL / BLB pair within the DRAM. The DRAM architecture also supports parallel operation of many SAs. For example, 256 of these SAs can be used simultaneously to convert an INT8 number into a 256-bit random bitstream.
[0093] Figure 5 In this paper, we include (a) the charging and discharging dynamic process of the capacitor and (b) the proposed time-weighted comparator in dynamic random access memory (DRAM). Figure 5 Chinese and English terms include:
[0094] Voltage——voltage, Time[ns]——time[nanoseconds], V_charge(t)——charging voltage(t)
[0095] V_discharge(t)——discharge voltage (t), C——capacitance, t——time, RC——RC time constant, V_0——initial voltage, t_8——time 8, t_1——time 1,
[0096] SA - Sense Amplifier
[0097] WL——Word Line
[0098] BL / BLB——Bit Line / Bit Line Bar
[0099] time-weight comparator proposed—— proposed time-delay weighted comparator
[0100] Binary Number
[0101] Random Number
[0102] Figure 5(b) shows the circuit structure of the charge sharing process in a DRAM memory array. 1. The memory array address selector selects the control signal for a specific memory array. In DRAM, data is stored in a matrix consisting of rows and columns, and the memory array address selector is responsible for selecting a specific column for read or write operations.
[0103] 2. Word lines (WLs) are wires connected to memory cells and used to transmit data. In the figure, WLs are represented by 8 rows, each row corresponding to a memory cell.
[0104] 3. During the charge sharing process, as shown in the figure, two binary numbers are stored in different rows WL on the same BL or BLB. For example, one binary digit (8 bits) is stored in WL1-WL8 on the BL, and another binary digit (8 bits) is stored in WL1-WL8 on the BLB (not limited to WL1-WL8, it can be WLn-WLn+8 / if the original binary digit is 4 bits, then WLn-WLn+4 can be used). By turning on the control transistor (usually controlled by the row address selector), the bits stored in these rows share charge with the BL capacitor. Since the bits stored under different WLs have different weights for binary digits, the weight of the high-order bits is greater. Therefore, during the charge sharing process, this method allocates a longer time for the high-order bits to ensure more charge sharing.
[0105] 4. Sense amplifier (SA). After charge sharing is completed, the DRAM's sense amplifier compares and amplifies the shared voltage, thereby achieving a comparison of two binary values and creating a time-weighted comparator without the need for additional hardware.
[0106] In this way, a comparison operation can be completed to obtain B2S( Figure 1 (a)) generates a bit of the bit stream. The required hardware includes a pair of BL / BLBs and a storage array address selector inside the DRAM, as well as 8 rows of WLs. At the same time, the structure of the DRAM in the embodiment of the present disclosure supports the parallel operation of many storage array address selectors. For example, 256 of the above structures can be operated simultaneously to complete the conversion of an INT8 number to a 256-bit random bit stream. Therefore, Figure 5 (b) shows a circuit design that uses the internal structure of DRAM to implement binary number comparison. Through the combination of charge sharing and sense amplifiers, efficient binary number comparison operations can be achieved without adding additional hardware.
[0107] Figure 5 (a) Reveals the curve of the voltage change over time during the charging and discharging process of the capacitor. The figure shows two key times: t8 and t1, which correspond to specific time points in the charging and discharging process, respectively.
[0108] The charging process is time t8, and the capacitor voltage gradually rises from the initial value V0 until it approaches 1V. This process can be described by the formula To describe, where C is the capacitance value (~140fF), RC is the RC time constant. Discharge process After time t1, the capacitor begins to discharge and the voltage gradually decreases from nearly 1V. The discharge process can be described by the formula To describe.
[0109] The present disclosure further verifies the reliability of the embodiments. This disclosure uses the TSMC 65nm process design library (PDK) for SPICE simulation, with simulation parameters of: Ccell = 15fF, CBL = 140fF, tbit1 = 8ns, tbit2 = 7ns, all the way to tbit8 = 1ns, and rise / fall time Tr = Tf = 0.1ns. This disclosure evaluates the accuracy of SA under three most stringent test scenarios, where the values in BL and BLB are very close, differing by only 1 bit: (1) Near minimum value: BLB stores "0000 0001" and BL stores "0000 0010" or "0000 0000", testing low value differences; (2) Middle value: BLB stores "01010101" and BL stores "0101 0110" or "0101 0100", testing intermediate accuracy; (3) Near maximum value: BLB stores "1111 1110" and BL stores "1111 1111" or "1111 1101", testing the robustness of high values.
[0110] Figure 6 The following are the SPICE simulation results of the sense amplifier (SA) performance for three 8-bit bit line (BL) and complementary bit line (BLB) configurations. Figure 6 The English terms in include:
[0111] BL Voltage——Bit line voltage
[0112] Time[ns]——time [nanoseconds]
[0113] BLB_stores - Complementary Bit Line Stores
[0114] BL_stores——Bitline stores
[0115] Figure 6Simulation results of the sense amplifier (SA) performance under three different 8-bit bitline (BL) and complementary bitline (BLB) configurations are presented. These simulation results verify the reliability of the sense amplifier under all scenarios, ensuring its ability to accurately compare the 8-bit binary values in the bitline (BL) and complementary bitline (BLB). The simulation results confirm the reliability of the SA under the three most demanding scenarios, ensuring its ability to accurately compare the 8-bit binary values in the BL and BLB. For an 8-bit input, the design generates an n-bit random bit stream by comparing n SAs in parallel, producing a probabilistic output without the need for additional comparators. By presenting voltage curves for different bitline configurations, this disclosure verifies the performance of the sense amplifier under various scenarios. The curves show how the voltage on the bitline changes over time during the read process and how time is controlled to achieve weighted comparison of different binary bits. This design leverages the native structure of DRAM, significantly reducing the need for additional hardware and improving efficiency. Through these test scenarios, the reliability and accuracy of the B2S conversion implemented in the sense amplifier are ensured.
[0116] Further through Figure 3 and Figure 4 , which verifies that the solution proposed in this disclosure not only saves the hardware area of B2S, but also greatly saves the cost of data transportation.
[0117] The present disclosure discloses a time-weighted B2S conversion: time-weighted voltage comparison is implemented through the Sense Amplifier (SA) within the DRAM, avoiding the use of an additional comparator module. Since the time-weighted SA voltage comparison method is used for B2S conversion, the B2S conversion implemented within the DRAM utilizes existing resources within the DRAM to replace traditional hardware. This allows random calculations to be performed directly in memory, reducing data transmission and external computing overhead. Therefore, the beneficial effects of the present disclosure include: 1. Reducing hardware overhead.
[0118] Traditional B2S converters typically require additional comparator modules, increasing hardware complexity and area overhead. However, the time-weighted internal B2S converter eliminates the need for additional comparator modules by utilizing the DRAM's internal sense amplifier (SA) for voltage comparison, significantly reducing hardware overhead. This approach fully utilizes the native DRAM structure, enabling B2S conversion without adding additional hardware.
[0119] 2. Improve energy efficiency.
[0120] Since no external hardware is required to perform the B2S conversion, the entire process occurs within the DRAM. This reduces energy consumption from unnecessary external circuitry and data transfer, thereby improving the overall energy efficiency of the system. In traditional designs, frequent data movement and external computation consume significant energy, but internal conversion avoids these issues.
[0121] 3. Reduce data transmission delay.
[0122] Traditional B2S conversion requires transferring data from DRAM to an external computing unit for processing, which not only increases latency but also leads to higher energy consumption. However, the time-weighted internal B2S converter completely integrates the conversion process within DRAM, reducing the need for data movement and significantly reducing latency. This is particularly important for high-performance computing and real-time processing tasks.
[0123] 4. Improve system reliability.
[0124] Traditional B2S converters can introduce additional points of failure or instability due to the complexity of external hardware. By performing the conversion directly within the DRAM, the entire process does not rely on additional external modules, which enhances the overall stability and reliability of the system. Utilizing the DRAM's inherent voltage comparison mechanism and time-weighted characteristics, more stable operation is achieved.
[0125] 5. Save chip area.
[0126] The absence of an additional B2S hardware module significantly reduces chip area. In traditional designs, the B2S module consumes a significant portion of the DRAM architecture, which can become a bottleneck in resource-constrained environments. The time-weighted internal B2S converter performs the conversion internally, significantly saving chip area and improving system area efficiency.
[0127] 6. Compatibility and scalability.
[0128] Because the converter utilizes existing resources within DRAM (such as the sense amplifier and BL / BLB capacitors), it is easily compatible with existing DRAM architectures without any hardware modifications. This makes this approach scalable and applicable to memory systems of different sizes and types.
[0129] Therefore, the present disclosure addresses this challenge by leveraging the sense amplifiers (SAs) of dynamic random access memory (DRAM) to achieve efficient in-memory binary-to-random (B2S) conversion. By utilizing the inherent voltage comparison mechanism in the sense amplifier (SA), the disclosed method eliminates the need for traditional comparator modules, enabling seamless integration of B2S functionality within dynamic random access memory (DRAM).
[0130] In summary, the present disclosure provides an efficient and feasible solution for modern DRAM-PIM architecture by reducing hardware overhead, improving energy efficiency, reducing latency, enhancing system reliability, saving chip area, and improving compatibility and scalability.
[0131] References:
[0132] [1] Z.Xia, J.Chen, Q.Huang, J.Luo, and J.Hu, "Neural synaptic plasticity-inspired computing: A high computing efficient deep convolutional neural network accelerator," IEEE Trans.Circuits Syst.I:Reg.Papers, vol.68, no.2, pp.728–740, 2020.
[0133] [2] Z.Chen, Y.Ma, and Z.Wang, "Hybrid stochastic-binary computing for low-latency and high-precision inference of cnns," IEEE Trans.Circuits Syst.I:Reg.Papers, vol.69, no.7, pp.2707–2720, 2022
[0134] [3]S.Li,A.O.Glova,X.Hu,P.Gu,D.Niu,K.T.Malladi,H.Zheng,B.Brennan,andY.Xie,“Scope:A stochastic computing engine for dram-based in-situaccelerator,”in 2018 51st Annual IEEE / ACM International Symposium onMicroarchitecture(MICRO).IEEE,2018,pp.696–709.
[0135] [4]X.Zhang,Y.Wang,Y.Zhang,J.Song,Z.Zhang,K.Cheng,R.Wang,and R.Huang,“Memory system designed for multiplyaccumulate(mac)engine based on stochasticcomputing,”in 2019International Conference on IC Design and Technology(ICICDT),2019,pp.1–4.
[0136] [5]S.Li,D.Niu,K.T.Malladi,H.Zheng,B.Brennan,and Y.Xie.DRISA:ADRAM-based Reconfigurable In-Situ Accelerator.In Proceedings of the 50th AnnualIEEE / ACM Inter-national Symposium on Microarchitecture,MICRO-50’17,pages 288–301,New York,NY,USA,2017.ACM.
[0137] [6]F.Devaux.The true Processing In Memory accelerator.In 2019 IEEEHot Chips 31Sympo-sium(HCS),pages 1–24,2019.
[0138] [7]Y.-C.Kwon,S.H.Lee,J.Lee,S.-H.Kwon,J.M.Ryu,J.-P.Son,O.Seongil,H.-S.Yu,H.Lee,S.Y.Kim,Y.Cho,J.G.Kim,J.Choi,H.-S.Shin,J.Kim,B.Phuah,H.Kim,M.J.Song,A.Choi,D.Kim,S.Kim,E.-B.Kim,D.Wang,S.Kang,Y.Ro,S.Seo,J.Song,J.Youn,K.Sohn,and N.S.Kim.25.4A 20nm 6GB Function-In-Memory DRAM,Based on HBM2 witha 1.2TFLOPS Programmable Computing Unit Using Bank-Level Parallelism,forMachine Learning Applications.In 2021IEEE International Solid-State CircuitsCon-ference(ISSCC),volume 64,pages 350–352,2021.
[0139] [8]S.Lee,K.Kim,S.Oh,J.Park,G.Hong,D.Ka,K.Hwang,J.Park,K.Kang,J.Kim,J.Jeon,N.Kim,Y.Kwon,K.Vladimir,W.Shin,J.Won,M.Lee,H.Joo,H.Choi,J.Lee,D.Ko,Y.Jun,K.Cho,I.Kim,C.Song,C.Jeong,D.Kwon,J.Jang,I.Park,J.Chun,and J.Cho.A 1ynm1.25V 8Gb,16Gb / s / pin GDDR6-based Accelerator-in-Memory supporting 1TFLOPS MACOperation and Various Activation Functions for Deep-Learning Applications.In2022IEEE International Solid-State Circuits Conference(ISSCC),volume 65,pages1–3,2022.
[0140] It should be understood that in the embodiments of the present invention, the term "and / or" merely describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent three possible situations: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0141] It is worth noting that although the foregoing content has described the spirit and principles of the present invention with reference to several specific embodiments, it should be understood that the present invention is not limited to the disclosed specific embodiments, and the division into various aspects does not mean that the features of these aspects cannot be combined. Such division is merely for the convenience of expression. The present invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A numerical comparator, characterized in that: The comparator includes a memory cell array, wherein the memory cells of the memory cell array are connected to form an array through word lines and bit lines. The first number involved in the comparison is sequentially stored in storage units of different word lines on the same bit line according to the bit order, and the different word lines on the same bit line constitute a first word line group. The second number involved in the comparison is sequentially stored in a storage unit of a word line different from any word line in the first word line group on the same complementary bit line corresponding to the bit line in bit order, and the word lines different from any word line in the first word line group constitute a second word line group. The bit line and the complementary bit line are connected to the input end of the sense amplifier corresponding to the bit line.
2. The numerical comparator according to claim 1, wherein: The memory cell is a DRAM cell.
3. The numerical comparator according to claim 1, wherein: The first number and the second number are binary numbers.
4. The numerical comparator according to claim 3, wherein: The word lines of the first word line group and the word lines of the second word line group form a sequential sequence.
5. The numerical comparator according to claim 3, wherein: The source or drain of the transistor corresponding to the first digit is connected to the bit line in parallel, and the source or drain of the transistor corresponding to the second digit is connected to the complementary bit line in parallel.
6. The numerical comparator according to claim 5, characterized in that The storage unit has a 1T1C structure.
7. The numerical comparator according to claim 6, wherein: The memory cell further includes a precharge circuit.
8. A B2S converter, characterized in that: The B2S converter comprises the numerical comparator as claimed in claim 1 , wherein the first number is an input binary number and the second number is a converted random number.
9. A random computing integrated circuit, characterized in that: Comprising the B2S converter as claimed in claim 8.
10. A numerical comparator design method, characterized in that: The following steps are involved: Based on a DRAM memory cell array, the memory cells of the memory cell array are connected through word lines and bit lines to form an array. A first number to be compared is input, and the number is stored in memory cells of different word lines on the same bit line in bit order, wherein the different word lines on the same bit line form a first word line group. A second number to be compared is input and stored in a memory cell of a word line different from any word line in the first word line group on the same complementary bit line corresponding to the bit line in bit order, wherein the word lines different from any word line in the first word line group constitute a second word line group. The bit line and the complementary bit line are connected to the input terminal of the sense amplifier corresponding to the bit line.
Citation Information
Patent Citations
Integrated semi-conductor memory circuit
EP0533996A1
Random number generation in ferroelectric random access memory (FRAM)
US20170277459A1
Cited By
In-memory B2S three-stage streamlined conversion method, integrated circuit and electronic equipment
CN122332348A